Switch your client in three lines
Drop your kimi api key into any OpenAI-compatible client to start generating text from an uncensored model. This guide covers the three lines you need to switch your endpoint.
Prerequisites: API Key
To use this uncensored api, you need an API key from kimiapi.cc. Visit the Get API key page and sign in with Google or an email and password. The key is displayed immediately and is the only credential required for authentication.
Each account holds one active key; generating a new one replaces the previous key. Your key grants access to the base URL and defines your rate limits. Keep this key secure, as it is tied directly to your prepaid credit balance.
Basic Chat Completion
The core endpoint is POST /v1/chat/completions. Point your client to the base URL and provide the model ID uncensored. This returns a standard text response. You can adjust parameters like temperature or top_p to control creativity.
curl https://api.kimiapi.cc/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'Ensure your request body includes the messages array with your prompt. The server processes the input and returns the generated text in the content field of the first choice.
Python SDK
Use the official OpenAI Python SDK to interact with the endpoint. Set the base_url to our API address and provide your kimi api key in the api_key field. This allows you to use familiar Python methods for chat completions.
from openai import OpenAI
client = OpenAI(base_url="https://api.kimiapi.cc/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)Instantiate the client with your credentials and call chat.completions.create(). The response object provides access to the generated content, token usage, and other metadata. This approach ensures you are using the same client code you would use for other OpenAI-compatible services.
Node SDK
For JavaScript developers, the @anthropic-ai/sdk or official openai Node package works with minor configuration. Configure the baseURL to point to our endpoint and set the apiKey to your key. This enables seamless integration into Node.js applications.
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.kimiapi.cc/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);Create a completion request by passing the model ID uncensored and your messages. The SDK handles the HTTP requests and parses the JSON response. This method is ideal for server-side rendering or API integrations where you need programmatic access to the model.
Streaming Responses
Enable streaming by setting stream: true in your request. The API returns a Server-Sent Events (SSE) stream, delivering tokens as they are generated. This reduces perceived latency for interactive applications.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)Each chunk contains partial content. The final chunk includes the complete token usage statistics. Handle the stream events to accumulate the full response. This feature is essential for chat interfaces where users expect to see text appear in real time.
Limits, Errors & Context
Each key is limited to 300 requests per minute and 8 concurrent requests. The request body must not exceed 8 MB. If your key is invalid, you receive a 401 error. A 402 error indicates insufficient prepaid credit. A 429 error signals that you have exceeded the rate limit.
The context window is 64,000 tokens total, including prompt and completion. If you do not specify max_tokens, the output is capped at 2,048 tokens. Adjust these limits based on your application's needs to ensure reliable performance.
Questions and answers
What happens if I exceed my rate limit?
The API returns a 429 Too Many Requests error. You must wait until the rate limit resets, which occurs on a sliding window basis per key. Exceeding the limit does not incur additional charges.
How is my credit charged?
Credit is charged based on actual token usage for both input and output. Errors and refusals do not consume credit. Prepaid credit never expires, so you can top up as needed without monthly fees.
Can I use this API for commercial projects?
Yes, the uncensored model allows lawful adult, fictional, and controversial topics without content refusals. The only hard limit is the prohibition of sexual content involving minors. Check the specific terms for your use case.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.