OpenAI Compatible API: Switch your client in three lines
Switch your client in three lines: change the base URL and API key, then keep your existing code. This drop-in openai compatible api delivers an uncensored llm api experience with transparent per-token pricing.
- Base URL
- https://api.openaicompatibleapi.com/v1
- Model
- uncensored
Prerequisites: Install OpenAI SDK
Before making requests, ensure your development environment has the official OpenAI SDK installed for your preferred language. This is the standard way developers interact with an openai compatible api, ensuring you get consistent typing and error handling. For Python, use pip install openai. For Node.js, run npm install openai. These libraries handle the underlying HTTP details, so you can focus on passing prompts and receiving text outputs.
Configure Base URL and API Key
The core advantage of our service is that you only need to update two configuration values. Point your client to our base URL: https://api.openaicompatibleapi.com/v1. Then, generate your API key on the signup page; it is shown immediately after registration with no credit card required. Set the base_url and api_key in your client initialization. This makes us a direct gpt api alternative for your existing projects without rewriting logic.
Make Your First Request
Send a standard chat completion request using the model ID uncensored. This model is an open-weight large language model tuned to answer without content refusals for lawful adult use. It is not GPT, Claude, or any other vendor's model. You can test connectivity immediately with a simple text prompt. The response will return the generated text in the content field.
curl https://api.openaicompatibleapi.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'
Enable Streaming Responses
For better user experience, enable streaming to receive tokens as they are generated. Set stream=True in your request. The SDK will yield chunks as they arrive, allowing you to display partial results instantly. This is particularly useful for long context windows up to 64,000 tokens. Streaming works identically to standard OpenAI clients, maintaining compatibility with your existing error-handling logic.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
Use Tool Calling (Function Calling)
Our API supports structured tool calling. Define your functions in the tools parameter, and the model will return tool calls in the response. You can then execute those tools and send the results back to the model for further reasoning. This feature works with any OpenAI-compatible client that supports the standard tool calling format. It is ideal for building agents or automating workflows without managing a complex llm proxy.
from openai import OpenAI
client = OpenAI(base_url="https://api.openaicompatibleapi.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)
Check Available Models
Query the GET /v1/models endpoint to verify available models. You will see the uncensored model listed. This confirms your API key and base URL are configured correctly. You can also use this endpoint to debug authentication issues. If the endpoint returns a list, your connection to the ai api for developers is active. No other models are offered, keeping the configuration simple.
Node.js
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.openaicompatibleapi.com/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);What the API supports
If your tool speaks the OpenAI API, these are the details that matter.
| Spec | Value |
|---|---|
| API format | OpenAI-compatible: any OpenAI SDK or client works — change the base URL and the key |
| Endpoints | POST /v1/chat/completions · GET /v1/models |
| Model ID | uncensored |
| API key | Bearer token in the Authorization header |
| Base URL | https://api.openaicompatibleapi.com/v1 |
| Streaming | Yes — server-sent events; the last chunk carries token usage |
| JSON mode | JSON object mode via response_format json_object |
| Function calling | Yes — tools, tool_choice; replies carry tool_calls, also when streaming; send results back as role: tool |
| Sampling parameters | temperature, top_p, stop, seed and the two penalties are passed through |
| Completion length | 16,000 tokens max; 2,048 if max_tokens is not set |
| Context window | 64,000 tokens (prompt + completion together) |
| Headers | X-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency |
| Requests per minute | 300 requests per minute per key |
| Request size | up to 8 MB per request |
| Parallel requests | up to 8 in parallel per key |
| Token prices | $0.25 per 1M input tokens · $1.00 per 1M output tokens |
| Volume bonus | +5% from $50, +10% from $100 |
| Billing | prepaid credit, charged by real token usage; errors and refusals are free |
| Credit expiry | no monthly fee; paid credit does not expire |
| Top-up | crypto: USDT on TRON or USDC on Base, $10–$500, any whole sum |
| Trial credit | $0.50 of credit valid 7 days, no card needed |
| Content policy | uncensored for adults; the only hard rule: no sexual content involving minors |
| Keys | one key per account, regenerate any time (the old one stops working) |
| Sign-in | Google or e-mail and password |
HTTP errors
Every error is JSON with a type you can switch on. You are never charged for an error.
| Code | Type | Meaning |
|---|---|---|
400 | bad_request | invalid JSON, empty messages, bad parameter, or prompt + max_tokens over the window — fix and resend |
401 | missing_key · invalid_key · key_revoked | no key, wrong key, or a key replaced by a newer one |
402 | no_credit | balance is empty — top up, requests resume at once |
403 | content_blocked | sexual content involving minors — refused, not billed |
404 | not_found | only /v1/chat/completions and /v1/models exist |
413 | request_too_large | request body larger than 8 MB |
429 | rate_limited · concurrency | over 300/min or 8 parallel — back off and retry |
503 | upstream_busy | model busy — retry in a few seconds |
Questions and answers
What are the rate limits?
You are limited to 300 requests per minute per API key. The request body size is capped at 8 MB. If you exceed the limit, you will receive a 429 error. You can regenerate your key at any time to reset the association, though the rate limit applies to the key itself.
Why do I get a 401 or 402 error?
A 401 error indicates an invalid or expired API key. A 402 error means your prepaid credit is exhausted. We do not offer a free unlimited tier; you must top up to continue using the uncensored llm api. Credit never expires, so you can top up from $10 by crypto (USDT or USDC).
Is the context window really 64k?
Yes, the model supports a strict 64,000 token context window for both prompt and completion combined. This allows for extensive document processing or long conversations. Ensure your inputs fit within this limit to avoid truncation or errors.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.