Integration Quickstart
Integrate our uncensored LLM API in minutes by swapping your base URL and API key. This quickstart guides you through authentication, chat completions, streaming, and function calling using standard OpenAI-compatible clients.
https://api.llmrouterapi.com/v1uncensored
Authentication and Base URL
Start by creating an account at Get API key. Upon signup, you receive a unique API key immediately. There is no need to enter a credit card for the trial, and prompts are not used for training. Your requests authenticate against the hosted endpoint using this key.
The base URL for all requests is https://api.llmrouterapi.com/v1. This URL is compatible with the official OpenAI SDKs and any standard OpenAI-compatible client. Simply configure your client to point to this base URL and set your header to Authorization: Bearer YOUR_API_KEY. No complex routing logic is required; we handle the infrastructure for you.
First Request: Chat Completions
Send your first request to the POST /v1/chat/completions endpoint. You must specify the model ID as "uncensored". This model is an open-weight large language model tuned to answer without content refusals for lawful adult use. It supports a context window of 100,000 tokens, covering both prompt and completion.
Below is a basic example using curl to demonstrate the standard JSON payload structure.
curl https://api.llmrouterapi.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'
Using the Python SDK
For Python developers, the official openai package works out of the box. Configure the client to use our base URL and your API key. This ensures your existing code works instantly by simply swapping the base URL and API key. The client handles serialization and error parsing automatically.
Ensure you specify the model as uncensored in your initialization or request. This approach eliminates integration friction and allows you to leverage the robust ecosystem of existing LLM tools.
from openai import OpenAI
client = OpenAI(base_url="https://api.llmrouterapi.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)
Node.js Integration
Node.js developers can use the openai npm package or any fetch-based client. Set the baseURL to our endpoint and provide your key in the headers. This method is ideal for server-side applications or edge functions where you need programmatic control over the request lifecycle.
The API supports tool and function calling, allowing you to define custom functions within your chat completions. This enables complex logic and data retrieval directly from the model's response structure.
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.llmrouterapi.com/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);
Streaming Responses
For lower latency and better user experience, enable streaming by setting stream: true in your request. The API returns a Server-Sent Events (SSE) stream, delivering tokens as they are generated. This is particularly useful for chat interfaces where users expect immediate feedback.
Each chunk contains partial content, allowing you to render text progressively. This feature works seamlessly with standard OpenAI-compatible streaming clients, requiring no additional parsing logic for the event structure.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
Limits, Errors, and Model Listing
Monitor your usage and handle errors gracefully. The API enforces a limit of 300 requests per minute per key and an 8 MB request body limit. If your key is invalid, you will receive a 401 error. A 402 status indicates insufficient prepaid credit. A 429 error signals that you have exceeded the rate limit.
You can list available models via GET /v1/models, which will return the uncensored model. Note that this service does not offer embeddings, image, or audio generation. For account management, you can regenerate your API key at any time, which revokes the previous one. Credits do not expire, and you can top up starting at $10.
Specs at a glance
A quick checklist for developers: format, limits, features, billing.
| Parameter | Details |
|---|---|
| API format | OpenAI Chat Completions schema; official openai SDKs work unchanged |
| Model ID | uncensored |
| Base URL | https://api.llmrouterapi.com/v1 |
| Methods | POST /v1/chat/completions · GET /v1/models |
| API key | Authorization: Bearer YOUR_KEY |
| JSON mode | response_format: {"type": "json_object"} |
| Max context | 100,000 tokens (prompt + completion together) |
| Streaming | Yes — server-sent events; the last chunk carries token usage |
| Completion length | 16,000 tokens max; 2,048 if max_tokens is not set |
| Sampling parameters | temperature, top_p, stop, seed, presence_penalty, frequency_penalty |
| Tools / tool calls | Yes — tools, tool_choice; replies carry tool_calls, also when streaming; send results back as role: tool |
| Headers | X-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency |
| Concurrency | up to 8 in parallel per key |
| Rate limit | 300 requests per minute per key |
| Request size | 8 MB request body |
| Subscription | no monthly fee; paid credit does not expire |
| Volume bonus | +5% on $50+, +10% on $100+ |
| Top-up | crypto: USDT on TRON or USDC on Base, $10–$500, any whole sum |
| Token prices | input $0.25 / 1M tokens, output $1.00 / 1M tokens |
| Free trial | $0.50 for 7 days, no card |
| Billing | pay as you go from prepaid credit; nothing is charged for failed or refused requests |
| Content policy | adult content allowed; sexual content involving minors is refused |
| Keys | one active key per account; a new key replaces the old one |
| Sign-in | sign in with Google or with e-mail + password |
Error reference
Errors come back as JSON with a stable type; failed and refused requests are not billed.
| HTTP | Type | What to do |
|---|---|---|
400 | bad_request | invalid JSON, empty messages, bad parameter, or prompt + max_tokens over the window — fix and resend |
401 | missing_key · invalid_key · key_revoked | no key, wrong key, or a key replaced by a newer one |
402 | no_credit | out of credit; add credit and retry |
403 | content_blocked | sexual content involving minors — refused, not billed |
404 | not_found | only /v1/chat/completions and /v1/models exist |
413 | request_too_large | body over 8 MB |
429 | rate_limited · concurrency | over 300/min or 8 parallel — back off and retry |
503 | upstream_busy | model busy — retry in a few seconds |
Questions and answers
What is the pricing structure?
We offer pay-as-you-go prepaid credit with no monthly fees. The cost is $0.25 per 1M input tokens and $1.00 per 1M output tokens. Credits never expire, and you can top up via crypto (USDT or USDC) starting at $10.
Does this API support tool calling?
Yes, the <code>POST /v1/chat/completions</code> endpoint supports function and tool calling. You can define custom functions in your request, and the model will return structured JSON data when appropriate.
How many tokens does the context window support?
The model supports a context window of 100,000 tokens. This total includes both the prompt (input) and the completion (output). Ensure your combined token count stays within this limit for optimal performance.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.