Integrate the Open Source LLM API in Minutes
Use our OpenAI-compatible chat completions API for uncensored inference with a single model, transparent pricing, and standard SDK integration.
Base URL and Authentication
Connect to the open source llm api by pointing your client to our base URL: https://api.opensourcellmapi.com/v1. Authentication is handled via a standard API key passed in the Authorization header. You receive this key immediately after signing up on the Get API key page with just an email and password. No phone number or credit card is required to start. If you need to rotate credentials, you can regenerate your key at any time, which immediately revokes the previous one. This ensures you maintain control over your access without relying on a complex dashboard interface.
Send Your First Request
Make a standard POST request to /v1/chat/completions. The endpoint expects a JSON body with a model field set to uncensored and a messages array. This endpoint supports the standard chat format, allowing you to switch from OpenAI or other providers with minimal code changes. Below is a basic cURL example to verify connectivity and receive a text response.
curl https://api.opensourcellmapi.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'
Python SDK Integration
Using the official OpenAI Python library simplifies integration. Install the package via pip, then initialize the client with your API key and our base URL. Set the model to uncensored and call the chat completions method. This approach lets you leverage existing Python LLM applications with zero modification to the inference logic. The library handles serialization and error parsing automatically.
from openai import OpenAI
client = OpenAI(base_url="https://api.opensourcellmapi.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)
Node.js SDK Integration
For JavaScript or TypeScript projects, use the OpenAI Node.js SDK. Install the package and configure the client with the same base URL and API key. The uncensored model ID ensures you are hitting our uncensored inference endpoint. This allows you to integrate unrestricted LLM capabilities directly into your web or server-side applications while maintaining compatibility with standard OpenAI client code.
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.opensourcellmapi.com/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);
Enable Streaming (SSE)
For lower latency and real-time user experiences, enable streaming by setting stream: true. The API returns Server-Sent Events (SSE), allowing you to process tokens as they are generated. This is ideal for chat interfaces where typing indicators improve perceived performance. The response structure remains consistent with standard OpenAI streaming formats, ensuring your event handlers work without custom parsing logic.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
Limits, Errors, and Context Window
Your API key is limited to 300 requests per minute and an 8 MB request body. The context window supports up to 64,000 tokens for the combined prompt and completion. If you receive a 401, your key is invalid or expired. A 402 indicates insufficient prepaid credit; you can top up from $10 via crypto (USDT or USDC). A 429 means you have exceeded the rate limit. All credits are prepaid and never expire, with bonuses available for larger top-ups.
Under the hood: specs
A quick checklist for developers: format, limits, features, billing.
| Spec | Value |
|---|---|
| Compatibility | OpenAI Chat Completions schema; official openai SDKs work unchanged |
| API key | Bearer token in the Authorization header |
| Model | uncensored |
| Endpoints | POST /v1/chat/completions · GET /v1/models |
| Base URL | https://api.opensourcellmapi.com/v1 |
| Completion length | up to 16,000 tokens per request (default 2,048) |
| Function calling | Yes — tools, tool_choice; replies carry tool_calls, also when streaming; send results back as role: tool |
| Other parameters | temperature, top_p, stop, seed, presence_penalty, frequency_penalty |
| JSON mode | response_format: {"type": "json_object"} |
| Context window | 64,000 tokens, input and output combined |
| Streaming | Yes — server-sent events; the last chunk carries token usage |
| Parallel requests | 8 requests at the same time per key |
| Response headers | X-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency |
| Rate limit | 300 requests per minute per key |
| Max body | 8 MB request body |
| Payment | crypto: USDT on TRON or USDC on Base, $10–$500, any whole sum |
| Trial credit | $0.50 for 7 days, no card |
| Bonus credit | +5% on $50+, +10% on $100+ |
| Credit expiry | no monthly fee; paid credit does not expire |
| Billing | prepaid credit, charged by real token usage; errors and refusals are free |
| Token prices | input $0.25 / 1M tokens, output $1.00 / 1M tokens |
| Key management | one key per account, regenerate any time (the old one stops working) |
| Content | uncensored for adults; the only hard rule: no sexual content involving minors |
| Sign-in | sign in with Google or with e-mail + password |
When a request fails
The type field is stable, the message is for humans. Errors cost nothing.
| Status | Type | Reason |
|---|---|---|
400 | bad_request | malformed request or too long for the context window |
401 | missing_key · invalid_key · key_revoked | check the Authorization header or use your current key |
402 | no_credit | out of credit; add credit and retry |
403 | content_blocked | refused by the content policy |
404 | not_found | only /v1/chat/completions and /v1/models exist |
413 | request_too_large | body over 8 MB |
429 | rate_limited · concurrency | slow down: rate or parallel limit reached |
503 | upstream_busy | model busy — retry in a few seconds |
Questions and answers
Does this API support function calling?
Yes, the <code>/v1/chat/completions</code> endpoint supports tool and function calling. You can define tools in your request and the model will return structured JSON for tool use, compatible with standard OpenAI client libraries.
What is the pricing for this uncensored LLM?
Pricing is $0.25 per 1M input tokens and $1.00 per 1M output tokens. There are no monthly fees or subscriptions. You pay for what you use from your prepaid credit balance.
Is the uncensored model the same as GPT-4 or Grok?
No. The model ID <code>uncensored</code> refers to an open-weight model hosted on our own GPU servers. It is not GPT, Claude, Gemini, Grok, or any other vendor's model. It is tuned specifically for fewer content refusals while remaining compatible with standard chat APIs.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.
Get API key