Get API key

Uncensored, OpenAI-compatible chat API

base_url
https://api.opensourcellmapi.com/v1
model
uncensored

Integrate the Open Source LLM API in Minutes

Use our OpenAI-compatible chat completions API for uncensored inference with a single model, transparent pricing, and standard SDK integration.

Base URL and Authentication

Connect to the open source llm api by pointing your client to our base URL: https://api.opensourcellmapi.com/v1. Authentication is handled via a standard API key passed in the Authorization header. You receive this key immediately after signing up on the Get API key page with just an email and password. No phone number or credit card is required to start. If you need to rotate credentials, you can regenerate your key at any time, which immediately revokes the previous one. This ensures you maintain control over your access without relying on a complex dashboard interface.

Send Your First Request

Make a standard POST request to /v1/chat/completions. The endpoint expects a JSON body with a model field set to uncensored and a messages array. This endpoint supports the standard chat format, allowing you to switch from OpenAI or other providers with minimal code changes. Below is a basic cURL example to verify connectivity and receive a text response.

curl https://api.opensourcellmapi.com/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "uncensored",
    "messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
  }'

Python SDK Integration

Using the official OpenAI Python library simplifies integration. Install the package via pip, then initialize the client with your API key and our base URL. Set the model to uncensored and call the chat completions method. This approach lets you leverage existing Python LLM applications with zero modification to the inference logic. The library handles serialization and error parsing automatically.

from openai import OpenAI

client = OpenAI(base_url="https://api.opensourcellmapi.com/v1", api_key="YOUR_KEY")

resp = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)

Node.js SDK Integration

For JavaScript or TypeScript projects, use the OpenAI Node.js SDK. Install the package and configure the client with the same base URL and API key. The uncensored model ID ensures you are hitting our uncensored inference endpoint. This allows you to integrate unrestricted LLM capabilities directly into your web or server-side applications while maintaining compatibility with standard OpenAI client code.

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.opensourcellmapi.com/v1", apiKey: process.env.API_KEY });

const resp = await client.chat.completions.create({
  model: "uncensored",
  messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);

Enable Streaming (SSE)

For lower latency and real-time user experiences, enable streaming by setting stream: true. The API returns Server-Sent Events (SSE), allowing you to process tokens as they are generated. This is ideal for chat interfaces where typing indicators improve perceived performance. The response structure remains consistent with standard OpenAI streaming formats, ensuring your event handlers work without custom parsing logic.

stream = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Tell the story in second person."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Limits, Errors, and Context Window

Your API key is limited to 300 requests per minute and an 8 MB request body. The context window supports up to 64,000 tokens for the combined prompt and completion. If you receive a 401, your key is invalid or expired. A 402 indicates insufficient prepaid credit; you can top up from $10 via crypto (USDT or USDC). A 429 means you have exceeded the rate limit. All credits are prepaid and never expire, with bonuses available for larger top-ups.

Under the hood: specs

A quick checklist for developers: format, limits, features, billing.

SpecValue
CompatibilityOpenAI Chat Completions schema; official openai SDKs work unchanged
API keyBearer token in the Authorization header
Modeluncensored
EndpointsPOST /v1/chat/completions · GET /v1/models
Base URLhttps://api.opensourcellmapi.com/v1
Completion lengthup to 16,000 tokens per request (default 2,048)
Function callingYes — tools, tool_choice; replies carry tool_calls, also when streaming; send results back as role: tool
Other parameterstemperature, top_p, stop, seed, presence_penalty, frequency_penalty
JSON moderesponse_format: {"type": "json_object"}
Context window64,000 tokens, input and output combined
StreamingYes — server-sent events; the last chunk carries token usage
Parallel requests8 requests at the same time per key
Response headersX-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency
Rate limit300 requests per minute per key
Max body8 MB request body
Paymentcrypto: USDT on TRON or USDC on Base, $10–$500, any whole sum
Trial credit$0.50 for 7 days, no card
Bonus credit+5% on $50+, +10% on $100+
Credit expiryno monthly fee; paid credit does not expire
Billingprepaid credit, charged by real token usage; errors and refusals are free
Token pricesinput $0.25 / 1M tokens, output $1.00 / 1M tokens
Key managementone key per account, regenerate any time (the old one stops working)
Contentuncensored for adults; the only hard rule: no sexual content involving minors
Sign-insign in with Google or with e-mail + password

When a request fails

The type field is stable, the message is for humans. Errors cost nothing.

StatusTypeReason
400bad_requestmalformed request or too long for the context window
401missing_key · invalid_key · key_revokedcheck the Authorization header or use your current key
402no_creditout of credit; add credit and retry
403content_blockedrefused by the content policy
404not_foundonly /v1/chat/completions and /v1/models exist
413request_too_largebody over 8 MB
429rate_limited · concurrencyslow down: rate or parallel limit reached
503upstream_busymodel busy — retry in a few seconds

Questions and answers

Does this API support function calling?

Yes, the <code>/v1/chat/completions</code> endpoint supports tool and function calling. You can define tools in your request and the model will return structured JSON for tool use, compatible with standard OpenAI client libraries.

What is the pricing for this uncensored LLM?

Pricing is $0.25 per 1M input tokens and $1.00 per 1M output tokens. There are no monthly fees or subscriptions. You pay for what you use from your prepaid credit balance.

Is the uncensored model the same as GPT-4 or Grok?

No. The model ID <code>uncensored</code> refers to an open-weight model hosted on our own GPU servers. It is not GPT, Claude, Gemini, Grok, or any other vendor's model. It is tuned specifically for fewer content refusals while remaining compatible with standard chat APIs.

Your key is one form away

Create an account, copy the key, change the base URL. That is the whole setup.

Get API key