Get API key

How to Use the Uncensored API for NSFW LLM Integration

The uncensored api provides a straightforward, OpenAI-compatible interface for integrating large language models that do not filter lawful adult, roleplay, or controversial content. By offering a single high-performance model with transparent prepaid pricing and crypto payments, it eliminates the friction of subscription locks and complex routing for developers building NSFW AI applications.

Updated

Key points

  • The API strictly follows the OpenAI chat-completions standard, allowing immediate integration with existing SDKs.
  • Pricing is prepaid and transparent at $0.25 per million input tokens and $1.00 per million output tokens.
  • The model supports a 64,000-token context window and includes streaming, function calling, and JSON mode.
  • Payments are processed exclusively via USDT (TRC20) or USDC (Base), with no monthly fees or subscription locks.

What is the Uncensored API?

The uncensored ai api is a hosted service designed for developers who need a reliable text-generation endpoint without content filters. Unlike platforms that aggregate multiple models or provide web interfaces, this service focuses on delivering a single, high-performance uncensored large language model. The model is open-weight and tuned to answer without refusals for lawful adult use, making it ideal for roleplay, creative writing, and NSFW applications.

The core functionality is simple: text in, text out. You send a prompt, and the model returns a completion. There are no embeddings, image generation, or fine-tuning capabilities to manage. This simplicity reduces integration complexity. You interact with the API via standard HTTP requests, using the base URL https://api.uncensoredaichat.cc/v1. The endpoint supports the same request structure as OpenAI's chat-completions endpoint, meaning you can often swap the base URL and API key in your existing code without major refactoring.

Privacy is a key design principle. Your prompts are not used for training data, and account creation requires only an email address. The model refuses only one hard content limit: sexual content involving minors. Everything else is fair game, including controversial topics, security research, and explicit fiction, as long as it is lawful.

Why Choose an Uncensored LLM?

Standard LLMs are optimized for broad appeal, which often means they are overly cautious. They may refuse to generate explicit content, critique political figures, or explore niche roleplay scenarios due to built-in safety filters. An uncensored LLM removes these arbitrary boundaries. This is crucial for applications where creative freedom or specific tone requirements are paramount.

  • Roleplay Fidelity: Characters can express emotions, desires, and conflicts without being cut off by a safety filter.
  • Content Flexibility: Developers can build apps for adult audiences without needing complex prompt engineering to bypass refusals.
  • Predictable Behavior: A single model means consistent behavior. You do not have to worry about the model switching personalities or filters when traffic is routed to different vendors.

For indie hackers and developers, this predictability is valuable. You know exactly what you are getting: a model that will generate the text you ask for, provided it meets the basic legal criteria. This reduces debugging time spent on why the model refused a specific prompt.

OpenAI Compatibility Explained

The term "uncensored openrouter" often comes up when discussing aggregation services, but those platforms route requests to various vendors. In contrast, this API serves a single, consistent model. However, it maintains strict compatibility with the OpenAI API structure. This means you can use the official OpenAI SDKs or any client that supports the OpenAI format.

The endpoint is POST /v1/chat/completions. You send a JSON body with messages, and the API returns a JSON response. The model ID is "uncensored". This simplicity allows for easy integration. For example, you can change the base URL in your Python or Node.js client and start sending requests immediately.

FeatureSupportNotes
Streaming (SSE)YesToken usage is in the last chunk.
Function CallingYesSupports tools and tool_choice.
JSON ModeYesUse response_format: json_object.
Top PYesStandard parameter supported.
TemperatureYesStandard parameter supported.

This compatibility means you do not need to learn a new protocol. You use the tools you already know. The only difference is the base URL and the model ID.

Integration Steps

Integrating the uncensored api is straightforward. First, sign up for an account using "Continue with Google" or by creating an email and password. You will receive an API key immediately. This key is used to authenticate your requests.

Next, configure your client. Set the base URL to https://api.uncensoredaichat.cc/v1. Set the API key in the Authorization header. Then, send a POST request to /v1/chat/completions with the model ID "uncensored".

Here is a basic example of how to structure the request:

from openai import OpenAI

client = OpenAI(base_url="https://api.uncensoredaichat.cc/v1", api_key="YOUR_KEY")

resp = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)

You can also use curl for quick testing:

curl https://api.uncensoredaichat.cc/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "uncensored",
    "messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
  }'

The response will contain the generated text in the choices array. You can then parse this text and display it in your application. The process is identical to using OpenAI's API, which reduces the learning curve for developers who are already familiar with the format.

Streaming and Performance

Streaming is essential for a good user experience, especially in chat applications. The API supports Server-Sent Events (SSE) for streaming responses. This allows you to display tokens as they are generated, reducing the perceived latency for the user.

When using streaming, the API sends multiple chunks of data. Each chunk contains a partial response. The final chunk includes the token usage statistics, so you can track your credit consumption accurately. This is important for prepaid accounts where you need to know when your balance is running low.

The API is designed for high performance. You can send up to 300 requests per minute per key. This is sufficient for most applications. If you need more concurrency, you can generate multiple keys, but note that only one key is active per account at a time. A new key replaces the old one.

Streaming also helps with error handling. If the connection drops, you can resume from the last chunk. This makes your application more robust. The uncensored ai api ensures that streaming is reliable and consistent.

Handling Context Windows

The context window is the amount of text the model can remember in a single request. The uncensored api supports a 64,000-token context window. This is the sum of the prompt and the completion. This large window allows for long conversations or large documents without losing context.

The maximum output per request is 16,000 tokens. If you do not specify the max_tokens parameter, the default is 2,048 tokens. This is useful for short responses, but for longer content, you should set max_tokens explicitly.

Managing context efficiently is key to performance. If your conversation grows beyond the context window, you may need to truncate older messages. This is a common pattern in chat applications. The API does not handle this for you; you must manage the message history in your client.

For example, if you are building a roleplay bot, you might keep the last 50 messages in the context. This ensures the model has enough history to maintain continuity without exceeding the limit. The 64,000-token window provides ample space for most use cases.

Pricing and Credits

The pricing model is prepaid and transparent. You pay for what you use. The cost is $0.25 per million input tokens and $1.00 per million output tokens. Errors and refusals are free, so you do not pay for failed requests.

Credits never expire. This means you can top up once and use the balance over time. There are no monthly fees or subscription locks. This is ideal for indie hackers who want to control their costs.

Payments are made via crypto only. You can top up with USDT (TRC20) or USDC (Base). The minimum top-up is $10, and the maximum is $500. If you top up $50, you get a 5% bonus. If you top up $100, you get a 10% bonus. This bonus credit is added to your account immediately.

New accounts get $0.50 of trial credit valid for 7 days. No card is needed. This allows you to test the API before committing to a payment. The trial is one per person.

Common Use Cases

The uncensored api is versatile. It can be used for a variety of applications, from roleplay bots to content generation tools.

  • NSFW Roleplay: The model does not filter adult content, making it perfect for character-driven games.
  • Creative Writing: Writers can generate stories without worrying about content filters blocking their ideas.
  • Data Generation: The model can generate large datasets of text for training or analysis.
  • API Wrapper Services: Developers can build their own services on top of this API, offering custom features or models.

For example, a developer building a dating app chatbot might use this API to generate responses that are more human-like and less restricted. Another developer might use it to generate adult-themed comic strips by providing text prompts to an image generation model.

The key is flexibility. The API provides the raw power of an uncensored LLM, leaving the creative decisions to you.

Questions and answers

Is the model the same as GPT or Claude?

No, the model is an open-weight model run on our own servers. It is not GPT, Claude, Gemini, Grok, DeepSeek, Qwen, Llama, or any other vendor's model. It is tuned specifically for uncensored responses.

Can I pay with a credit card?

No, payments are crypto only. You can top up with USDT (TRC20) or USDC (Base). There is no PayPal or bank transfer option.

What happens if I make a mistake in my request?

Errors are free. You do not pay for failed requests or refusals. If you have a billing issue, such as a double charge, you can contact support through the Support page.

Is there a limit on how many requests I can make?

Yes, the limit is 300 requests per minute per key. You can also have 8 requests at the same time per key. The request body must be under 8 MB.

Your key is one form away

Create an account, copy the key, change the base URL. That is the whole setup.