HomeDocs

Claude Code Proxy: Switch your client in three lines

Connect your existing client to an uncensored LLM proxy by updating the base URL and API key. No code changes required for standard chat completions, streaming, or tool calling.

Base URL
https://api.claudecodeproxyapi.com/v1
Model
uncensored
Get API key

Prerequisites: Install Dependencies

Before making requests, ensure your environment has the necessary tools. For command-line testing, you need curl. For Python projects, install the official OpenAI SDK via pip install openai. For Node.js, run npm install openai. These libraries handle the OpenAI-compatible protocol automatically, allowing you to swap the endpoint without altering your application logic.

Configure Base URL and Key

Point your client to our specialized claude code proxy endpoint. Replace the standard OpenAI base URL with our host and provide your unique API key. This single configuration change directs all traffic to our uncensored model server while maintaining full compatibility with standard SDK methods.

curl https://api.claudecodeproxyapi.com/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "uncensored",
    "messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
  }'

Set BASE_URL to https://api.claudecodeproxyapi.com/v1 and pass your key in the Authorization header. This setup works for any client that supports OpenAI-compatible endpoints.

Test with Chat Completions

Verify connectivity with a simple text request. Send a POST request to /v1/chat/completions with the uncensored model ID. This confirms your key is valid and the proxy is routing requests correctly. The response will contain the model's direct answer without standard content filters for lawful adult content.

from openai import OpenAI

client = OpenAI(base_url="https://api.claudecodeproxyapi.com/v1", api_key="YOUR_KEY")

resp = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)

Use the openai Python client to send a message. The code mirrors standard OpenAI usage, ensuring zero friction when switching from other providers.

Enable Streaming Responses

For lower latency and real-time token display, enable streaming. Set stream=True in your client configuration. The API returns Server-Sent Events (SSE) as tokens are generated. This is ideal for CLI interfaces or applications where immediate feedback improves user experience.

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.claudecodeproxyapi.com/v1", apiKey: process.env.API_KEY });

const resp = await client.chat.completions.create({
  model: "uncensored",
  messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);

Implement streaming in Node.js by iterating over the async generator. This approach handles partial responses efficiently, keeping your UI responsive even for longer outputs.

Use Tool Calling Functions

The proxy supports function calling for enhanced agent capabilities. Define your tools in the request payload and the model will return structured JSON for tool execution. This feature works seamlessly with the uncensored model, allowing for robust agent workflows without vendor lock-in.

Manage Your API Key and Limits

Your account includes one API key, which can be regenerated at any time. Note that regeneration revokes the old key immediately. Monitor your usage to avoid interruptions; requests return a 402 error when credit is insufficient and a 429 error if you exceed 300 requests per minute. The context window is 64,000 tokens for prompt plus completion.

Streaming

stream = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Tell the story in second person."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

What the API supports

The numbers below are the real limits of this API, not marketing. Compare them with what your app needs.

FeatureSupport
ProtocolOpenAI-compatible: any OpenAI SDK or client works — change the base URL and the key
AuthenticationBearer token in the Authorization header
MethodsPOST /v1/chat/completions · GET /v1/models
Base URLhttps://api.claudecodeproxyapi.com/v1
Model IDuncensored
JSON modeJSON object mode via response_format json_object
Tools / tool callsSupported: tools + tool_choice, tool_calls in the reply (streamed too), tool results as role: tool messages
Max output16,000 tokens max; 2,048 if max_tokens is not set
Other parameterstemperature, top_p, stop, seed and the two penalties are passed through
Max context64,000 tokens, input and output combined
StreamingYes — server-sent events; the last chunk carries token usage
Response headersX-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency
Concurrencyup to 8 in parallel per key
Rate limit300 requests per minute per key
Request sizeup to 8 MB per request
Volume bonus+5% from $50, +10% from $100
How you payprepaid credit, charged by real token usage; errors and refusals are free
Subscriptionpaid credit never expires, no subscription
Priceinput $0.25 / 1M tokens, output $1.00 / 1M tokens
PaymentUSDT (TRC20) or USDC (Base), any whole amount from $10 to $500
Trial credit$0.50 for 7 days, no card
Accountsign in with Google or with e-mail + password
Content policyadult content allowed; sexual content involving minors is refused
Keysone active key per account; a new key replaces the old one

Errors and what to do

Every error is JSON with a type you can switch on. You are never charged for an error.

StatusTypeReason
400bad_requestinvalid JSON, empty messages, bad parameter, or prompt + max_tokens over the window — fix and resend
401missing_key · invalid_key · key_revokedcheck the Authorization header or use your current key
402no_creditbalance is empty — top up, requests resume at once
403content_blockedrefused by the content policy
404not_foundunknown endpoint
413request_too_largerequest body larger than 8 MB
429rate_limited · concurrencyover 300/min or 8 parallel — back off and retry
503upstream_busymodel busy — retry in a few seconds

Questions and answers

Is this an official Claude or OpenAI service?

No, it is an independent service running an open-weight uncensored model. It is compatible with the OpenAI API format but does not serve GPT or Claude models.

What happens if I run out of credit?

Subsequent requests will return a 402 Payment Required error. You can top up your account at any time via crypto (USDT or USDC) to restore service.

Can I use this with the Claude Code CLI?

Yes, by configuring the CLI to use our base URL and API key, you can route requests through this uncensored proxy without modifying the core application code.

Your key is one form away

Create an account, copy the key, change the base URL. That is the whole setup.