LLM API is a multi-model API relay for developers: Claude, GPT, Gemini, DeepSeek, GLM, Qwen and Kimi behind one key, natively speaking the Anthropic, OpenAI and Gemini protocols. Subscription or metered billing, direct access, set up in two environment variables.
# one command, fully configured
PS> irm llmapi.pro/setup.ps1 | iex
# or two env vars, manually
$ export ANTHROPIC_BASE_URL=https://llmapi.pro
$ export ANTHROPIC_API_KEY=your-key
$ claude
Works with all Claude Code terminals and IDE plugins
01 — Models
Flagship models from 10+ vendors, channels and pricing synced live
Flagship coding & reasoning · powers Claude Code
claude-opus-4-8 · claude-sonnet-5
All-round · Codex coding agents
gpt-5.6 · codex
Multimodal · long context
gemini-3.1-pro · flash
Best value
Domestic flagship
Tongyi lineup
Long context
Fast multimodal
Realtime
Per request
02 — Billing
Two independent billing modes — run both on one account.
Heavy coding · capped cost
Pay per use · every model
03 — The Problem
One account, one invoice and one SDK per vendor — and per-token bills that add up fast.
Rate limits kill your flow. You hit the cap right when you need it most.
LiteLLM breaks, proxies fail, tool calling doesn't work. Hours wasted debugging.
We solved all three.
04 — How It Works
Three steps. No configuration files. No proxy servers.
Get your API key in 30 seconds. Free tier included.
Add these to your shell and you're done.
Claude Code works exactly the same. No changes needed.
# Point your client at LLM API
$ export ANTHROPIC_BASE_URL=https://llmapi.pro
$ export ANTHROPIC_API_KEY=your-key
# That's it. Start coding.
$ claude
05 — Included
Built for Claude Code, Codex CLI and any SDK.
Anthropic / OpenAI / Gemini protocol compatible: messages, streaming, tool calling, extended thinking — any client connects with zero changes.
Control your Claude Code from phone — scan QR, chat syncs, seamless cross-device. Not available on official Claude.
Track tokens, monitor costs, manage API keys. All in one place.
Multi-provider failover. If one backend goes down, we switch automatically. Your work never stops.
06 — Why Us
Official API charges per token — one feature can cost tens of yuan. Our fixed monthly fee, unlimited tokens, one price no matter how much you use. Plans start at ¥59/mo, while the equivalent official experience runs into the hundreds. Making AI coding affordable.
Multi-node load balancing, automatic failover, smart 429 retry. Your workflow won't be interrupted by backend fluctuations. 24/7 uptime to support your development rhythm.
No VPN, no credit card, no overseas phone number. Two env vars and you're ready. Alipay payment in CNY. Solving the biggest barrier for Chinese developers using frontier models.
07 — Billing
Subscriptions with unlimited tokens; metered API with a balance that never expires.
08 — Voices
Switched from the official API and saved $150 in the first month. Setup took literally two minutes. Everything just works.
Marcus T.
Full-stack Developer
I was burning through my Pro plan in under an hour. Now I code all day without worrying about limits. The tool support is flawless.
Sarah K.
Backend Engineer
Our team of 4 moved to the Team plan. We went from $800/month combined to $99. The failover means we've had zero downtime.
James R.
Engineering Lead
09 — FAQ