Go faster
More tokens, no limits.
Up to 10x more tokens than your Claude or Codex subscription. Never let a subscription limit break your focus and slow you down.
View plansAbsurdly fast LLM inference for AI agents. Compatible with Claude, Codex, Cursor, Pi, Hermes, and more. Up to 10x faster than normal.
Go faster
Up to 10x more tokens than your Claude or Codex subscription. Never let a subscription limit break your focus and slow you down.
View plansGo smarter
Fast inference on the top open-source models. Just as smart as Claude and Codex at a fraction of the cost. And did we mention much faster?
View integrationsOne API
One click to switch between the best open-source models and closed-source models like Claude and Codex. Find the right model for you.
Explore modelsRoute Claude Code, Codex, Grok, OpenCode, Pi, Hermes, OpenClaw, Fx, and Cursor through the Inference.net gateway. One shared key. Native surfaces. Safe reversal.
Native Anthropic /v1/messages. Six model slots for every Claude Code workload.
fast claude onOpenAI-compatible Responses API with a reversible Codex configuration.
fast codex onChat completions with a catalog-backed section for each supported alias.
CLI ADAPTER IN PREVIEWChat completions with Fast Inference models registered as an OpenCode provider.
fast opencode onChat completions with model and provider settings written for Pi Agent.
fast pi onHermes provider configuration is documented while its CLI adapter is in preview.
CLI ADAPTER IN PREVIEWOpenClaw provider routing is documented while its CLI adapter is in preview.
CLI ADAPTER IN PREVIEWFx gateway configuration is documented while its CLI adapter is in preview.
CLI ADAPTER IN PREVIEWCursor base URL routing is documented while its CLI adapter is in preview.
CLI ADAPTER IN PREVIEWKimi, GLM, DeepSeek, Fable, and GPT — routed through Fast Inference. Open the full catalog for every model we serve.
View all models$9 / mo
$10 of Gateway credit, renewed each month.
Pay as you go after the credit. Same rates as going direct.
$49 / mo
$60 of Gateway credit, renewed each month.
Pay as you go after the credit. Same rates as going direct.
Prompts and outputs are never used to train models. Data retention is kept to the minimum needed to operate the service.
Your data never enters a training set.
Every key belongs to one project.
See exactly where tokens and spend go.
Disable a credential in one click.
The same endpoint, every model, every provider. No markup. Answers to the questions teams ask before they switch.
Fast Inference is an OpenAI-compatible API that routes to every model and provider. One key, with no per-request markup.
npm install openai◆baseURL: "https://api.inference.net/v1"◆ship it◆