Go faster
More tokens, no limits.
Up to 10x more tokens than your Claude or Codex subscription. Never let a subscription limit break your focus and slow you down.
View plansAbsurdly fast LLM inference for AI agents. Compatible with Claude, Codex, Cursor, Pi, Hermes, and more. Up to 10x faster than normal.
Go faster
Up to 10x more tokens than your Claude or Codex subscription. Never let a subscription limit break your focus and slow you down.
View plansGo smarter
Fast inference on the top open-source models. Just as smart as Claude and Codex at a fraction of the cost. And did we mention much faster?
View integrationsOne API
One click to switch between the best open-source models and closed-source models like Claude and Codex. Find the right model for you.
Explore modelsRoute Claude Code, Codex, Grok, OpenCode, Pi, Hermes, OpenClaw, Fx, and Cursor through the Inference.net gateway. One shared key. Native surfaces. Safe reversal.
Native Anthropic /v1/messages. Six model slots for every Claude Code workload.
fast claude onOpenAI-compatible Responses API with a reversible Codex configuration.
fast codex onChat completions with a catalog-backed section for each supported alias.
fast grok onChat completions with Fast Inference models registered as an OpenCode provider.
fast opencode onChat completions with model and provider settings written for Pi Agent.
fast pi onHermes provider configuration is documented while its CLI adapter is in preview.
CLI ADAPTER IN PREVIEWOpenClaw provider routing is documented while its CLI adapter is in preview.
CLI ADAPTER IN PREVIEWFx gateway configuration is documented while its CLI adapter is in preview.
CLI ADAPTER IN PREVIEWCursor base URL routing is documented while its CLI adapter is in preview.
CLI ADAPTER IN PREVIEWKimi, GLM, DeepSeek, Fable, and GPT — routed through Fast Inference. Open the full catalog for every model we serve.
View all models$9 / mo
$10 of Gateway credit, renewed each month.
Pay as you go after the credit. Same rates as going direct.
$49 / mo
$60 of Gateway credit, renewed each month.
Pay as you go after the credit. Same rates as going direct.
Prompts and outputs are never used to train models. Data retention is kept to the minimum needed to operate the service.
Your data never enters a training set.
Every key belongs to one project.
See exactly where tokens and spend go.
Disable a credential in one click.
Quick answers to help you get started
Fast Inference provides access to low-cost, high-performance AI models to power AI applications and coding agents.
$ npm install -g @inference/fastset baseURL: "https://api.inference.net/v1"ship it