Fast open-source models

Fast Inferencefor AI agents

Absurdly fast LLM inference for AI agents. Compatible with Claude, Codex, Cursor, Pi, Hermes, and more. Up to 10x faster than normal.

Go faster

daily spend0% markup
490.4K tokens$4.47

More tokens, no limits.

Up to 10x more tokens than your Claude or Codex subscription. Never let a subscription limit break your focus and slow you down.

View plans

Go smarter

head to headkimi ≈ opus
KIMI-K3-Fast
Opus 5
  • SWE
    72
    74
  • GPQA
    81
    83
  • tok/s
    184
    38
  • $/1M
    0.6
    15

Smart, fast, & affordable.

Fast inference on the top open-source models. Just as smart as Claude and Codex at a fraction of the cost. And did we mention much faster?

View integrations

One API

model catalogopen + closed
open
  • kimi-k3-fast
  • glm-5.2-fast
  • deepseek-v4-fast
closed
  • opus-5
  • gpt-5.6
  • fable-5

All the best models.

One click to switch between the best open-source models and closed-source models like Claude and Codex. Find the right model for you.

Explore models
Coding agents

Fast inference with any agent.

Route Claude Code, Codex, Grok, OpenCode, Pi, Hermes, OpenClaw, Fx, and Cursor through the Inference.net gateway. One shared key. Native surfaces. Safe reversal.

Models

Fast frontier models.

Kimi, GLM, DeepSeek, Fable, and GPT — routed through Fast Inference. Open the full catalog for every model we serve.

View all models
Pricing

Simple, predictable pricing

Operator

$10 monthly allowance

$9 / mo

$10 of Gateway credit, renewed each month.

Pay as you go after the credit. Same rates as going direct.

  • Claude Code and OpenClaw ready
  • Automatic provider failover
  • Personal usage meter
  • Cancel any time
Get started

Studio

$60 monthly allowance

$49 / mo

$60 of Gateway credit, renewed each month.

Pay as you go after the credit. Same rates as going direct.

  • Everything in Operator
  • More monthly Gateway allowance
  • Hosted and provider model access
  • Cancel any time
Get started
Private by default

Your data is
Your data.

Prompts and outputs are never used to train models. Data retention is kept to the minimum needed to operate the service.

01No training

Your data never enters a training set.

02Scoped access

Every key belongs to one project.

03Clear usage

See exactly where tokens and spend go.

04Easy revocation

Disable a credential in one click.

Frequently asked questions.

The same endpoint, every model, every provider. No markup. Answers to the questions teams ask before they switch.

faq.sh7 questions

Fast Inference is an OpenAI-compatible API that routes to every model and provider. One key, with no per-request markup.

npm install openaibaseURL: "https://api.inference.net/v1"ship it