Fast open-source models

Fast Inferencefor AI agents

Absurdly fast LLM inference for AI agents. Compatible with Claude, Codex, Cursor, Pi, Hermes, and more. Up to 10x faster than normal.

Go faster

daily spend0% markup
490.4K tokens$4.47

More tokens, no limits.

Up to 10x more tokens than your Claude or Codex subscription. Never let a subscription limit break your focus and slow you down.

View plans

Go smarter

head to headkimi ≈ opus
KIMI-K3-Fast
Opus 5
  • SWE
    72
    74
  • GPQA
    81
    83
  • tok/s
    184
    38
  • $/1M in
    4.5
    5

Smart, fast, & affordable.

Fast inference on the top open-source models. Just as smart as Claude and Codex at a fraction of the cost. And did we mention much faster?

View integrations

One API

model catalogopen + closed
open
  • kimi-k3-fast
  • glm-5.2-fast
  • deepseek-v4-fast
closed
  • opus-5
  • gpt-5.6
  • fable-5

All the best models.

One click to switch between the best open-source models and closed-source models like Claude and Codex. Find the right model for you.

Explore models
Coding agents

Fast inference with any agent.

Route Claude Code, Codex, Grok, OpenCode, Pi, Hermes, OpenClaw, Fx, and Cursor through the Inference.net gateway. One shared key. Native surfaces. Safe reversal.

Models

Fast frontier models.

Kimi, GLM, DeepSeek, Fable, and GPT — routed through Fast Inference. Open the full catalog for every model we serve.

View all models
Pricing

Simple, predictable pricing

Operator

$10 monthly allowance

$9 / mo

$10 of Gateway credit, renewed each month.

Pay as you go after the credit. Same rates as going direct.

  • Works with Claude Code and Codex
  • Automatic provider failover
  • Personal usage meter
  • Cancel any time
Get started

Studio

$60 monthly allowance

$49 / mo

$60 of Gateway credit, renewed each month.

Pay as you go after the credit. Same rates as going direct.

  • Everything in Operator
  • More monthly Gateway allowance
  • Hosted and provider model access
  • Cancel any time
Get started
Private by default

Your data is
Your data.

Prompts and outputs are never used to train models. Data retention is kept to the minimum needed to operate the service.

01No training

Your data never enters a training set.

02Scoped access

Every key belongs to one project.

03Clear usage

See exactly where tokens and spend go.

04Easy revocation

Disable a credential in one click.

Frequently asked questions.

Quick answers to help you get started

faq.sh9 questions

Fast Inference provides access to low-cost, high-performance AI models to power AI applications and coding agents.

Step 01$ npm install -g @inference/fast
Step 02set baseURL: "https://api.inference.net/v1"
Step 03ship it