Inference pricing

Use the models.Skip the markup.

Choose a monthly inference credit, then keep running at the listed model rates. Every request, token, and dollar stays visible in one project ledger.

Plans

Pick your monthly credit.

Start with the allowance that fits your workload. Usage beyond the included credit continues at the listed model rate.

Operator

$10 monthly allowance

$9 / mo

$10 of Gateway credit, renewed each month.

Pay as you go after the credit. Same rates as going direct.

  • Claude Code and OpenClaw ready
  • Automatic provider failover
  • Personal usage meter
  • Cancel any time
Get started

Studio

$60 monthly allowance

$49 / mo

$60 of Gateway credit, renewed each month.

Pay as you go after the credit. Same rates as going direct.

  • Everything in Operator
  • More monthly Gateway allowance
  • Hosted and provider model access
  • Cancel any time
Get started
How billing moves

One balance. One clear trail.

01

Fund

Your plan arrives as inference credit every month.

02

Run

Every model request draws from one shared balance.

03

Inspect

Track spend and tokens by model from the Usage dashboard.

Frequently asked questions.

The same endpoint, every model, every provider. No markup. Answers to the questions teams ask before they switch.

faq.sh7 questions

Fast Inference is an OpenAI-compatible API that routes to every model and provider. One key, with no per-request markup.

npm install openaibaseURL: "https://api.inference.net/v1"ship it