Section
Tokens, pricing & quotas

Tokens, pricing & quotas

How per-token billing works, how tokens are counted, and how workspace quotas work.

Serverless inference is billed per token. This page covers what you are billed on, how tokens are counted, and how workspace quotas behave. The rates themselves are published on the pricing page.

#What you're billed on

Per-token (token-based pricing) — pay-as-you-go per 1M tokens, priced by model. Input and output tokens are metered separately and carry their own rates, so a request costs the two added together, not one rate applied to a combined count. Best fit for variable traffic and quick prototyping. It applies to every serverless call, /v1/chat/completions and /v1/embeddings alike.

The rest of this page covers per-token billing and the monthly quotas that govern it.

#What a token is

A token is roughly a piece of a word — about 4 characters or 0.75 English words on average. As a rough rule of thumb, "Hello, world!" lands around 4 tokens on a typical OSS chat tokenizer (Qwen-family and OpenAI-shape vocabs sit in that range). Different languages and code tokenize differently; for non-English text or heavily formatted code, expect more tokens per character. The exact tokenization depends on the model. The response always tells you the count used via usage.prompt_tokens and usage.completion_tokens.

#Current rates

Every model is priced individually, and rates are the same for every workspace — there's no per-account rate card to look up.

Current per-token rates are published on the pricing page — every model quoted for serverless inference, per 1M tokens, with input and output priced separately. Embedding models are metered the same way, per token, and their rates are quoted on request.

#How spend is calculated

Per-request cost follows one formula:

Text
cost = (input_tokens / 1_000_000 × input_rate_per_1M) + (output_tokens / 1_000_000 × output_rate_per_1M)

Input is everything you send: system prompt, conversation history, the new user message, tool definitions, schemas. Output is what the model generates.

#Estimating cost before you call

For a pre-call estimate, run your prompt through a tokenizer locally. The tiktoken library covers OpenAI-shape tokenizers — most OSS chat models in the same vocab family produce counts within a few percent of the actual billed total.

Python
import tiktoken
enc = tiktoken.get_encoding("cl100k_base")  # an OpenAI-shape encoding
n_tokens = len(enc.encode("your prompt here"))

This is an estimate — actual billing always uses the model's own tokenizer. If the served model ships its own tokenizer (e.g. a Qwen variant), expect small drift from the tiktoken count.

#Quotas

Quotas are per workspace and monthly. One resource is tracked:

ResourceWhat countsReset
TokensSum of prompt + completion tokens across all per-token API calls.Monthly, on the workspace reset date.

For the quota you can see how much you're allowed this month, how much you've used, how much is left, that figure as a percentage, and the date the count goes back to zero. Past 80% it's flagged as approaching its limit; at 100% it's flagged as over.

#Enforcement vs. tracking-only

A workspace-level enforcement setting controls behavior at the limit. With enforcement on (the default), requests past the limit are rejected. With it off, usage is still counted and shown under a Tracking only banner, without blocking traffic — useful for audit and pre-production workspaces.

#What happens at the limit

When enforcement is on and the token quota hits its monthly limit, further calls fail with 429 Too Many Requests. That is the same status a per-minute rate limit returns, but it does not clear the same way: a rate limit passes in seconds, while a quota holds until the workspace reset date. So if backing off doesn't help, you're over quota, not over rate — reduce spend or request a higher limit. Watch for the approaching-limit flag described above — it appears once you pass 80% — well before the hard cap.

Per-key controls (scopes, RPM/TPM) are not implemented today; an optional per-key expiry is — see authentication → workspace-level controls.

#Reducing spend

  • Use the smallest model that does the job. Profile against your real workload before reaching for a pricier model — premium-by-default burns spend you didn't need.
  • Cap max_tokens on every request. Runaway loops are real.
  • Batch embeddings — one call with 100 inputs beats 100 calls with 1.
  • Cache embeddings at ingest time — re-embedding the same document at query time is waste.
  • Trim system prompts — they're paid for every turn.

For ongoing visibility, watch your spend in real time via the Observability metrics dashboard.

#What next

Was this page helpful?