Every chart and number on the Observability dashboard is defined here in plain English. It covers that dashboard and no other — other views in the console have panels of their own, described where those views are. If something on the Observability dashboard isn't covered here, file an issue.
#Throughput
#Total requests
What it is: Count of API calls — both successful and failed — recorded over the selected window.
Unit: count.
How it's computed: Every completed request in the window counts once, success or error. It is a total for the window you selected, not a rate: widen the window and the number grows.
When this is high: Heavy traffic. Watch for quota headroom and rate-limit pressure on the relevant key.
When this is low: Quiet traffic. Confirm clients are actually online before assuming health — a flat zero often means a misrouted base URL, not calm.
#Tokens per second (TPS)
What it is: Output-token rate, averaged across the requests in the window.
Unit: tokens/sec.
How it's computed: Per request, completion tokens divided by that request's own generation time — for a streaming request, the time after the first token arrives; for a non-streaming one, its full duration. Those per-request rates are then averaged, and the tile's headline figure is that average, with the window's p50 and p95 on the line beneath it. It is an average of rates, not total tokens divided by total time, so a single fast short request counts as much as a long one. The rate is also measured across the gaps between tokens, of which there is one fewer than there are tokens, so very short completions read slightly high.
When this is high: The model and the path between client and provider are decoding briskly. Good.
When this is low: Slow generation. Causes range from a saturated upstream provider to congested proxy pipes to clients that aren't draining the stream fast enough.
#Latency
#Time to first token (TTFT)
What it is: How long your caller waits between sending a request and seeing the first piece of the answer come back.
Unit: ms.
How it's computed: Milliseconds from the arrival of your request to the first byte written back to you, then surfaced as p50, p95, and (optionally) p99 percentiles over the window. It covers everything that happens in between — waiting in line, warming up, and the model reading your prompt before it writes a single token. For a reasoning model it also covers the reasoning tokens, not just the wait for the visible answer.
Streaming requests only. TTFT is recorded when a response streams. Non-streaming requests are excluded from this metric entirely rather than counted as zero — so a workload that never streams shows no TTFT at all. Use Total request duration for those.
When this is high: The user is waiting before they see anything. Compare against total request duration — if the two are close, the wait is nearly all pre-generation; if duration is much larger, generation itself is the slow part.
When this is low: Snappy starts. Users perceive the model as responsive even before generation completes.
#Total request duration
What it is: Milliseconds from request received to response fully written.
Unit: ms.
How it's computed: Milliseconds from request received to response fully written. Surfaced as average and as p50, p95, p99 percentiles.
When this is high: Long-running requests. Could be large completions, slow upstream, or stalled streams. Compare against TTFT to see whether the time went before generation or during it.
When this is low: Fast end-to-end completion. Pair with token counts to confirm the request actually did work and didn't truncate early.
#Reliability
#Error rate
What it is: Percentage of requests that returned a 4xx or 5xx response.
Unit: %.
How it's computed: error_count / total_requests × 100, where a request counts as an error when its HTTP status is 400 or above. Note that finish_reason does not enter this calculation: a request that returned 200 but finished with an error reason is not counted here.
When this is high: Investigate the dominant status class. 401/403 point at credentials; 429 points at quota or rate limits; 5xx points at upstream provider trouble. The status code of an individual request is on its row in the request explorer — use that to localize, since the dashboard does not break the error rate down by code.
Error rate above is the only reliability number on this dashboard. An overall success rate, a missing-completion rate and a finish_reason breakdown are not shown, and neither are per-status-code or per-finish_reason charts — see Filters and dimensions.
#Spend
#Input and output tokens
What it is: The prompt-token and completion-token counts for an individual request.
Unit: tokens.
Where to find them: These are per-request fields, not dashboard aggregates. Each row in the request explorer carries its own prompt and completion counts; the dashboard's rolled-up spend figures report total tokens, cost and request counts without splitting input from output. To analyse the input/output split across many requests, export or page through the request rows.
Why the split matters: Output tokens are typically priced higher than input, so they are usually the bigger cost lever. And a long system prompt is charged again on every turn — trimming it pays back per request, which is visible on the rows even though it is not charted.
#Total tokens
What it is: Sum of input and output tokens — the headline usage number.
Unit: tokens.
How it's computed: Sum of total_tokens across all matching requests.
When this is high: Heavy usage. Pair with cost to see whether the spend is concentrated on expensive models or simply spread across high volume.
When this is low: Quiet workload.
#Cost
What it is: Estimated dollar spend across all requests in the window.
Unit: USD ($).
How it's computed: Each request is priced at log time against the rate sheet for its model — input tokens at the model's input rate, output tokens at its output rate — and those per-request amounts are summed across the window. The dashboard surfaces this as a single total, a trend over time, and a top-models leaderboard ordered by cost descending. Because pricing is applied and stored per request, a later rate-sheet change does not retroactively alter what you see here. Models billed on something other than tokens — image or audio work — are priced on their own terms rather than by this token arithmetic.
When this is high: Concentrated on a few expensive models, or broad high-volume traffic. Drill into the top-models view to see which model owns the spend.
When this is low: Quiet, or running cheap models exclusively.
#Composition
#Modality mix
What it is: Breakdown of requests by request_type — text, image, audio, or embedding.
Unit: count (and % of total).
How it's computed: Requests are grouped and counted by request_type, with text as the fallback for older records that predate modality tagging.
#Filters and dimensions
There is no per-API-key filter, no status code as a chart dimension, and no finish_reason filter. Use the per-request rows for a status code or a finish reason on a single request.
The dashboard filter bar exposes three core controls today:
- Model — multi-select; populated from the workspace's active model set in the selected window.
- Request type — multi-select;
text,image,audio,embedding. - Time range — presets
now-15m,now-1h,now-24h,now-7d,now-30don this dashboard. Other views in the console offer longer windows and custom date ranges.
The same bar has a collapsible Advanced row carrying four more filters. These apply to everything on the dashboard, not only the request list:
- Request ID — exact-match lookup for a single request.
- Streaming — restrict to streaming or non-streaming requests.
- Tokens used — min/max bounds on a request's total token count. Note this filters tokens actually consumed, not the
max_tokensyou asked for — there's no column to filter that on, though the value does appear in a captured request body. - Latency (ms) — min/max bounds on request duration.
The request explorer carries one control of its own, separate from the bar above: Include prompt & response content, which shows or hides the request and response bodies on the rows you already have. It doesn't re-fetch the list.
Sort orders on the requests list: time_desc (default), time_asc, lat_desc, lat_asc, cost_desc, cost_asc, tokens_desc, tokens_asc.
Per-request rows carry: id, startTime, endTime, model, latency, ttftMs, tokensPerSecond, totalTokens (alongside promptTokens and completionTokens), totalCost, finishReason, requestType, stream, and statusCode. Image, audio and embedding requests add fields of their own — imageCount, inputLength, audioVoice, embeddingCount and embeddingDims — each present only when it applies. The dashboard composes the metrics above from this row set.
Metric values can lag actual usage briefly, so treat the dashboard as near-real-time rather than real-time. A request you just made may not be counted yet. To confirm a specific request landed, look it up by its Request ID in the filter bar's advanced filters rather than waiting for the aggregates to catch up — that lookup reads the rows directly. Its finishReason is on the row too, if you need to know how it ended.