Latency crept up, errors appeared, or you want to confirm the fleet is healthy: Monitoring → Dedicated is where those numbers live, fleet-wide or one endpoint at a time. The fleet view's filter row searches by endpoint or model name, filters by status, and picks a time range from the last 15 minutes out to 30 days, with a manual Refresh and an Auto 30s toggle for a live view. The drill-down keeps the time range, Refresh and Auto 30s, and drops the search and status filters — on one endpoint there is nothing to filter down to.
The dedicated monitoring surface has its own switch, separate from both the Dedicated Endpoints feature itself and from the Cost lens. Where it is off, the link isn't shown and going to the URL directly lands you on the serverless view instead — there is no error to read. If you expect it and don't see it, ask your Omniva contact.
There is no alerting on this surface today — no thresholds you can set, nothing that emails or pages you when a
number turns bad. Watching means this page: leave Auto 30s on during anything sensitive. The numbers are also
console-only — they are not served on the public /v1 API and an API key cannot fetch them. The one
machine-readable export is the CSV on the Cost lens.
#Fleet overview
A status strip across the top counts endpoints by state and doubles as a filter for the table below it. Under it, fleet-wide tiles cover request volume, successes and failures, error rate, average and P95 latency, and throughput. Three of them — request volume, error rate and P95 latency — also carry a trend against the prior window; the rest show the window's figure alone.
A Requests over time chart shows volume with an errors band — the bucket width follows the range you picked, from a minute on the shortest window out to several hours on the longest, so don't read a bar as an hour, and a GPUs in use — live panel counts GPUs actually backing running replicas right now, broken out by type. That's a live count, not your committed capacity — for the latter, see GPU allotment.
A Needs attention panel appears only when something is in an error state, worst first. Below it, a table lists every endpoint with its status, GPUs, and request/error/latency numbers for your selected window; an idle endpoint reads "idle" rather than a misleading zero. Throughout, request counts reflect completed responses — some client-disconnect cases aren't counted. Read the error figures as a floor rather than a total: a request that fails before any response is recorded never lands here at all, so real-world failures can be higher than the error rate shown.
Fleet liveness and request telemetry come from two different sources, and they fail differently. The page itself always loads either way — neither outage takes it down, so a blank screen is never the signal.
If the fleet source goes quiet, status and GPU columns fall back to —, the request, error and latency figures beside them stay good, and the table's footer names the half that's missing rather than printing a zero it can't stand behind. One thing the dashes don't tell you: an endpoint with no traffic in your selected window disappears from the table altogether while this lasts, because the only source that knew it existed is the one that's down. A shorter list than you expect is part of the same outage, not a set of deleted endpoints.
If the telemetry source goes quiet, the affected panel says so in place and offers Retry; the rest of the page renders normally. A refresh that fails leaves the last good numbers on screen rather than blanking them, so check the panel's own message before trusting a figure that hasn't moved. When a number looks wrong below is the symptom-by-symptom version.
#Per-endpoint drill-down
Click an endpoint for the same view scoped to it alone, plus finer-grained latency: Total Requests, Successful, Failed, Error Rate, and p50 / p95 / p99 latency.
Where the underlying engine exposes them, a Performance & engine — live panel adds time-to-first-token, inter-token latency, in-flight and queued requests, KV-cache use, and GPU utilization and memory — and says so plainly when it can't, distinguishing an engine that isn't reporting from one that simply has no traffic. Two more panels round it out: Error classification (client vs. server errors) and Finish reasons (how completions actually ended — stop, length, a filter, or an error).
#When a number looks wrong
Two of the eight symptoms below are a source that isn't answering. The other six are not: an endpoint that was deleted, a window with no traffic in it, two different GPU-allotment states, the endpoint's own configuration, and one case where the HTTP API and the console simply use different words for the same thing. Find the symptom, run the check in the third column before acting on it, then follow the link for the full picture.
| Surface | Exact symptom | Discriminating check | Next action | Full detail |
|---|---|---|---|---|
| Console · Monitoring | Status and GPUs read — on every row at once | The table's footer reads "fleet source unreachable — status and GPU columns unavailable" | The console can't tell what state these endpoints are in — the dashes are not a verdict either way, so don't read them as trouble or as reassurance. Endpoints with no traffic in the window drop out of the list entirely until this clears. Read status from the Endpoints page instead, and retry. Request, error, and latency figures in the same table come from a different source and are still good. | Status meanings |
| Console · Monitoring | Status reads Unavailable on one row | That endpoint has traffic in the window but no row at all on the Endpoints page | That endpoint no longer exists — deletion is the usual cause. Its traffic history stays queryable, but the endpoint doesn't come back; create a new one if you still need it. | Delete |
| Console · Monitoring | GPUs in use — live reads — under "Live GPU telemetry unavailable" | The same panel says "Loading live GPU telemetry…" instead while a read is still in flight — the caption is what separates the two | Refresh. The dash is deliberate: the panel won't print a zero it can't verify. Request, error, and latency figures come from the other source and are unaffected. | Observability metrics |
| Console · Monitoring | Requests reads "idle" and Error rate reads — next to "no traffic" | The row is present and its status isn't Error | Zero requests in the selected window — not a fault. Widen the time range before investigating anything else. | Observability metrics |
| Console · Endpoints | Balance reads over allotment with a negative figure | The Committed cell still reads "of N" with N above zero, so a cap exists and your commitment has passed it | Reduce a commitment of that type, or ask for a larger cap — in that order. | Recovering, in order |
| Console · Endpoints | Balance reads unentitled with a negative figure | The Committed cell reads "· no allotment" rather than "of N" — the cap for that type is zero, not merely exceeded | Ask your Omniva contact to provision that GPU type. A cap of zero means there is no headroom for that type at all, whatever you do to your own endpoints — suspending, deleting, or shrinking others of that type will clear the negative figure, but it buys you no room. | Requesting more capacity |
| Console · Endpoints | Replicas never rise above N | Which kind of endpoint is it? The row's own scaling summary tells you: a plain replica count for a fixed endpoint, a scale min–max range for an autoscaled one. Fixed — N is the count you set; there is no ceiling to raise. Autoscaled — check N against maxReplicas first, then check demand. The metric you chose is one demand signal among several the platform watches, and whichever asks for the most replicas wins — so your metric being quiet proves nothing on its own. Read them all in the drill-down's Performance & engine — live panel: in-flight requests, queued requests, KV-cache use | Fixed: set a higher count with Edit scaling. Autoscaled and sitting at maxReplicas: raise maxReplicas, whichever signal is driving it. Autoscaled, below maxReplicas, and every signal on that panel quiet: demand is the limit, not the ceiling, and raising maxReplicas changes nothing. Either increase is measured against your allotment and can be rejected. | Edit scaling |
| HTTP API | status.state comes back Failed for an endpoint you suspended | Open that endpoint's Conditions panel and look for a Suspended condition reading True. The console's own column reads Suspended either way, because it puts an asserted Suspended ahead of everything else — so the column alone doesn't settle it. A Degraded condition alongside it is not conclusive either: one raised before the suspend survives it, so an old fault and a current one look identical here | Usually nothing is broken; Resume when you want it serving. To tell a stale Degraded from a live one, check whether it predates the suspension — and if you can't establish that, resume and see whether it clears. If it reasserts on a running endpoint, treat it as a real fault and follow the Conditions panel guidance. | The Endpoint object |
#GPU cost
Where GPU-cost visibility is available to you, a GPU history — last 7d panel on the drill-down adds GPU-hours and average utilization for that endpoint. If you expect it and don't see it, ask your Omniva contact.
For cost across your whole dedicated fleet — billable spend, trends, and a per-endpoint breakdown, not just one endpoint's last 7 days — see GPU cost.