Section
GPU cost

GPU cost

Billable GPU spend, hours, and utilization across your dedicated fleet — filters, trends, per-endpoint detail, and CSV export.

What did the dedicated fleet actually cost, and how much of that went on idle GPUs? The Cost tab on Monitoring → Dedicated — the console labels the page GPU spend — reports billable spend, GPU-hours, and utilization across your whole workspace. Dedicated endpoints are billed on GPU-hours alone — the tokens they serve are not metered against your workspace's monthly token quota, which governs serverless only.

Don't see a Cost tab?

GPU-cost visibility is a separate switch from the rest of dedicated-endpoint monitoring — the same switch that also controls the drill-down's GPU history panel. If you expect this tab and don't see it, ask your Omniva contact.

#Filters

The Cost view has its own filter row — separate from the search-by-name and status filters on Monitor dedicated endpoints:

FilterWhat it does
Time rangePresets from the last 24 hours out to the last 90 days, Lifetime for as far back as cost data is kept — about a year, not literally forever — or exact start and end dates of your own.
GPU typeNarrow to one or more GPU types.
ModelNarrow to one or more base models.
Billable onlyHide rows and totals that aren't actually billed.

#Reading the cost KPIs

Four tiles summarize the selected window:

  • GPU cost — total billable GPU spend, with a delta against the prior window and a projected monthly run rate (extrapolating the window's daily average out to 30 days).
  • GPU-hours — the GPUs an endpoint asked for, integrated over how long they were held. Two GPUs held for three hours is 6 GPU-hours, the same whether they were busy or idle the whole time. What counts is the GPUs requested and placed on hardware, not the ones ready to serve: a replica still pulling its image, warming up, or crash-looping accrues them exactly like a healthy one — which is why this can read higher than the live GPU count on Monitoring.
  • Avg $/GPU-hr — billable cost divided by GPU-hours for the window.
  • Avg utilization — how busy those GPUs were, averaged evenly across collection windows rather than weighted by size, so a one-GPU hour counts the same as a thirty-GPU hour. Treat it as a rough shape, not an exact figure, and note that a genuine 0% and a missing measurement look identical here.

When a window has no cost or GPU-hours at all, the tiles dim the figure rather than presenting a confident zero; average utilization is the one tile that shows a dash instead.

Two transparency chips appear only when they have something to report:

  • Captured (unbillable) — cost that was priced but not attributed to a billable endpoint, which in practice means a pricing mismatch. Usage the platform could not price at all — an endpoint with no history behind it, or a hardware type it could not resolve — carries no cost figure to show and so does not appear here. Note that the GPU-hours tile counts billable hours only, so none of the unbillable usage is in it — including the mismatched usage whose dollars this chip does report. That is why the two can't be reconciled against each other: this chip's cost has no matching hours anywhere on the page. A zero here means no mismatches, not that everything was accounted for.
  • Idle spend (util < 10%) — the slice of billable GPU cost that ran at very low utilization, so you can see how much of your spend is sitting mostly idle. Read it as an upper bound: a window whose utilization was never measured counts as idle here rather than being left out, so capacity that simply wasn't measured sits in this figure alongside capacity that genuinely sat idle. The Avg utilization tile shows a dash when it has nothing to report, and this chip stays hidden when utilization was never recorded at all — but it gives you no signal for the mixed case, where some windows were measured and some weren't, so check the Util column before acting on it.

Below the tiles:

  • Cost & GPU-hours trend — GPU-hours and GPU cost over time, shown as two stacked charts sharing one time axis rather than one chart with two different scales. A gap in either line means no sample for that point, not zero.
  • Cost by GPU type and Cost by model — donut breakdowns of the same window's billable cost, so you can see which GPU type or model is driving spend.
  • Effective rate by GPU type — your actual blended rate per GPU type for the window. The page states its own formula: "$/GPU-hr = billable GPU cost ÷ GPU-hours · GPU-hours integrate requested GPUs over time" — it's a blended actual, not your contracted rate, so it moves with your mix of flavors. The panel itself is just GPU type and rate — utilization is not part of the calculation and is not shown here; read it from the Avg utilization tile or the Util column below.

#Per-endpoint detail

The By endpoint table lists your endpoints by cost in the selected window, highest first, up to two hundred rows. Its figures are not filtered to billable usage the way the tiles above are, so the column totals will not tie back to them whenever any of the window was unbillable. Each row carries the GPU type, GPU-hours, GPU cost, Total cost (GPU plus CPU and RAM), Util, and a Billing column carrying a Billable badge — or, if a row isn't billable, a badge naming why instead.

Export CSV downloads the table exactly as filtered, as gpu-cost-endpoints-<date>.csv. It carries more than the screen does: GPU, CPU and RAM costs as separate columns rather than the single Total cost, plus the specific GPU model, the GPU count, and the reason behind any non-billable badge.

Click a row for its cost drill-down: the same figures broken out interval by interval, with utilization reported per window rather than averaged across the range. It also carries provenance — the GPU model actually observed, which rate applied, and the workspace the cost belongs to — for reconciling against your own records.

The drill-down has a horizon of about six weeks

A drill-down returns at most a thousand intervals, which at hourly sampling is roughly 41 days. Select a longer range — the 90-day preset, or Lifetime — and the oldest intervals fall outside that limit: they are dropped silently, with no notice on the view, and the drill-down's own summary figures are computed from what survived, so they will read lower than the true total for the range. The By endpoint table and the tiles above are not affected. For reconciliation beyond about six weeks, work in shorter ranges and add them up, rather than trusting a single long-range drill-down.

#What next

Was this page helpful?