GPU allotment caps how many GPUs of each type your organization can commit across its dedicated endpoints. Your allotment is set up with Omniva when your organization is onboarded, per GPU type, and every endpoint you create draws on it. It's checked on create, resume, and scaling changes — never on suspend, delete, or a plain read.
Your create, resume, or scaling change would have put your organization over its cap for that GPU type — nothing was changed. Jump to What a 409 means for the response fields and the fix.
#What an endpoint commits
A commitment is a count of GPUs, not of replicas — one replica can hold several GPUs. Each endpoint commits GPUs per replica × replicas, and which replica number counts depends on how the endpoint is sized:
- Fixed replicas — commits
GPUs per replica × replicas. - Autoscaling — commits
GPUs per replica × maxReplicas: the ceiling you configured, not the number of replicas running at any given moment. - Suspended — commits nothing, whatever its size.
Your organization's committed total for a GPU type is the sum of that figure over every endpoint holding that type.
A worked example. A flavor with 2 GPUs per replica, autoscaling between 1 and 4 replicas, commits 2 × 4 = 8 GPUs of its type. It commits 8 while all four replicas are serving, and it still commits 8 after traffic drops away and it settles back to one.
When an endpoint scales itself down, nothing is released — not even when it reaches zero replicas. The ceiling is what is reserved; what runs beneath it moves freely.
Changing that ceiling yourself is a different thing, and it does release capacity. Lower a fixed endpoint's replicas, or an autoscaled endpoint's maxReplicas, and its commitment drops the moment you save. Suspend and Delete release the commitment in full. Those three are the levers in the recovery ladder below.
#Reading the GPU allotment ledger
The GPU allotment panel on the Endpoints page in the console shows one row per GPU type in your allotment or in use, with four columns:
- Type — the GPU type this row budgets. Each type is capped on its own, so headroom on
H100does nothing for aB200endpoint. - Committed — how many GPUs of that type are reserved right now, shown as
{committed} of {allotment}(for example3 of 8) — or with an · no allotment suffix if you've committed against a type with a zero cap. - Utilization — at a glance, how close this type is to its cap.
- Balance — how much headroom is left, or how far over you are.
Committed can be larger than the endpoints you can see. It sums across your whole organization, while the Endpoints page and the endpoints API both show only the workspace you are in — so endpoints your organization runs in its other workspaces raise the figure without ever appearing in your list. If the ledger and your endpoints don't reconcile, that gap is the first thing to check.
Balance is the column to watch — its caption changes as you approach the cap:
- Under 80% used — remaining count, captioned remaining.
- 80% or more used — remaining count, captioned remaining · N% (e.g. "remaining · 83%").
- At the cap — 0, still captioned remaining.
- Over the cap — a negative count, captioned over allotment.
- No allotment at all for this type — a negative count, captioned unentitled.
Note the panel's footnote: "Suspended endpoints don't count toward your allotment."
If the panel doesn't appear, the numbers just aren't available to show right now — your cap itself is unchanged. Either way, treat the 409 as the answer that counts, not what the console did or didn't show you.
#Per-GPU-type caps
Allotment is capped per GPU type, not as one combined pool — nvidia/h100, nvidia/b200 and nvidia/b300. Your own allotment won't necessarily cover all three: the ledger shows one row per type you hold or have in use, and that is what you can draw on. Ask your Omniva contact about a type you need and don't see. The ledger and the create form track each type separately.
When you create or resize an endpoint, the form works from whatever's left for the flavor's GPU type: replica and autoscaling fields are capped at your remaining headroom where it can read them, and a flavor that doesn't fit even one replica is labeled Doesn't fit. Create a dedicated endpoint walks through what that looks like in the form itself.
The form isn't the last word, though, because headroom moves: another endpoint in your organization can claim the capacity you were shown between the page loading and your submit. So a form that let you proceed is not a promise there was still room — handle the 409 on every mutating call.
#What a 409 means
A 409 GPU_ALLOTMENT_EXCEEDED response means the create, resize, or resume you attempted was rejected because it would put your organization over its allotment for that GPU type. The response's extensions object carries:
| Field | Type | Description |
|---|---|---|
code | string | Always GPU_ALLOTMENT_EXCEEDED. |
gpuType | string | Which GPU type hit its cap, for example nvidia/h100. |
allotment | integer | Your organization's cap for this GPU type. |
committed | integer | Your organization's commitment for this GPU type at the moment of the request. |
projected | integer | What your commitment would have become had this request been allowed. |
need | integer | How many more GPUs above the cap this request would need. |
remediation | string | A link back to this page. |
Nothing changed: the request was rejected before it was applied, so your commitment is still exactly what committed reports.
How you recover depends on which of the three operations raised it. A flavor is chosen once, at create, and is fixed for the life of the endpoint — so "pick a smaller flavor" is a create-time move only, and never the answer to a rejected scaling change or resume:
| 409 raised on | First safe action | If that still doesn't fit |
|---|---|---|
| Create | Lower the replica count or maxReplicas, or choose a smaller flavor — the flavor is still yours to pick at this point. See Create a dedicated endpoint. | Work down the ladder below. |
| Scaling change | Lower the requested replica count or maxReplicas. The endpoint's own current reservation is credited back before the check, so only a net increase can be rejected — shrinking or staying the same never is. | The flavor can't change on an existing endpoint; that takes a delete and a fresh create under a new configuration. Otherwise, the ladder below. |
| Resume | Use Edit scaling on the suspended endpoint to bring the count or ceiling down, save, then resume. A resume carries no size of its own — it is measured at whatever is stored on the endpoint. See Lifecycle: Edit scaling. | Work down the ladder below. |
#Recovering, in order
need is a shortfall in GPUs, not in replicas. At 2 GPUs per replica a need of 4 is two replicas' worth; at 8 GPUs per replica the same 4 is not even one. Whichever step you take, reduce the requested commitment by at least that many GPUs.
The order below is by reversibility, not by disruption — how disruptive a step is depends on your workload, but the last one cannot be taken back.
- Adjust the rejected request. Lower the replica count or
maxReplicasyou asked for by at leastneedGPUs' worth, or at create time pick a smaller flavor. This is the only step that leaves everything you already run untouched. - Lower another endpoint's ceiling. Use Edit scaling on an endpoint of the same GPU type to bring its
replicasormaxReplicasdown: it releases the difference and keeps serving. What you give up is that endpoint's room to grow — and, if you set the ceiling below what it is running right now, some of the capacity it is using today. - Suspend an endpoint you can stop. This releases its whole commitment for that GPU type. Suspending is not blocked by GPU allotment, so capacity is never what stands in the way — but the endpoint stops serving, and bringing it back is itself a capacity check that can be refused.
- Ask for a larger allotment. The only route when the type's cap is zero, or when everything you hold is genuinely needed. See Requesting more capacity.
- Delete an endpoint you no longer need. Last, because it is the one step with no way back: deleting releases its allotment commitment, and the endpoint's name and configuration go with it. Suspending releases the same commitment and keeps both — compare the three before you choose.
Headroom moves underneath you: another endpoint in your organization can claim the capacity you just freed before your retry lands. Keep handling the 409 on every mutating call rather than assuming a cleared check stays cleared — see GPU allotment headroom for why a headroom read is a snapshot, not a reservation.
A 503 GPU_ALLOTMENT_UNAVAILABLE response is different: the capacity check itself couldn't complete. Retry — it carries no usage numbers, since the check never completed.
#Suspend, resume, and scale-to-zero
Capacity behaves a little differently at each lifecycle stage:
- Suspend releases its allotment commitment, and is not blocked by GPU allotment — a suspended endpoint doesn't count toward your allotment at all.
- Resume re-checks capacity from scratch. A suspended endpoint's capacity may already have been claimed by another endpoint, so resuming can be blocked — or return the same 409 — if there's no longer room.
- Autoscaling to zero replicas still reserves your maximum. An autoscaled endpoint that has scaled down to zero live replicas keeps its full
maxReplicascommitted against your allotment. Nothing is running, but the headroom stays reserved so it's there the moment traffic returns. - Delete releases its allotment commitment — an endpoint being deleted stops counting straight away.
#Requesting more capacity
GPU allotment is provisioned for your organization by Omniva, not something you configure yourself. If your organization isn't provisioned yet, or needs more than its current cap, ask your Omniva contact to adjust the allotment.
Worth knowing before you spend time debugging a create that won't go through: if your organization has no allotment provisioned at all, every GPU type reads a cap of zero — so every create is refused with 409 GPU_ALLOTMENT_EXCEEDED and an allotment of 0. Nothing is wrong with the request; there is simply no capacity to draw on yet. That is a provisioning conversation, not a smaller flavor.
#GPU allotment vs. token quotas
GPU allotment and your workspace's monthly token quota are two independent systems — allotment caps dedicated-endpoint GPU capacity, the quota caps serverless usage — and hitting one never affects the other. See Tokens, pricing & quotas for the quota model.