An advisory, read-only look at your organization's GPU allotment: current commitment, cap, and remaining headroom per GPU type. Nothing here is reserved or locked by calling it — use it to check headroom before you create, update, or resume an endpoint. Use it as a preflight hint, but always handle 409 — headroom can change before your mutation is checked.
#Authentication
Bearer token in the Authorization header — the same API key you use for /v1 requests.
Authorization: Bearer $OMNIVA_API_KEYSee Authentication for how to mint and rotate keys.
A missing, malformed, revoked, or out-of-workspace key answers with a 401 or 403 whose body carries neither the shape below nor the /v1 one. Treat both as opaque and read the HTTP status line — see 401 Unauthorized.
#Request
| Field | Type | Description |
|---|---|---|
target | string | An endpoint name in your workspace. When set and the endpoint exists, the response's target field reports what it would commit if it were running right now — useful for clamping an edit or resume form's ceiling before you submit. Omitted from the response if no endpoint by this name exists in your workspace. |
No request body.
#Response
{
"state": "provisioned",
"perType": {
"nvidia/h100": {
"allotment": 8,
"committed": 6,
"remaining": 2
}
}
}The fields below are the response's supported contract. Anything else in the body is unsupported and may change without notice, so don't branch on it.
| Field | Type | Description |
|---|---|---|
state | string | provisioned (your organization has a cap, even an all-zero one) or unprovisioned (no cap has been set for your organization yet). |
perType | object | GPU type (e.g. nvidia/h100) to usage. Empty when no allotment is tracked for your organization. |
perType.<type>.allotment | integer | Your organization's cap for this GPU type. |
perType.<type>.committed | integer | Current committed usage for this GPU type, summed across your organization. |
perType.<type>.remaining | integer | allotment − committed, floored at 0 — never negative, even over cap. |
target | object | Present only when the request set ?target= and an endpoint by that name exists in your workspace. |
target.gpuType | string | Which perType entry to measure target.commitment against. |
target.commitment | integer | What the target endpoint would commit if it were running — gpu × replicas at its current or maximum scale, per What an endpoint commits. A suspended target reports what resuming it would commit, not its current (zero) reservation. |
Because this read holds nothing, the numbers it returns are a snapshot, not a reservation: another endpoint in your organization can commit the remaining headroom between your read and your next mutation. Clamp your forms against it, then still handle the 409.
{
"state": "provisioned",
"perType": {
"nvidia/h100": {
"allotment": 8,
"committed": 6,
"remaining": 2
}
},
"target": {
"gpuType": "nvidia/h100",
"commitment": 4
}
}#Errors
This read has one error of its own, and it uses the RFC 7807 problem-detail field shape — served as application/json, like every response on this surface. See The error envelope for the base shape and for the smaller body that early request validation returns.
{
"type": "about:blank",
"title": "GPU Allotment Check Unavailable",
"status": 503,
"detail": "The GPU allotment capacity check is temporarily unavailable. Please try again.",
"instance": "/v1/gpu-allotment/headroom",
"extensions": {
"code": "GPU_ALLOTMENT_UNAVAILABLE"
}
}| Status | Cause |
|---|---|
| 503 | GPU_ALLOTMENT_UNAVAILABLE — the capacity read couldn't be completed, so no numbers are reported. Retry. |
#Code samples
import requests
import os
resp = requests.get(
"https://api.tokenfactory.omniva.com/v1/gpu-allotment/headroom",
headers={"Authorization": "Bearer " + os.environ["OMNIVA_API_KEY"]},
params={"target": "chat-prod-endpoint"},
)
resp.raise_for_status()
headroom = resp.json()
for gpu_type, usage in headroom["perType"].items():
print(gpu_type, usage["remaining"], "remaining of", usage["allotment"])