Section
GPU allotment headroom

/v1/gpu-allotment/headroom

Read your organization's GPU allotment usage and remaining headroom per GPU type.

GET/v1/gpu-allotment/headroom

An advisory, read-only look at your organization's GPU allotment: current commitment, cap, and remaining headroom per GPU type. Nothing here is reserved or locked by calling it — use it to check headroom before you create, update, or resume an endpoint. Use it as a preflight hint, but always handle 409 — headroom can change before your mutation is checked.

#Authentication

Bearer token in the Authorization header — the same API key you use for /v1 requests.

HTTP
Authorization: Bearer $OMNIVA_API_KEY

See Authentication for how to mint and rotate keys.

A missing, malformed, revoked, or out-of-workspace key answers with a 401 or 403 whose body carries neither the shape below nor the /v1 one. Treat both as opaque and read the HTTP status line — see 401 Unauthorized.

#Request

FieldTypeDescription
target
stringAn endpoint name in your workspace. When set and the endpoint exists, the response's target field reports what it would commit if it were running right now — useful for clamping an edit or resume form's ceiling before you submit. Omitted from the response if no endpoint by this name exists in your workspace.

No request body.

#Response

200application/json
{
"state": "provisioned",
"perType": {
  "nvidia/h100": {
    "allotment": 8,
    "committed": 6,
    "remaining": 2
  }
}
}

The fields below are the response's supported contract. Anything else in the body is unsupported and may change without notice, so don't branch on it.

FieldTypeDescription
statestringprovisioned (your organization has a cap, even an all-zero one) or unprovisioned (no cap has been set for your organization yet).
perTypeobjectGPU type (e.g. nvidia/h100) to usage. Empty when no allotment is tracked for your organization.
perType.<type>.allotmentintegerYour organization's cap for this GPU type.
perType.<type>.committedintegerCurrent committed usage for this GPU type, summed across your organization.
perType.<type>.remainingintegerallotment − committed, floored at 0 — never negative, even over cap.
targetobjectPresent only when the request set ?target= and an endpoint by that name exists in your workspace.
target.gpuTypestringWhich perType entry to measure target.commitment against.
target.commitmentintegerWhat the target endpoint would commit if it were running — gpu × replicas at its current or maximum scale, per What an endpoint commits. A suspended target reports what resuming it would commit, not its current (zero) reservation.

Because this read holds nothing, the numbers it returns are a snapshot, not a reservation: another endpoint in your organization can commit the remaining headroom between your read and your next mutation. Clamp your forms against it, then still handle the 409.

{
"state": "provisioned",
"perType": {
  "nvidia/h100": {
    "allotment": 8,
    "committed": 6,
    "remaining": 2
  }
},
"target": {
  "gpuType": "nvidia/h100",
  "commitment": 4
}
}

#Errors

This read has one error of its own, and it uses the RFC 7807 problem-detail field shape — served as application/json, like every response on this surface. See The error envelope for the base shape and for the smaller body that early request validation returns.

503application/json
{
"type": "about:blank",
"title": "GPU Allotment Check Unavailable",
"status": 503,
"detail": "The GPU allotment capacity check is temporarily unavailable. Please try again.",
"instance": "/v1/gpu-allotment/headroom",
"extensions": {
  "code": "GPU_ALLOTMENT_UNAVAILABLE"
}
}
StatusCause
503GPU_ALLOTMENT_UNAVAILABLE — the capacity read couldn't be completed, so no numbers are reported. Retry.

#Code samples

import requests
import os

resp = requests.get(
  "https://api.tokenfactory.omniva.com/v1/gpu-allotment/headroom",
  headers={"Authorization": "Bearer " + os.environ["OMNIVA_API_KEY"]},
  params={"target": "chat-prod-endpoint"},
)
resp.raise_for_status()
headroom = resp.json()
for gpu_type, usage in headroom["perType"].items():
  print(gpu_type, usage["remaining"], "remaining of", usage["allotment"])

#What next

Was this page helpful?