Token Factory's public /v1 API and the console's API-key management use different error shapes today. This page covers the public /v1 contract and notes the differences inline.
#On the public /v1 API
Errors raised by the platform itself on /v1 use the OpenAI-style nested envelope:
{
"error": {
"type": "invalid_request_error",
"code": "invalid_payload",
"message": "Missing required field: model",
"param": "model"
}
}The HTTP status code is the primary signal — pivot off that, not off the message text. The error.code is a stable machine-readable string; error.message is human-readable and may change.
Other surfaces use different shapes. If you're parsing errors programmatically, branch by surface — don't assume one envelope.
| Surface | Envelope shape | Status range |
|---|---|---|
/v1/chat/completions | { error: { type, code, message, param } } (OpenAI-shape) | 4xx, 5xx |
/v1/embeddings | Same OpenAI-shape envelope | 4xx, 5xx |
/v1/models | Same OpenAI-shape envelope | 4xx, 5xx |
| API key management routes | { code, error, ... } (nested with a stable string code) | 4xx, 5xx |
/v1/endpoints (endpoint management) | Mostly the RFC 7807 problem-detail fields ({ type, title, status, detail, instance }); early request validation returns a smaller { detail, code } instead. Both are served as application/json | 4xx, 5xx |
#At a glance
| Status | Cause | What to do |
|---|---|---|
400 Bad Request | Malformed JSON or missing required parameter | Read the error string, fix the request shape |
401 Unauthorized | Missing, malformed, unknown, revoked, or expired API key — all of them answer alike | Check Authorization header; see Authentication |
403 Forbidden | The request authenticated, but the identity behind it isn't allowed here — most often a session carrying no organization context. Not what a revoked or expired key returns | Sign in again, or switch to the workspace that owns the resource |
404 Not Found | Unknown model, unknown endpoint | Check the model value against the dashboard catalog |
408 Request Timeout | Request took too long to read or upstream worker timed out | Retry with a shorter prompt or honor Retry-After |
409 Conflict | Resource-state conflict (e.g. creating an API key past your workspace's key limit, revoking a key an administrator has disabled, modifying a workspace setting that another admin just changed, or creating a dedicated endpoint with a name that's already in use) | Read the error string; refresh state and retry |
409 GPU_ALLOTMENT_EXCEEDED | Creating, resuming, or changing the scaling of a dedicated endpoint would exceed your organization's GPU allotment for that GPU type | Reduce the request or free capacity; see GPU allotment |
422 Unprocessable Entity | Validation failed (e.g. invalid value range) | Read the error string for the offending field |
429 Too Many Requests | Rate limit, or a serverless quota exceeded — dedicated traffic is never quota-limited | Back off, then retry; see 429 Too Many Requests |
500 Internal Server Error | Unhandled server error | Retry with exponential backoff and jitter; if persistent, file a support ticket |
502 Bad Gateway | Upstream proxy issue | Retry with exponential backoff and jitter |
503 Service Unavailable | The model is not serving requests right now | Retry; if persistent, the model is degraded — pick an alternate from the catalog. For a dedicated endpoint, check it is serving |
#401 Unauthorized
Causes: header missing, header malformed, key revoked, key expired, key from a different workspace.
- Confirm the header is present and shaped
Authorization: Bearer sk-.... - Confirm the key starts with
sk-and matches the masked prefix in the dashboard. - If the key was rotated recently, your service may still have the old value cached — restart the worker.
#409 GPU_ALLOTMENT_EXCEEDED
Creating, resuming, or changing the scaling of a dedicated endpoint would put your organization over its GPU allotment for that GPU type. Distinct from the generic 409 Conflict above — this one is capacity-specific and carries extra fields in the response's top-level extensions object:
| Field | Meaning |
|---|---|
code | Always GPU_ALLOTMENT_EXCEEDED. |
gpuType | Which GPU type hit its cap, for example nvidia/h100. |
allotment | Your organization's cap for this GPU type. |
committed | Your organization's commitment for this GPU type at the moment of the request. |
projected | What your commitment would have become had this request been allowed. The request was rejected before changing anything, so committed above is unchanged. |
need | How many more GPUs above the cap this request would need. |
remediation | A link back to GPU allotment. |
Reduce the replica count or autoscaling ceiling and try again. Only create, resume, and scaling changes can return this — suspend, delete, and reads never do — and which of the three raised it changes what else is open to you: at create you can also choose a smaller flavor, but on an endpoint that already exists the flavor is fixed, so a scaling change or a resume has to come down in count or ceiling instead. See GPU allotment: What a 409 means for the per-operation table and the full capacity model.
A 503 with extensions.code GPU_ALLOTMENT_UNAVAILABLE means the capacity check itself couldn't complete — it is not an allotment violation. Retry; the response carries no usage numbers, since the check never completed.
#429 Too Many Requests
You hit either a per-minute rate limit or your workspace quota.
Which it is depends on where the request went. Serverless calls are subject to both. A call to a dedicated endpoint is never counted against the monthly token quota — dedicated capacity is billed on GPU-hours instead — so a 429 from one is not your quota running out. Back off and retry.
- The response includes
Retry-After(seconds) — honor it. - Implement exponential backoff with jitter — start at 1s ±50%, double each retry (1s, 2s, 4s, 8s, 16s ±50% per step), capped at 30s. Jitter prevents thundering-herd retries when a transient upstream failure clears.
- For sustained 429s, request a quota increase or shift load to a smaller, cheaper model in the catalog.
See Tokens, pricing & quotas for the full quota model.
#503 Service Unavailable
A specific model is temporarily not serving requests. The response is a 503; treat the message as diagnostic detail that may vary rather than a string to match on.
- Retry — most 503s resolve within seconds.
- Fall back to another model of similar capability from the Model Library — use the same capability chips to find peers.
- Check status — persistent 503 on a single model is a model-availability issue, not a platform issue.
A minimal fallback chain in Python:
from openai import OpenAI, APIError
import os
client = OpenAI(
api_key=os.environ["OMNIVA_API_KEY"],
base_url="https://api.tokenfactory.omniva.com/v1",
)
# Primary first, then a peer model from the Model Library
models = ["Omniva/glm-5.3", "<alternate-model-id>"]
for model in models:
try:
resp = client.chat.completions.create(
model=model,
messages=[{"role": "user", "content": "Hello"}],
)
break
except APIError as e:
if e.status_code == 503:
continue
raise#Gateway timeouts
Not every failure comes from the platform, so not every failure body is JSON. If a request can't be routed to a serving replica, an edge timeout can end it before the platform returns anything — the response then carries a body that is not JSON rather than an error envelope, and code that calls .json() on it will throw while parsing.
This is the shape to expect from a dedicated endpoint that isn't currently serving — suspended, or still provisioning. The request is held open for minutes and then times out. Set a client timeout of your own, short enough that your deadline fires well before the gateway's and you get a clean, catchable error, and treat a non-JSON error body as a transport-level failure rather than a contract violation. See Lifecycle for which states serve and which don't.
#SSL
CERTIFICATE_VERIFY_FAILED on the Python openai SDK (and most TLS-using clients) means your runtime can't verify the server's certificate chain. The cause is almost always a corporate proxy or VPN intercepting TLS with its own root CA. The CA needs to be installed in the trust store your runtime uses.
#Python
Use certifi's bundle plus your corporate root.
export REQUESTS_CA_BUNDLE=/path/to/corp-root.pem
export SSL_CERT_FILE=$REQUESTS_CA_BUNDLEBoth must be set: requests reads REQUESTS_CA_BUNDLE; standard-library urllib/http.client reads SSL_CERT_FILE.
If you control the bundle file, append your corp root to certifi's bundle rather than replacing it:
cat $(python -m certifi) /path/to/corp-root.pem > /tmp/combined-ca.pem
export REQUESTS_CA_BUNDLE=/tmp/combined-ca.pem
export SSL_CERT_FILE=/tmp/combined-ca.pem#Node.js
export NODE_EXTRA_CA_CERTS=/path/to/corp-root.pem#curl
curl --cacert /path/to/corp-root.pem https://api.tokenfactory.omniva.com/v1/chat/completionsrequests.get(..., verify=False), curl -k, NODE_TLS_REJECT_UNAUTHORIZED=0 all silence the symptom but leave you wide open to MITM. Install the root CA correctly instead.