Section
Errors & status codes

Errors & status codes

Every error code with cause and fix, including SSL and proxy issues.

Token Factory's public /v1 API and the console's API-key management use different error shapes today. This page covers the public /v1 contract and notes the differences inline.

#On the public /v1 API

Errors raised by the platform itself on /v1 use the OpenAI-style nested envelope:

JSON
{
  "error": {
    "type": "invalid_request_error",
    "code": "invalid_payload",
    "message": "Missing required field: model",
    "param": "model"
  }
}

The HTTP status code is the primary signal — pivot off that, not off the message text. The error.code is a stable machine-readable string; error.message is human-readable and may change.

Other surfaces use different shapes. If you're parsing errors programmatically, branch by surface — don't assume one envelope.

SurfaceEnvelope shapeStatus range
/v1/chat/completions{ error: { type, code, message, param } } (OpenAI-shape)4xx, 5xx
/v1/embeddingsSame OpenAI-shape envelope4xx, 5xx
/v1/modelsSame OpenAI-shape envelope4xx, 5xx
API key management routes{ code, error, ... } (nested with a stable string code)4xx, 5xx
/v1/endpoints (endpoint management)Mostly the RFC 7807 problem-detail fields ({ type, title, status, detail, instance }); early request validation returns a smaller { detail, code } instead. Both are served as application/json4xx, 5xx

#At a glance

StatusCauseWhat to do
400 Bad RequestMalformed JSON or missing required parameterRead the error string, fix the request shape
401 UnauthorizedMissing, malformed, unknown, revoked, or expired API key — all of them answer alikeCheck Authorization header; see Authentication
403 ForbiddenThe request authenticated, but the identity behind it isn't allowed here — most often a session carrying no organization context. Not what a revoked or expired key returnsSign in again, or switch to the workspace that owns the resource
404 Not FoundUnknown model, unknown endpointCheck the model value against the dashboard catalog
408 Request TimeoutRequest took too long to read or upstream worker timed outRetry with a shorter prompt or honor Retry-After
409 ConflictResource-state conflict (e.g. creating an API key past your workspace's key limit, revoking a key an administrator has disabled, modifying a workspace setting that another admin just changed, or creating a dedicated endpoint with a name that's already in use)Read the error string; refresh state and retry
409 GPU_ALLOTMENT_EXCEEDEDCreating, resuming, or changing the scaling of a dedicated endpoint would exceed your organization's GPU allotment for that GPU typeReduce the request or free capacity; see GPU allotment
422 Unprocessable EntityValidation failed (e.g. invalid value range)Read the error string for the offending field
429 Too Many RequestsRate limit, or a serverless quota exceeded — dedicated traffic is never quota-limitedBack off, then retry; see 429 Too Many Requests
500 Internal Server ErrorUnhandled server errorRetry with exponential backoff and jitter; if persistent, file a support ticket
502 Bad GatewayUpstream proxy issueRetry with exponential backoff and jitter
503 Service UnavailableThe model is not serving requests right nowRetry; if persistent, the model is degraded — pick an alternate from the catalog. For a dedicated endpoint, check it is serving

#401 Unauthorized

Causes: header missing, header malformed, key revoked, key expired, key from a different workspace.

  1. Confirm the header is present and shaped Authorization: Bearer sk-....
  2. Confirm the key starts with sk- and matches the masked prefix in the dashboard.
  3. If the key was rotated recently, your service may still have the old value cached — restart the worker.

#409 GPU_ALLOTMENT_EXCEEDED

Creating, resuming, or changing the scaling of a dedicated endpoint would put your organization over its GPU allotment for that GPU type. Distinct from the generic 409 Conflict above — this one is capacity-specific and carries extra fields in the response's top-level extensions object:

FieldMeaning
codeAlways GPU_ALLOTMENT_EXCEEDED.
gpuTypeWhich GPU type hit its cap, for example nvidia/h100.
allotmentYour organization's cap for this GPU type.
committedYour organization's commitment for this GPU type at the moment of the request.
projectedWhat your commitment would have become had this request been allowed. The request was rejected before changing anything, so committed above is unchanged.
needHow many more GPUs above the cap this request would need.
remediationA link back to GPU allotment.

Reduce the replica count or autoscaling ceiling and try again. Only create, resume, and scaling changes can return this — suspend, delete, and reads never do — and which of the three raised it changes what else is open to you: at create you can also choose a smaller flavor, but on an endpoint that already exists the flavor is fixed, so a scaling change or a resume has to come down in count or ceiling instead. See GPU allotment: What a 409 means for the per-operation table and the full capacity model.

503 GPU_ALLOTMENT_UNAVAILABLE is different

A 503 with extensions.code GPU_ALLOTMENT_UNAVAILABLE means the capacity check itself couldn't complete — it is not an allotment violation. Retry; the response carries no usage numbers, since the check never completed.

#429 Too Many Requests

You hit either a per-minute rate limit or your workspace quota.

Which it is depends on where the request went. Serverless calls are subject to both. A call to a dedicated endpoint is never counted against the monthly token quota — dedicated capacity is billed on GPU-hours instead — so a 429 from one is not your quota running out. Back off and retry.

  • The response includes Retry-After (seconds) — honor it.
  • Implement exponential backoff with jitter — start at 1s ±50%, double each retry (1s, 2s, 4s, 8s, 16s ±50% per step), capped at 30s. Jitter prevents thundering-herd retries when a transient upstream failure clears.
  • For sustained 429s, request a quota increase or shift load to a smaller, cheaper model in the catalog.

See Tokens, pricing & quotas for the full quota model.

#503 Service Unavailable

A specific model is temporarily not serving requests. The response is a 503; treat the message as diagnostic detail that may vary rather than a string to match on.

  1. Retry — most 503s resolve within seconds.
  2. Fall back to another model of similar capability from the Model Library — use the same capability chips to find peers.
  3. Check status — persistent 503 on a single model is a model-availability issue, not a platform issue.

A minimal fallback chain in Python:

from openai import OpenAI, APIError
import os

client = OpenAI(
  api_key=os.environ["OMNIVA_API_KEY"],
  base_url="https://api.tokenfactory.omniva.com/v1",
)

# Primary first, then a peer model from the Model Library
models = ["Omniva/glm-5.3", "<alternate-model-id>"]

for model in models:
  try:
      resp = client.chat.completions.create(
          model=model,
          messages=[{"role": "user", "content": "Hello"}],
      )
      break
  except APIError as e:
      if e.status_code == 503:
          continue
      raise

#Gateway timeouts

Not every failure comes from the platform, so not every failure body is JSON. If a request can't be routed to a serving replica, an edge timeout can end it before the platform returns anything — the response then carries a body that is not JSON rather than an error envelope, and code that calls .json() on it will throw while parsing.

This is the shape to expect from a dedicated endpoint that isn't currently serving — suspended, or still provisioning. The request is held open for minutes and then times out. Set a client timeout of your own, short enough that your deadline fires well before the gateway's and you get a clean, catchable error, and treat a non-JSON error body as a transport-level failure rather than a contract violation. See Lifecycle for which states serve and which don't.

#SSL

CERTIFICATE_VERIFY_FAILED on the Python openai SDK (and most TLS-using clients) means your runtime can't verify the server's certificate chain. The cause is almost always a corporate proxy or VPN intercepting TLS with its own root CA. The CA needs to be installed in the trust store your runtime uses.

#Python

Use certifi's bundle plus your corporate root.

Bash
export REQUESTS_CA_BUNDLE=/path/to/corp-root.pem
export SSL_CERT_FILE=$REQUESTS_CA_BUNDLE

Both must be set: requests reads REQUESTS_CA_BUNDLE; standard-library urllib/http.client reads SSL_CERT_FILE.

If you control the bundle file, append your corp root to certifi's bundle rather than replacing it:

Bash
cat $(python -m certifi) /path/to/corp-root.pem > /tmp/combined-ca.pem
export REQUESTS_CA_BUNDLE=/tmp/combined-ca.pem
export SSL_CERT_FILE=/tmp/combined-ca.pem

#Node.js

Bash
export NODE_EXTRA_CA_CERTS=/path/to/corp-root.pem

#curl

Bash
curl --cacert /path/to/corp-root.pem https://api.tokenfactory.omniva.com/v1/chat/completions
Do not disable verification

requests.get(..., verify=False), curl -k, NODE_TLS_REJECT_UNAUTHORIZED=0 all silence the symptom but leave you wide open to MITM. Install the root CA correctly instead.

#What next

Was this page helpful?