Section
Chat completions

/v1/chat/completions

Create a chat completion.

POST/v1/chat/completions

Generate a chat completion. OpenAI-compatible wire shape — existing OpenAI SDKs work with a base_url swap to https://api.tokenfactory.omniva.com/v1.

#Authentication

Bearer token in the Authorization header. See Authentication for how to mint a key.

HTTP
Authorization: Bearer YOUR_API_KEY

#Parameters

FieldTypeDescription
modelrequired
stringID of the model to use (e.g. Omniva/glm-5.3), or, if your workspace has dedicated endpoints, one of their Model IDs — the form is dedicated/ followed by the endpoint name. Must match a model available to your workspace.
messagesrequired
Message[]Conversation so far, as an ordered array of { role, content } objects. role is one of system, user, assistant, or tool.
temperature
numberSampling temperature. Higher values produce more random output; lower values are more deterministic. Typical range 0 to 2.
Default: 1
max_tokens
integerHard cap on the number of tokens generated in the response. If unset, the model may fill its full context window.
stream
booleanIf true, the response is streamed as Server-Sent Events. See the streaming section below.
Default: false
n
integerNumber of completion choices to generate. Not supported today — the handler ignores this field and returns a single choice.
top_p
numberNucleus-sampling cutoff in the range 0 to 1. Not supported today — the handler ignores this field; use temperature instead.
presence_penalty
numberPenalty in the range -2 to 2 for tokens already present in the prompt. Not supported today — the handler ignores this field.
frequency_penalty
numberPenalty in the range -2 to 2 for tokens by their frequency so far. Not supported today — the handler ignores this field.
logit_bias
objectMap of token-id to bias value. Not supported today — the handler ignores this field.
logprobs
booleanIf true, return log probabilities of the output tokens. Not supported today — the handler ignores this field.
top_logprobs
integerNumber of most-likely tokens to return at each position when logprobs is set. Not supported today — the handler ignores this field.
seed
integerSeed for best-effort deterministic sampling. Not supported today — the handler ignores this field.
stop
string | string[]Up to 4 strings where the model will stop generating. Not supported today — the handler ignores this field.
response_format
objectConstrain the output format (e.g. JSON mode). Not supported today — the handler ignores this field.
tools
Tool[]Tool definitions the model may call. Not supported today — the handler ignores this field and returns no tool_calls.
tool_choice
string | objectControl which tool, if any, the model picks. Not supported today — the handler ignores this field.
user
stringEnd-user identifier for abuse-tracking. Not supported today — the handler ignores this field.
service_tier
stringService-tier hint (e.g. auto, default). Not supported today — the handler ignores this field.

Where dedicated endpoints are in use, their Model IDs route to that endpoint's own GPU capacity instead of the shared pool — everything else about the request (auth, shape, streaming) stays the same. See Call a dedicated endpoint for where to find that id and how it behaves.

#Message object

FieldTypeDescription
rolestringOne of system, user, assistant, tool.
contentstringThe message text.
namestring (optional)Identifies the speaker in named conversations. Most clients leave this unset.
tool_call_idstringRequired when role is "tool". References the tool_calls[N].id from the assistant message that requested the tool call.
tool_callsToolCall[]On an assistant message, the array of tool invocations the model wants to perform. Each entry has id, type (typically "function"), and a function: { name, arguments }. Token Factory does not yet emit tool_calls — see Parameters above.

#Response

200application/json
{
"id": "cmpl-1748044800000",
"object": "chat.completion",
"created": 1748044800,
"model": "Omniva/glm-5.3",
"choices": [
  {
    "index": 0,
    "message": {
      "role": "assistant",
      "content": "A vector database stores high-dimensional embeddings and retrieves nearest neighbors by similarity."
    },
    "finish_reason": "stop"
  }
],
"usage": {
  "prompt_tokens": 18,
  "completion_tokens": 19,
  "total_tokens": 37
}
}

The top-level fields follow OpenAI's chat.completion shape: id, object, created (epoch seconds, UTC), model, choices[], usage. Each choices[i] has index, message ({ role, content }), and finish_reason (typically stop).

#Streaming response shape

When stream: true, the response is text/event-stream. Each event is a data: line carrying a JSON chunk in OpenAI's chat.completion.chunk format. Chunks accumulate via delta.content; the final chunk sets finish_reason: "stop". The stream terminates with a sentinel data: [DONE] frame — clients should stop reading at that point.

The on-the-wire shape is one SSE event per chunk, blank-line separated:

Text
data: {"id":"cmpl-1748044800000","object":"chat.completion.chunk","created":1748044800,"model":"Omniva/glm-5.3","choices":[{"index":0,"delta":{"role":"assistant","content":"Hello"},"finish_reason":null}]}

data: {"id":"cmpl-1748044800000","object":"chat.completion.chunk","created":1748044800,"model":"Omniva/glm-5.3","choices":[{"index":0,"delta":{"content":" there"},"finish_reason":null}]}

data: {"id":"cmpl-1748044800000","object":"chat.completion.chunk","created":1748044800,"model":"Omniva/glm-5.3","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}

data: [DONE]

When a model emits tool calls (not yet supported by Token Factory — included here for wire-shape reference), the tool_calls array streams incrementally. Each chunk carries a partial function.arguments string that the client concatenates by index:

Text
data: {"id":"cmpl-...","object":"chat.completion.chunk","created":1748044800,"model":"Omniva/glm-5.3","choices":[{"index":0,"delta":{"tool_calls":[{"index":0,"id":"call_abc","type":"function","function":{"name":"get_weather","arguments":""}}]},"finish_reason":null}]}

data: {"id":"cmpl-...","object":"chat.completion.chunk","created":1748044800,"model":"Omniva/glm-5.3","choices":[{"index":0,"delta":{"tool_calls":[{"index":0,"function":{"arguments":"{\"city\":"}}]},"finish_reason":null}]}

data: {"id":"cmpl-...","object":"chat.completion.chunk","created":1748044800,"model":"Omniva/glm-5.3","choices":[{"index":0,"delta":{"tool_calls":[{"index":0,"function":{"arguments":"\"SF\"}"}}]},"finish_reason":"tool_calls"}]}

data: [DONE]

See the Streaming guide for client-side consumption patterns.

#Errors

Errors raised by the platform return a JSON body with an error field; a request that can't reach a serving replica can instead end in a gateway timeout whose body is not JSON. The HTTP status indicates the class.

StatusCauseWhat to do
400 Bad RequestValidation error — missing model or messages, or malformed payloadInspect error.param when present; fix the request shape
401 UnauthorizedMissing or invalid Bearer tokenCheck Authorization header; see Authentication
403 ForbiddenAuthenticated, but the identity isn't allowed here — typically no organization context on the session. A revoked or expired key returns 401, not thisSign in again, or switch to the owning workspace
404 Not FoundUnknown modelCheck the model value against the dashboard catalog
408 Request TimeoutRequest took too long or upstream worker timed outRetry with a shorter prompt or honor Retry-After
409 ConflictResource-state conflict on workspace or key stateRead the error string; refresh state and retry
422 Unprocessable EntityValidation failed (e.g. invalid value range)Read the error string for the offending field
429 Too Many RequestsRate limit or quota exceededBack off, then retry; see Tokens, pricing & quotas
500 Internal Server ErrorUnhandled server errorRetry with exponential backoff and jitter; file a support ticket if persistent
502 Bad GatewayUpstream proxy issueRetry with exponential backoff and jitter
503 Service UnavailableThe model is not serving requests right nowRetry; if persistent, pick an alternate from the catalog

See Errors for the full error catalog and retry guidance.

#Code samples

from openai import OpenAI
import os

client = OpenAI(
  api_key=os.environ["OMNIVA_API_KEY"],
  base_url="https://api.tokenfactory.omniva.com/v1",
)

resp = client.chat.completions.create(
  model="Omniva/glm-5.3",
  messages=[{"role": "user", "content": "Hello in one sentence."}],
)
print(resp.choices[0].message.content)

#Not yet supported

The Parameters table above marks each unsupported OpenAI field with "Not supported today" in its Description. The handler accepts these fields in the request body without erroring, but ignores them — the response will not change. Track support status in Models overview.

#What next

Was this page helpful?