Short definitions for the terms that recur across these docs. If a term you saw on a page isn't covered here, file an issue.
#API key
A secret credential of the form sk-oais-... that authenticates requests to the public /v1 API. Keys are scoped to a single workspace; quotas, billing, and Observability all attribute to the owning workspace.
#Bearer token
The HTTP authorization scheme used by /v1. Pass your API key as Authorization: Bearer sk-oais-... — there is no separate token-exchange step.
#Cold start
The gap between an autoscaled Dedicated Endpoint sitting at zero replicas and being able to serve again. What brings it back, and what happens to a call that arrives in the meantime, is covered under Scale-to-zero vs a warm floor. How long the gap lasts is not something the platform reports. Keeping a warm floor means the endpoint never sits at zero.
#Dedicated Endpoint
Reserved GPU capacity serving a single model, called through the same /v1 API using a dedicated/<endpoint-name> Model ID. Distinct from a REST endpoint such as /v1/chat/completions.
See What are Dedicated Endpoints?.
#Endpoint flavor
A pre-defined hardware shape — GPU count, type, and precision — offered for a model on a Dedicated Endpoint. Which flavors are offered depends on the model you're deploying. Your GPU allotment doesn't change that list — a shape that won't fit your remaining headroom is shown and marked as not fitting, rather than hidden.
#GPU allotment
The per-organization, per-GPU-type cap on GPU capacity committed across your Dedicated Endpoints. Checked on create, resume, and scaling changes; never on suspend, delete, or a read. See GPU allotment.
#GPU-hour
The unit dedicated-endpoint GPU usage is measured in — how many GPUs an endpoint holds, and for how long. The GPU cost view states the arithmetic in full and reports it per window. Distinct from GPU allotment, which caps what you may reserve rather than recording what you used.
#Model ID
The string identifier for a model in API requests, formatted as provider/model-name. The prefix carries meaning: an Omniva/… id (for example Omniva/glm-5.2) is an Omniva-optimized build of an open model — tuned and quantized for low-latency serving — while the raw upstream open weights are offered under the author prefix (for example zai-org/GLM-5.2). Echoed back in the model field of responses.
#Model Library
The browseable catalog of available models in the Token Factory app — filterable by provider, modality, and status. The same set is exposed programmatically through GET /v1/models.
#Playground
The in-app chat and completion UI for trying a model interactively without writing code. The Playground calls the same /v1 gateway as your API key, but authenticates with your signed-in session instead of a key — not a substitute for an API key in production.
#Production gateway
The deployed /v1 endpoint at https://api.tokenfactory.omniva.com/v1. The contract documented across these pages.
#Replica
One running instance of a Dedicated Endpoint's model server. minReplicas and maxReplicas bound how many can run at once; a fixed endpoint runs an exact count instead.
#Scale-to-zero
minReplicas: 0 on an autoscaled Dedicated Endpoint: once the endpoint is idle past its scale-to-zero window, the last replica is removed. See cold start for what a call that arrives while it sits at zero runs into. Contrast with warm floor.
#Token-based pricing
The pay-as-you-go pricing mode: billed per 1M tokens, with input and output priced separately per model. Best fit for variable traffic and prototyping. GA today via /v1/chat/completions and /v1/embeddings.
#Token Factory
Omniva's managed inference product with a console and an OpenAI-compatible /v1 API. The production gateway is live for chat and embeddings.
#Utilization
Three different figures wear this name, and they measure different things:
- On the GPU cost view — how busy the GPUs were, averaged evenly across collection windows rather than weighted by GPU count, so a one-GPU hour counts the same as a thirty-GPU one. A true 0% and a missing measurement render identically there.
- On the GPU allotment ledger — how much of your organization's per-type cap is committed. Capacity accounting; says nothing about how busy anything is.
- As
kv-cache-utilizationin Autoscaling — the fraction of the serving engine's KV-cache memory in use, the alternative scaling metric.
#Warm floor
minReplicas ≥ 1 on an autoscaled Dedicated Endpoint: at least one replica always stays resident, so the endpoint never sits at zero. That floor is held around the clock and accrues GPU-hours the whole time. Contrast with scale-to-zero.
#Workspace
The scope unit for everything that costs money or carries quota: API keys, monthly token caps, billing, and Observability dashboards. Users belong to one or more workspaces; each request attributes to exactly one.