Section
Models overview

Models overview

Choose, compare, and launch the right model from the Model Library and Token Factory API.

Token Factory models are both dashboard-browsable and API-discoverable. Use this page when you need to pick a model, verify the facts that matter, and launch it through the right path.

Browse

Model Library in the Token Factory dashboard

List

GET /v1/models

Launch

Token-based pricing

#Pick by job, then verify by facts

Start from the workload. The catalog is useful because it connects task fit, model facts, benchmark context, and launch path in one place.

If you needStart withVerify before launch
Chat or agentsChat-capable live modelsTool-use support, context window, price, and benchmark fit
Embeddings or RAGEmbedding models/v1/embeddings support, vector dimensions, retrieval quality, and cost
Long contextContext filter ≥ 128K or ≥ 256KWhether your workload actually needs the headroom
Profile, don't guess

Pick the smallest model that plausibly fits the task, then test against your real prompts, latency budget, and cost target. Premium-by-default burns spend you may not need.

#Compare live models in the dashboard

Open Model Library to browse the workspace catalog — it is part of the signed-in console, so it asks you to sign in first. The Model Library is a workspace-scoped catalog of every model available to your team. Each entry carries:

  • A model ID in the form provider/model-name. The prefix carries meaning: an Omniva/… id (for example, Omniva/glm-5.2) is an Omniva-optimized build of an open model — tuned and quantized for low-latency serving — while the same base model may also be offered under its upstream author prefix (for example, zai-org/GLM-5.2) as standard open weights. Always copy the exact id from the catalog.
  • Modality — text or embedding.
  • Context window length — token capacity for the model's input plus output.
  • Cost class — Token-based pricing.
  • Lifecycle stage — GA, Beta, or Coming Soon.
  • Deployment-option toggles — which deploy paths the model supports.
  • Benchmark and model-card notes — when published by the provider.

Filter by modality, sort by recently added, or open a model entry to see its complete specifications, including supported context length, function-calling capability, and any model-card notes from the provider.

Benchmarks are selection evidence, not marketing claims

Use benchmark sections to compare live models when data is present. If a specific metric is missing, do not infer it from another model or provider. Treat missing metrics as unknown until Token Factory publishes a sourced value.

#Discover callable models from the API

Use GET /v1/models when your app needs to discover callable models programmatically from the public Token Factory API surface.

Bash
curl https://api.tokenfactory.omniva.com/v1/models \
  -H "Authorization: Bearer $OMNIVA_API_KEY"

The public response is an OpenAI-compatible list, shaped for client discovery. The example below is trimmed to one entry for brevity — the actual response includes every model in your workspace catalog:

JSON
{
  "object": "list",
  "data": [
    {
      "id": "Omniva/glm-5.2",
      "object": "model",
      "created": 0,
      "owned_by": "Omniva"
    }
  ]
}

That API response is intentionally smaller than the dashboard catalog. The dashboard presents additional comparison fields — capability chips, example pricing, benchmark sections, and deploy affordances — for comparing models visually. Build API clients against the public /v1/models response for programmatic model discovery.

Public list and dashboard catalog are related, not identical

GET /v1/models tells your app which model IDs are callable through the public API. The dashboard catalog tells a human how to compare, benchmark, and launch those models inside Token Factory.

#Launch with Token-based pricing

During Early Access, models launch through Token-based pricing: OpenAI-compatible /v1 traffic billed per token, with no dedicated endpoint to manage. It keeps the request shape close to OpenAI-compatible calls and removes dedicated-endpoint management overhead — the right default for variable traffic and quick prototyping.

Dedicated Endpoints — reserved capacity for eligible models

For steadier, latency-sensitive traffic, eligible models also support Dedicated Endpoints — GPUs reserved for one model rather than tokens consumed per call, deployed and managed from the console. See What are Dedicated Endpoints? to compare the two paths.

#Models you can try today

The models available through Token Factory Early Access are listed on the open models catalog — what each one is for, its context window, and how it is served.

For embeddings, the embeddings guide uses nomic-ai/nomic-embed-text as its canonical example.

Filter the Model Library by modality or capability chip to see which of those models your own workspace can call, and copy the model ID from there for GET /v1/models.

#What next

Was this page helpful?