Section
Quickstart

Quickstart

Make your first OpenAI-compatible API call to Token Factory.

Make one OpenAI-compatible /v1 chat request against Token Factory's live endpoint.

Already onboarded and signed in? You're one API key away from your first call.

Endpoint

https://api.tokenfactory.omniva.com/v1

Env var

OMNIVA_API_KEY

Example model

Omniva/glm-5.3

  1. Get an API key

    First, make sure you have workspace access and you're signed in — see Workspace access (new customers request access).

    Once signed in, open API keys and click Generate. The full value is shown once at creation time — copy it now and keep it out of source control.

    For rotation and leak-response, see Authentication.

  2. Set your environment

    Export the key under its canonical name. Every snippet below reads from it.

    export OMNIVA_API_KEY="sk-oais..."
  3. Install a client

    Token Factory's wire shape matches OpenAI's at https://api.tokenfactory.omniva.com/v1. Use the official OpenAI SDK or any HTTP client.

    pip install openai
  4. Make your first call

    Send a chat completion against https://api.tokenfactory.omniva.com/v1.

    from openai import OpenAI
    import os
    
    client = OpenAI(
      api_key=os.environ["OMNIVA_API_KEY"],
      base_url="https://api.tokenfactory.omniva.com/v1",
    )
    
    resp = client.chat.completions.create(
      model="Omniva/glm-5.3",
      messages=[
          {
              "role": "user",
              "content": "Give me one sentence on retrieval-augmented generation.",
          }
      ],
    )
    print(resp.choices[0].message.content)

#Verification

That's it — first call worked.

You should see a short answer explaining retrieval-augmented generation, for example:

Retrieval-augmented generation combines a retrieval step over an external knowledge source with a generative model so the output is grounded in cited evidence.

Exact wording varies between runs. If you got something back, the rest of Token Factory is one call away.

#What just happened?

Request

Your OpenAI-compatible client sent a request to Token Factory's /v1/chat/completions endpoint. Read more →

Auth

OMNIVA_API_KEY authenticated the request without putting the secret in source code. Read more →

Model

model selected the hosted chat model. Replace it with any available chat model from Models overview.

#Stream tokens (optional)

Pass stream: true when you want tokens to arrive incrementally — typical for chat UIs that render a typing effect. Deeper concerns (reconnects, partial-tool calls, cancellation, plus TypeScript and cURL variants) live in the Streaming guide.

stream = client.chat.completions.create(
  model="Omniva/glm-5.3",
  messages=[{"role": "user", "content": "Count to five slowly."}],
  stream=True,
)
for chunk in stream:
  print(chunk.choices[0].delta.content or "", end="", flush=True)

#What next

#Troubleshooting

Common first-run fixes
  • 401 Unauthorized: confirm OMNIVA_API_KEY is set in the same shell where you run the command.
  • 403 Forbidden: your key is valid but doesn't have access to the requested resource. Check workspace membership.
  • 400 or 404 for the model: choose a model from Models overview and paste its exact ID.
  • 429 Too Many Requests: check your quota and retry after the rate-limit window.
  • No output from streaming cURL: include -N so cURL does not buffer the stream.

Hitting an SSL trust-chain error? See Errors → SSL. Model returning 503 instead of generating? See the fallback pattern.

Was this page helpful?