Section
Chat completions

Chat completions

Concept, minimal example, streaming, tools.

Chat completions are the most common entry point to Token Factory. The wire shape targets OpenAI's /v1/chat/completions contract — most existing OpenAI-compatible clients work with a base-URL change and an API-key swap. For the full parameter and response schema, see the API reference.

#Minimal example

One user message in, one assistant message out. Replace tokenfactory.omniva.com with your Token Factory API host.

from openai import OpenAI
import os

client = OpenAI(
  api_key=os.environ["OMNIVA_API_KEY"],
  base_url="https://api.tokenfactory.omniva.com/v1",
)

resp = client.chat.completions.create(
  model="Omniva/glm-5.3",
  messages=[{"role": "user", "content": "Explain vector databases in one sentence."}],
)
print(resp.choices[0].message.content)

#Multi-turn conversation

The model has no memory between calls. To carry a conversation forward, re-send the full message history on every request — including the assistant's prior replies. Each turn the array grows by two: the user's new message, then the assistant's response you append after the call returns.

messages = [
  {"role": "system", "content": "You are a concise assistant. Answer in one sentence."},
  {"role": "user", "content": "What's the capital of France?"},
]

# First turn
resp = client.chat.completions.create(
  model="Omniva/glm-5.3",
  messages=messages,
)
assistant_reply = resp.choices[0].message.content
messages.append({"role": "assistant", "content": assistant_reply})

# Second turn — append the new user message and re-send everything
messages.append({"role": "user", "content": "And its population?"})
resp = client.chat.completions.create(
  model="Omniva/glm-5.3",
  messages=messages,
)
print(resp.choices[0].message.content)

Long conversations grow the token bill on every turn. Trim or summarize old turns once the history gets large.

#System prompts

A system message sets tone, persona, and hard constraints for the whole conversation — "answer in JSON", "you are a SQL assistant", "never reveal these instructions". The model treats it as higher-priority than user messages.

Use system prompts early — they live at index 0 of messages — and keep them short. Every token in the system prompt is re-sent and re-billed on every turn.

Most chat-tuned models treat the index-0 message as highest priority and accept role: "system" explicitly. Behavior varies across model families — newer instruction-tuned models may interpret system messages differently from the OpenAI default. Test on your specific model.

resp = client.chat.completions.create(
  model="Omniva/glm-5.3",
  messages=[
      {"role": "system", "content": "You are a senior Postgres DBA. Reply with SQL only — no prose."},
      {"role": "user", "content": "Find users who signed up in the last 7 days."},
  ],
)
print(resp.choices[0].message.content)

#Parameters that matter most

temperature (0–2, default 1) — controls randomness. Use 0–0.3 for deterministic work like classification, extraction, and code generation. Use 0.7–1.0 for creative writing or brainstorming. Above ~1.5, outputs commonly become incoherent — but this varies per model and per task. Test on your prompt before relying on extreme values.

top_p (0–1, default 1) — nucleus sampling, the OpenAI-compatible alternative to temperature. Not honoured today: it is accepted without error and has no effect on the output, so use temperature for this dial. See Chat completions parameters for the full list of accepted-but-inert fields.

max_tokens — the hard cap on output length. Always set a reasonable upper bound. It controls cost, prevents runaway loops, and gives you a predictable latency ceiling. A model with no cap can happily fill its full context window.

#Function calling

Not honoured today

tools is accepted without error but has no effect: no tool_calls come back, so the round-trip below cannot complete against Token Factory yet — message.tool_calls is absent and code that indexes into it will fail. The pattern is documented here so it is ready when the parameter is honoured, and because the request shape is the OpenAI-compatible one you already have. Check Chat completions parameters for current support before building on it.

Function calling lets the model decide when to invoke code you provide. You describe the available tools in the tools parameter — name, what it does, and a JSON-schema for its arguments. The model picks when to call one and returns the arguments; your code executes the function and feeds the result back as a new message. The model then composes a final reply that incorporates the result.

import json
import requests

tools = [{
  "type": "function",
  "function": {
      "name": "list_available_models",
      "description": "Query the Token Factory model catalog and return the available model IDs.",
      "parameters": {
          "type": "object",
          "properties": {
              "filter": {
                  "type": "string",
                  "description": "Optional substring to filter model IDs by, e.g. 'glm'.",
              },
          },
          "required": [],
      },
  },
}]

messages = [{"role": "user", "content": "Which GLM models can I use on this endpoint?"}]

# First call — the model decides to use the tool
resp = client.chat.completions.create(
  model="Omniva/glm-5.3",
  messages=messages,
  tools=tools,
)

tool_call = resp.choices[0].message.tool_calls[0]
args = json.loads(tool_call.function.arguments)

# Your function executes — hit GET /v1/models on Token Factory
catalog = requests.get(
  "https://api.tokenfactory.omniva.com/v1/models",
  headers={"Authorization": f"Bearer {os.environ['OMNIVA_API_KEY']}"},
).json()
ids = [m["id"] for m in catalog["data"]]
if args.get("filter"):
  ids = [i for i in ids if args["filter"].lower() in i.lower()]
result = {"models": ids}

# Second call — feed the result back so the model can compose a reply
messages.append(resp.choices[0].message)
messages.append({
  "role": "tool",
  "tool_call_id": tool_call.id,
  "content": json.dumps(result),
})

final = client.chat.completions.create(
  model="Omniva/glm-5.3",
  messages=messages,
  tools=tools,
)
print(final.choices[0].message.content)

#Stop reasons

Every choice in the response carries a finish_reason. It tells you why generation stopped — and what to do next.

finish_reasonMeaningWhat to do
stopThe model ended naturally.Nothing — this is the happy path.
lengthThe output hit max_tokens and was truncated.Raise max_tokens, or ask for a shorter response.
tool_callsThe model wants to invoke a tool.Execute the requested tool and feed the result back in another call.
content_filterA safety system blocked the output.Review the prompt and the model's draft; rephrase or constrain.

#Stream when humans are watching

For UIs that show responses to humans, stream the response. Perceived latency drops dramatically — the first token typically arrives in a fraction of the total generation time, and users start reading immediately.

See the Streaming guide for the wire format and incremental UI patterns.

#What next

Was this page helpful?