Section
Streaming

Streaming

Server-Sent Events format and language parsers.

Streaming returns the response token-by-token as Server-Sent Events (SSE) — a long-lived HTTP response where the server pushes a sequence of data: events as the model generates them. For human-facing UIs, it dramatically reduces perceived latency: users see the first words within a few hundred milliseconds instead of waiting for the full response.

#Why stream

  • Chat interfaces — show tokens as they arrive so the user starts reading immediately.
  • Long-form generation (summaries, articles, drafts) — partial output is useful even before the model is done.
  • Cost-sensitive ops — you can inspect the early tokens and stop mid-response if the output is heading the wrong way.

When not to stream: batch processing, structured-output extraction where you need the full JSON before parsing, and any machine-to-machine call where perceived latency doesn't apply. Non-streaming responses are simpler to handle — use them by default unless a human is watching the output appear.

#SSE format

Each event is data: followed by a JSON object identical in shape to a non-streaming response, except delta replaces message and contains only the new content for that chunk. The stream ends with the literal data: [DONE].

Text
data: {"id":"chatcmpl-...","choices":[{"delta":{"role":"assistant"},"index":0}]}

data: {"id":"chatcmpl-...","choices":[{"delta":{"content":"Hello"},"index":0}]}

data: {"id":"chatcmpl-...","choices":[{"delta":{"content":" there"},"index":0}]}

data: [DONE]

The first chunk usually carries the role; subsequent chunks carry content fragments. Concatenate the delta.content values in order to reconstruct the full message.

Side-by-side, the differences from a non-streaming response:

FieldNon-streamingStreaming
choices[N].message{ role, content }(absent)
choices[N].delta(absent){ role?, content? } (partial)
choices[N].finish_reasonpopulated on the responsepopulated on the last delta
Top-level wire formatsingle JSON objectdata: <json>\n\n ... data: [DONE]\n\n

#Python

The openai SDK exposes streaming as an iterator. Set stream=True on the create call and loop with for chunk in stream.

from openai import OpenAI
import os

client = OpenAI(
  api_key=os.environ["OMNIVA_API_KEY"],
  base_url="https://api.tokenfactory.omniva.com/v1",
)

stream = client.chat.completions.create(
  model="Omniva/glm-5.3",
  messages=[{"role": "user", "content": "Write a haiku about streaming."}],
  stream=True,
)

for chunk in stream:
  print(chunk.choices[0].delta.content or "", end="", flush=True)
print()

end="" keeps tokens flowing on a single line; flush=True forces the terminal to render each chunk as it arrives instead of buffering.

#TypeScript

The Node SDK exposes streaming as an async iterable. Use for await and write each fragment to process.stdout.

import OpenAI from "openai";

const client = new OpenAI({
apiKey: process.env.OMNIVA_API_KEY,
baseURL: "https://api.tokenfactory.omniva.com/v1",
});

const stream = await client.chat.completions.create({
model: "Omniva/glm-5.3",
messages: [{ role: "user", content: "Write a haiku about streaming." }],
stream: true,
});

for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}
process.stdout.write("\n");

In Node, process.stdout.write() doesn't buffer like print — each call goes straight to the terminal. In browsers, you'll typically write into a DOM node directly (innerHTML / textContent / a streaming React component) — no flush semantics, but be aware of re-render costs if you're appending to a long string on each chunk.

#cURL

Pass "stream": true in the JSON body and -N (alias for --no-buffer) to curl so it doesn't hold output back. Raw SSE prints to the terminal; pipe it into whatever your shell can do.

curl -N https://api.tokenfactory.omniva.com/v1/chat/completions \
-H "Authorization: Bearer $OMNIVA_API_KEY" \
-H "Content-Type: application/json" \
-d '{
  "model": "Omniva/glm-5.3",
  "messages": [{"role": "user", "content": "Write a haiku about streaming."}],
  "stream": true
}'

# Crude token extractor for quick inspection:
# ... | grep -o 'content":"[^"]*' | sed 's/content":"//'

#Error handling mid-stream

If a chunk arrives with an error field set, the stream is terminating with an error — the model did not finish. Stop accumulating output, surface the error to the caller, and do not treat the partial text as a valid response.

for chunk in stream:
  if hasattr(chunk, "error") and chunk.error:
      # stream is dying; surface and stop
      raise RuntimeError(chunk.error.message)
  print(chunk.choices[0].delta.content or "", end="", flush=True)

Common mid-stream errors include the model losing capacity mid-generation, content-filter trips after partial output, and timeout on the model side. Treat partial output as diagnostic, never as the answer. See the error reference for status-code semantics — in particular, the 503 fallback pattern for upstream-loss recovery.

#Backpressure

If you break out of the loop or close the stream client-side, the server-side request continues to run for a short while before being interrupted. Don't rely on client disconnect as a cost-control mechanism — by the time the server notices and aborts, you may have already paid for most of the tokens. Set max_tokens to bound generation up front instead.

Reverse proxies may buffer SSE

Some corporate proxies and load balancers buffer SSE responses, collapsing the token-by-token experience into a single large chunk at the end. If you see the response arrive all at once instead of streaming, the proxy is the culprit. Workarounds: route the request directly (bypass the proxy), configure the proxy to flush SSE (proxy_buffering off in nginx, equivalent in your gateway), or reproduce outside the proxy to confirm the cause before debugging the client. A buffered stream that never delivers data: [DONE] still returns a successful status, so it won't stand out in the dashboard's error rate — which is why it's worth ruling the proxy out directly rather than hunting for it in aggregate.

#What next

Was this page helpful?