A dedicated endpoint isn't a separate URL. It's served through the same OpenAI-compatible /v1 API as every other model on the platform — the model field does the routing. Every dedicated endpoint gets a Model ID of the form dedicated/<endpoint-name>; pass it as model and the request lands on your reserved GPUs instead of the shared pool.
#Where to find the Model ID
You never construct this value by hand — copy it from the console.
Open the endpoint on the Endpoints page and expand its detail view. Under Identity, the Model ID row shows the exact string in monospace with its own copy button. If the endpoint is still provisioning and the platform hasn't published a Model ID yet, that row shows a placeholder dash until one is published.
There's no separate host or path to note down for a dedicated endpoint. The Model ID is the entire hand-off — base URL and auth stay exactly what you already use.
#Make the request
Wait for the row to read Available before your first request — that's the point at which the endpoint is fully up, and it's the simple rule to follow. What actually decides whether a call is served is narrower than the row status: it's the Ready count in the endpoint's detail view, which is how many replicas can currently serve. A Suspended endpoint has none and never serves, and a freshly created one doesn't serve until its first replica is ready. Progressing on its own doesn't mean a request fails, though — an endpoint scaling from two replicas to four reads Progressing while those two ready replicas keep serving normally. See Lifecycle: Suspend and resume for what a call to an endpoint with no ready replicas does instead of failing fast.
Authentication, streaming, and request/response shapes are identical to shared models — only the model value changes. Using the always-warm example from Autoscaling, called as dedicated/chat-prod-endpoint:
from openai import OpenAI
import os
client = OpenAI(
api_key=os.environ["OMNIVA_API_KEY"],
base_url="https://api.tokenfactory.omniva.com/v1",
)
resp = client.chat.completions.create(
model="dedicated/chat-prod-endpoint",
messages=[{"role": "user", "content": "Hello from my dedicated endpoint."}],
)
print(resp.choices[0].message.content)Streaming works the same way — add stream: true exactly as you would for any other model; see Streaming for the format and language parsers.
#Calling an endpoint that has scaled to zero
If the endpoint is autoscaled with minReplicas: 0, it sits at zero replicas once the idle window has passed, and a call arriving then has nothing ready to answer it. It behaves like any other call with no ready replica behind it: the request is held, and if none becomes ready it ends in the gateway timeout described under Lifecycle: Suspend and resume. Size your client timeouts for that before you put scale-to-zero in front of user-facing traffic.
An endpoint kept at minReplicas ≥ 1 never sits at zero in the first place. See Scale-to-zero vs a warm floor for what brings a zeroed endpoint back and the windows that control when scale-to-zero kicks in.