Section
What are Dedicated Endpoints?

What are Dedicated Endpoints?

Reserved GPU capacity for a model you choose — what changes, what stays the same, and when to use one instead of the shared serverless pool.

Should you reserve GPUs for a model, or keep calling it from the shared serverless pool? A dedicated endpoint holds capacity for one model you choose — GPUs set aside for your traffic alone — and is reached through the same OpenAI-compatible /v1 API you already call.

#Dedicated vs. serverless

Two questions separate them: what you reserve, and what you use.

Dedicated endpointServerless
CapacityReserved GPUs for your models onlyShared pool across every workspace
What you reserveA set number of GPUs, held for as long as the endpoint is deployedNothing — there is no reservation to hold
What you useGPU-hoursTokens, per request
What you're billed onGPU-hours — the capacity you hold, whether or not it is busyTokens, per request
Latency & quotaPredictable latency on capacity that is yours alone. Your monthly token quota does not apply — you are billed for the hardware you hold, not for what you send through itGoverned by your workspace's monthly token quota
ScalingYou choose fixed replicas or autoscaling, including scale-to-zeroNo provisioning step — call any model in the Model Library directly

The reserve and use rows move independently, and that is the thing to carry with you. Usage follows the replicas that actually ran — GPU cost is where you see what a window came to. The reservation does not follow them: see GPU allotment for what an endpoint still holds while it sits at zero replicas.

#When to choose a dedicated endpoint

Reach for a dedicated endpoint when:

  • Traffic is steady enough that reserved GPUs stay busy rather than idle — the Cost view's utilization figure is how you check that after the fact.
  • The workload is latency-sensitive — keeping a warm floor (minReplicas ≥ 1 — see Autoscaling) means the endpoint never sits at zero replicas, which matters for production traffic your users are waiting on.
  • You want GPUs that are yours for the duration rather than drawn from a pool shared with other customers' traffic — the capacity is reserved, and your monthly token quota stops governing it.

Two things here aren't self-serve. Your organization's GPU allotment is provisioned by Omniva; with none provisioned for a GPU type, there is no capacity to deploy onto it. And a custom or fine-tuned model, rather than one from the Model Library, is arranged through your Omniva contact.

#What stays the same

  • The same base URL and OpenAI-compatible /v1 API.
  • The same Authorization: Bearer authentication.
  • The same request and response shapes — streaming, message format, everything. Only the model field changes.

#What you manage

Creating and running a dedicated endpoint means you're responsible for:

  • Its name and base model.
  • Its endpoint flavor — a pre-defined hardware shape (GPU count, type, and precision) chosen from what the console offers for that model.
  • Its scaling — fixed replicas or autoscaling, never both.
  • Its lifecycle — suspend, resume, change its replica count with Edit scaling, and delete.
  • Staying within your organization's GPU allotment — the capacity cap all of the above draws from.

#Before you start

  • Workspace access and a signed-in session.
  • A GPU allotment provisioned for your organization — it caps how much dedicated capacity you can hold. See GPU allotment.
  • A model in the Model Library that supports dedicated deployment. Not every model does, and the create flow only offers the eligible ones. Two things decide what appears in the picker: the model is one you can already call serverlessly — dedicated capacity is drawn from the same set of models, not a separate one — and at least one hardware shape is available for it. A third condition is checked when you deploy rather than when you pick: the model has to be live rather than paused. So a model that was paused after the page loaded can still be offered and then refused at deploy, with a 400 and the code MODEL_NOT_LIVE. The hardware-shape condition is why a model can leave the picker without anything changing about the model itself: shapes can be retired, and retiring the last one for a model takes it out of the create flow while leaving it in the Model Library. Ask your Omniva contact if a model you need isn't offered.

None of this requires writing infrastructure config by hand — naming, model choice, hardware shape, and scaling are all console steps, covered next in Create a dedicated endpoint.

#What next

Was this page helpful?