Should you reserve GPUs for a model, or keep calling it from the shared serverless pool? A dedicated endpoint holds capacity for one model you choose — GPUs set aside for your traffic alone — and is reached through the same OpenAI-compatible /v1 API you already call.
#Dedicated vs. serverless
Two questions separate them: what you reserve, and what you use.
| Dedicated endpoint | Serverless | |
|---|---|---|
| Capacity | Reserved GPUs for your models only | Shared pool across every workspace |
| What you reserve | A set number of GPUs, held for as long as the endpoint is deployed | Nothing — there is no reservation to hold |
| What you use | GPU-hours | Tokens, per request |
| What you're billed on | GPU-hours — the capacity you hold, whether or not it is busy | Tokens, per request |
| Latency & quota | Predictable latency on capacity that is yours alone. Your monthly token quota does not apply — you are billed for the hardware you hold, not for what you send through it | Governed by your workspace's monthly token quota |
| Scaling | You choose fixed replicas or autoscaling, including scale-to-zero | No provisioning step — call any model in the Model Library directly |
The reserve and use rows move independently, and that is the thing to carry with you. Usage follows the replicas that actually ran — GPU cost is where you see what a window came to. The reservation does not follow them: see GPU allotment for what an endpoint still holds while it sits at zero replicas.
#When to choose a dedicated endpoint
Reach for a dedicated endpoint when:
- Traffic is steady enough that reserved GPUs stay busy rather than idle — the Cost view's utilization figure is how you check that after the fact.
- The workload is latency-sensitive — keeping a warm floor (
minReplicas≥ 1 — see Autoscaling) means the endpoint never sits at zero replicas, which matters for production traffic your users are waiting on. - You want GPUs that are yours for the duration rather than drawn from a pool shared with other customers' traffic — the capacity is reserved, and your monthly token quota stops governing it.
Two things here aren't self-serve. Your organization's GPU allotment is provisioned by Omniva; with none provisioned for a GPU type, there is no capacity to deploy onto it. And a custom or fine-tuned model, rather than one from the Model Library, is arranged through your Omniva contact.
#What stays the same
- The same base URL and OpenAI-compatible
/v1API. - The same
Authorization: Bearerauthentication. - The same request and response shapes — streaming, message format, everything. Only the
modelfield changes.
#What you manage
Creating and running a dedicated endpoint means you're responsible for:
- Its name and base model.
- Its endpoint flavor — a pre-defined hardware shape (GPU count, type, and precision) chosen from what the console offers for that model.
- Its scaling — fixed replicas or autoscaling, never both.
- Its lifecycle — suspend, resume, change its replica count with Edit scaling, and delete.
- Staying within your organization's GPU allotment — the capacity cap all of the above draws from.
#Before you start
- Workspace access and a signed-in session.
- A GPU allotment provisioned for your organization — it caps how much dedicated capacity you can hold. See GPU allotment.
- A model in the Model Library that supports dedicated deployment. Not every model does, and the create flow only offers the eligible ones. Two things decide what appears in the picker: the model is one you can already call serverlessly — dedicated capacity is drawn from the same set of models, not a separate one — and at least one hardware shape is available for it. A third condition is checked when you deploy rather than when you pick: the model has to be live rather than paused. So a model that was paused after the page loaded can still be offered and then refused at deploy, with a
400and the codeMODEL_NOT_LIVE. The hardware-shape condition is why a model can leave the picker without anything changing about the model itself: shapes can be retired, and retiring the last one for a model takes it out of the create flow while leaving it in the Model Library. Ask your Omniva contact if a model you need isn't offered.
None of this requires writing infrastructure config by hand — naming, model choice, hardware shape, and scaling are all console steps, covered next in Create a dedicated endpoint.