Creating a dedicated endpoint reserves GPUs for one model and hands you back a Model ID to call it with. The Endpoints page in the console walks you through four steps — name and model, flavor, scaling (Fixed replicas or Autoscaling), then review and deploy — and assembles the exact configuration as you go.
#Your first endpoint, end to end
- Check that your organization has a GPU allotment
Sign in, then check GPU allotment. With none provisioned for the GPU type you want, there is no capacity to deploy onto — the flavor step shows a notice rather than a list of shapes, and a deploy would be refused for capacity anyway.
- Create the endpoint
Name, model, flavor, scaling, deploy. Every field is in the create flow below, the eligible models in the Model Library, the scaling knobs in Autoscaling.
- Wait until at least one replica is ready
The row on Endpoints reaches Available — see Lifecycle: Status meanings.
- Copy the Model ID and make your call
The endpoint's detail view hands it to you — see Call a dedicated endpoint.
- Suspend it when you are finished
Optional; it stops serving and releases its GPU-allotment commitment — see Lifecycle: Suspend and resume.
#The create flow, step by step
- Name your endpoint and pick a base model
Endpoint names use lowercase letters, numbers, and hyphens — optionally split into dot-separated labels, each starting and ending with a letter or number — up to 253 characters. (This grammar is why an endpoint name can never contain a slash.)
Once the name is valid, the model picker unlocks: search the Model Library by name and choose the model to deploy. If you follow a link to a specific model that isn't eligible for dedicated deployment, the console says so and asks you to pick a different one.
Names are also unique within your workspace and permanent for the endpoint's life: creating one with a name that's already in use is rejected, and there's no rename control later — delete the endpoint and create a new one if you need a different name. That also means a new Model ID, since it's built from the name — see Call a dedicated endpoint.
- Choose an endpoint flavor
A flavor is a pre-defined hardware shape — GPU count and type, plus precision, for example "2× B200 · FP8" — with a one-line description of what it's tuned for. The flavors on offer come from the model you picked, so the choices change with the model.
Cards are labeled Recommended for the model's suggested default, Only flavor when it's the sole option, and Doesn't fit when your remaining GPU allotment for that type can't cover even one replica. A "Doesn't fit" card is dimmed rather than hidden, with the reason attached: "This shape doesn't fit — try a smaller flavor." See GPU allotment for what's behind that limit.
Occasionally a model resolves to zero flavors: it isn't eligible for a dedicated GPU endpoint yet, so pick another base model. If the flavors fail to load instead, the step offers a retry in place.
- Choose fixed replicas or autoscaling
Scaling is one or the other, never both. Fixed replicas runs exactly the count you set; Autoscaling adds and removes replicas to match demand between a minimum and a maximum.
Picking a flavor pre-fills this step from that flavor's own recommendation — a replica count for fixed, or a min/max envelope for autoscaling — labeled "suggested by" plus the flavor's name. A flavor only ever suggests the replica numbers, and not every flavor carries a suggestion — where it doesn't, the form starts you at a minimum of 1 and a maximum of 3, with no "suggested by" label.
Every number here is yours to change, including the minimum. The rest of the policy is a platform default whether you picked a flavor or not: the
concurrencymetric at a threshold of 75, and scale-up, scale-down and scale-to-zero windows of 30, 900 and 1800 seconds — all adjustable, and all covered in Autoscaling.Setting the minimum to 0 enables scale-to-zero: no replica is left running once the endpoint goes idle, and nothing is ready to answer the next call until it has scaled back up. That is about what runs, not about what you have reserved — see GPU allotment for what an idle endpoint still holds, and Autoscaling: Scale-to-zero vs a warm floor for the trade-off.
Calling the API directly? Send both replica numbersThe 1 and 3 above are the form's starting values, not the API's. Over the API the replica numbers aren't pre-filled: an omitted
minReplicasis read as 0, which is scale-to-zero — fine if that's what you want, surprising if you expected a warm floor — andmaxReplicasis required, so leaving it out is rejected rather than defaulted. Set both explicitly and the two surfaces behave identically. See Endpoints API. - Review the manifest and deploy
The Endpoint manifest panel assembles the exact configuration as you complete each step — identity (name, model, flavor) and scaling (policy and windows) — and its Copy config button copies precisely what deploying will send. Once every step is complete, the panel reads "manifest ready" and Deploy endpoint is enabled.
#What clamping looks like near capacity
If your organization is close to its GPU allotment for a flavor's type, the form surfaces that where it can, rather than leaving you to find out at submit time:
- The replica (or min/max) input's upper bound is capped at your remaining headroom for that type.
- A flavor that doesn't fit at all is dimmed with a Doesn't fit badge rather than removed — you can still see it exists.
- If your organization's remaining allotment can't fit any flavor, the deploy button explains why:
Org GPU allotment fully committed (<committed> of <allotment>). - If your organization isn't provisioned yet, the flavor step is replaced with a notice instead of a form.
All of this draws on the same allotment the GPU allotment ledger tracks — nothing here is a separate limit.
#After you create it
Deploying keeps you on this form, confirms the endpoint is being created, and offers a link on to the Endpoints list. The endpoint then moves from Progressing to Available as its GPUs are provisioned; Suspended and Error are the other two states you'll see on the row.
Open the row to expand its detail view. Once the endpoint is ready, the Model ID appears under Identity — a copyable string you pass as the model field on any OpenAI-compatible request. That's the hand-off: no separate URL to note down, just the Model ID. See Call a dedicated endpoint for how to use it, and Lifecycle for what the detail view's Replicas and Ready counts mean.