Section
Lifecycle

Lifecycle

What Suspend, Resume, Delete, and Edit scaling actually do to a dedicated endpoint — and to your GPU allotment.

Stop an endpoint, bring it back, resize it, or delete it: four decisions that differ in whether the configuration survives, whether your GPU allotment is released, and whether you can undo them. Suspend, Resume, and Delete are on the endpoint's row on the Endpoints page, Edit scaling inside its expanded detail view. That's the full set.

There's no restart

The console doesn't have a standalone restart action. If you want an endpoint to fully re-provision, suspend it and then resume it — two separate, explicit actions rather than one button.

#Status meanings

StatusWhat it means
AvailableServing traffic normally.
ProgressingDeploying, scaling, or otherwise catching up to a new target — not yet steady.
SuspendedStopped by request — not serving, and reversible with Resume.
ErrorSomething needs attention — open the endpoint's detail view for specifics (see Reading the Conditions panel below).

Those four are the console's own vocabulary. Driving the same endpoint over HTTP, the HTTP API reports a coarser status.state off the same conditions — a different set of words, not a translation of these. See The Endpoint object.

Ready isn't a status you'll see on the row — it's one of a pair of fields inside the detail view's Runtime section, distinct from the status above it: Replicas is how many replicas currently exist, and Ready is how many of those can serve traffic. The two diverge while the endpoint is Progressing, and settle back together once it's Available.

Both figures are observations, not the number you asked for. If capacity is short, Replicas reports what actually got created — which can sit below the count you configured, and stay there. On a fixed-count endpoint, a Replicas figure below the replicas you set is itself the signal: it is not merely starting slowly, it could not get what it asked for. On an autoscaled endpoint the comparison is different — sitting below maxReplicas is the normal case, because that number is a ceiling rather than a target. Compare against the scaling you set in Edit scaling rather than assuming Replicas reflects it. The monitoring view (see Monitor dedicated endpoints) adds a fifth state, Unavailable, for an endpoint that's been deleted or can no longer be resolved — you won't see that one on the Endpoints table itself.

#Reading the Conditions panel

Open an endpoint's expanded detail view and look at the Runtime section: alongside the replica counts is a Conditions panel — this is the detail the Error row above sends you to.

Each row there is one condition: a type, a status of True, False, or Unknown, a reason, and the time that condition last changed. The types you can actually see are Available, Progressing, Degraded, Suspended, ReplicaFailure, and — on multi-node endpoints only — UpdateInProgress. Read status literally — True means that condition currently holds, False means it's cleared, and Unknown means it hasn't been determined yet, so read it as neither. That matters because a Degraded condition reading False is good news, not bad: it means the endpoint is not currently degraded. Only conditions reading True are asserting anything; the rest of the list is informational.

The panel never shows a condition's full message — only its structured type, status and reason, plus when it last changed. A Degraded condition asserting True is what turns the row's status to Error. Occasionally a condition arrives with no reason at all; when that happens the row still tells you the type, the status and the timing, and those are what to go on. If its reason doesn't point to something you can fix yourself (an oversized scaling target, for instance), try a Suspend and Resume — see below — and if it persists, contact Omniva support with the endpoint name and the condition's type, status, and reason.

Two things about this panel are worth knowing before you rely on it.

Degraded is usually rebuilt, but it can persist. On the ordinary refresh path it is cleared and rebuilt as the endpoint's state is re-evaluated, so a fault that resolved itself leaves nothing behind — if a row turned Error and is no longer Error, the condition that caused it may already be gone. It does not always clear, though: a Degraded raised while the endpoint was still initializing can stay asserted, and one raised before a suspend survives the suspend. If you see Degraded on an endpoint you deliberately suspended, compare the time the condition last changed against when you suspended it — if the condition is the older of the two, it predates the suspension and isn't a live fault.

An endpoint that cannot start says so in this panel — don't wait for the status column. Whatever the row reads, an endpoint that has sat provisioning far longer than its peers is not necessarily still starting. Which condition carries the news depends on what went wrong:

  • No capacity to run on, or the model's image cannot be pulled — the common cases — show up as Progressing reading False with the reason ProgressDeadlineExceeded. This is the pair to look at first.
  • ReplicaFailure is narrower than it sounds: it appears when the replicas could not be created at all, rather than created but with nowhere to run. It also only surfaces on single-node endpoints.

In every case the condition's reason is the specific cause, and it is what to quote to support along with the endpoint name.

#Suspend and resume

Suspend and Resume share one button on the row, with no confirmation dialog either way.

Suspend stops the endpoint from serving and its commitment stops counting against your organization's allotment immediately — as the allotment panel puts it: "Suspended endpoints don't count toward your allotment." On success a toast names what stopped counting — Freed {commitment}× {TYPE}. — or, when the ledger can't supply those figures, just confirms the suspend. Suspending does not wait for in-flight work: replicas get a short shutdown window and are then stopped, so a long generation running when you click can be cut off mid-stream — the same caution Delete carries below. If that matters, drain traffic off the Model ID first.

Resume puts the endpoint back and re-checks capacity from scratch: while it sat suspended, the capacity it used to hold may already have been claimed by another endpoint. Where the row can tell that a resume would put it over the allotment, it says so and disables the button: Resuming needs {commitment}× {TYPE} — only {remaining} remaining in your GPU allotment., or, when the specifics aren't available, Resuming would exceed your GPU allotment. That is the row being helpful, not the check itself — an over-cap resume comes back as a 409 on submit, which is also what happens if the capacity moved after the row was drawn.

Occasionally a resume clears that check but still loses a capacity race a moment later — that shows up as a short GPU allotment exceeded (or Capacity check unavailable) badge on the row, with the full reason on hover.

A suspended endpoint's Model ID stops serving the moment Suspend completes, and a newly created endpoint isn't serving yet either until its first replica is ready. It's the Ready count that settles this, not the row's status: a Progressing endpoint that already has ready replicas — one scaling up, say — keeps serving throughout. Calling a Model ID with no ready replicas does not fail fast: the request is held open for minutes and then returns a gateway timeout whose body is not JSON, so a client that assumes an error envelope will fail while parsing it.

Set a client timeout of your own, short enough that your deadline fires well before the gateway's and you get a clean, catchable timeout instead. Resume a suspended endpoint, or wait for one with no ready replicas yet to reach Available, before retrying — a resume can take a minute or two to serve again.

#Delete

Delete opens a confirmation first — there's no accidental one-click delete. Confirming starts deletion and releases its allotment commitment. The row goes straight away, but how long the endpoint takes to finish shutting down isn't reported back — don't read the confirmation as proof that in-flight requests have drained.

If there's any chance you'll want this endpoint again, suspend it instead: suspension preserves its name and configuration and releases its allotment commitment just as deletion does. Delete is the only one of the two you can't take back.

#Edit scaling

Edit scaling lives inside the endpoint's expanded detail view, not the row itself. It opens the same scaling controls you saw at create time — Scaling mode (Fixed replicas or Autoscaling), the replica envelope, metric, threshold, and the three scale windows — see Autoscaling for what each one does. Scaling is the one part of a deployed endpoint you can still change. Nothing else about the endpoint (name, model, flavor) is editable here.

An edit is checked against your organization's GPU allotment the same way a create is — with one difference: the endpoint's own current reservation is credited back before the new request is measured against your remaining headroom. Shrinking a running endpoint, or leaving it the same size, is never blocked by capacity the endpoint already holds. Only a net increase beyond what your remaining headroom (plus that credited-back reservation) can cover gets capped or rejected. One more thing shrinking shares with suspend: a removed replica gets the same short shutdown window and is then stopped, so a request in flight on it can be cut off rather than run to completion.

Edit scaling works identically on a suspended endpoint — same button, same fields. Since a suspended endpoint holds no GPU reservation, saving the edit provisions nothing by itself; the new size takes effect, and is checked against your allotment, at Resume. That makes shrinking a suspended endpoint's envelope the standard move when Resume is disabled for capacity — reduce it, save, then resume against the smaller size.

#Scale to zero, suspend, or delete

Three ways to stop an endpoint doing work, with three different consequences. Note that the second and third columns are not the same question: an endpoint can keep its name and every setting and still have nothing able to answer a request.

DecisionConfiguration retainedCan serve nowAllotment commitmentReversal
Scale to zeroYes — nothing about the endpoint changesNo, while it sits at zero replicasHeld in full — see GPU allotmentNothing to reverse; the endpoint is still deployed and still autoscaling.
SuspendYes — name, model, flavor, and scaling settings are all keptNoReleasedResume, which is checked against your allotment and can be refused.
DeleteNo — the name and its configuration are goneNoReleasedNone.

Of these, only Resume — and an Edit scaling that raises the ceiling — can add to your committed total, so only those two can be blocked for capacity. See GPU allotment for the arithmetic behind that column and what a 409 looks like when a create, resume, or scale-up is rejected.

#What next

Was this page helpful?