Autoscaling sets how much capacity your endpoint holds and how fast that changes — the trade between holding idle GPUs and making users wait. One choice dominates: whether to keep a replica warm at all times. The rest tunes how eagerly the platform adds and sheds capacity, on inference signals rather than raw CPU or memory.
#Scale-to-zero vs a warm floor
That choice is minReplicas.
minReplicas: 0(scale-to-zero) — when the endpoint is idle past the scale-to-zero window, the last replica is removed and nothing is held warm between bursts. An arriving request is what triggers scale-up again; while no replica is ready, a call is held rather than refused, and can end in the timeout described under Lifecycle: Suspend and resume. GPU-hours stop accruing while it sits at zero; best for spiky, latency-tolerant, or dev/staging traffic.minReplicas: ≥ 1(warm floor) — at least one replica is always resident, so the endpoint never sits at zero and never has to come back from it. Those replicas are held around the clock and accrue GPU-hours the whole time. Use this for latency-sensitive production traffic.
Start at minReplicas: 1 for anything user-facing. Drop to 0 only when you have measured that the traffic is bursty enough to leave the floor idle most of the time, and that your clients tolerate the wait when the endpoint has to come back from zero.
#The knobs
| Field | Type | Description |
|---|---|---|
minReplicas | integer | Lower bound on replicas. 0 enables scale-to-zero; ≥1 keeps a warm floor. Default: 1 |
maxReplicasrequired | integer | Upper bound on replicas. This is your capacity ceiling — the endpoint never scales past it. |
metric | enum | The inference signal to scale on: concurrency or kv-cache-utilization. Default: concurrency |
threshold | number | Per-replica setpoint for the chosen metric. The platform adds capacity to keep the measured value near this target. Lower = scale out sooner. For kv-cache-utilization this is a 0–1 fraction. Default: 75 (concurrency) / 0.8 (kv-cache) |
scaleUpWindow | integer (s) | How long sustained load is observed before adding capacity. Keep small so you react quickly to traffic spikes. Range 0–300. Default: 30 |
scaleDownWindow | integer (s) | How long reduced load must hold steady before a replica is removed — a dip that recovers inside the window removes nothing. Longer avoids thrashing on bursty traffic. Range 30–3600. Default: 900 |
scaleToZeroWindow | integer (s) | Idle time before the last replica is removed. Only relevant when minReplicas = 0. Range 300–86400. Default: 1800 |
Those defaults are what the console's scaling form pre-fills for you. Over the HTTP API three of them differ: omit minReplicas inside an autoscaling block there and it is 0, not 1 — scale-to-zero, not a warm floor. Send it explicitly to pin one. Omit threshold or scaleDownWindow there and the API falls back to this platform's prior defaults — 10 and 600 — not the console's 75 and 900; send them explicitly to get the console's numbers. The metric, scale-up, and scale-to-zero window defaults above apply on both paths unchanged.
#Why three separate windows
Scaling is deliberately asymmetric — reacting too slowly hurts differently in each direction:
- Scale up fast. Under-provisioning is felt immediately as queueing and rising latency. A short
scaleUpWindow(tens of seconds) means a genuine spike gets capacity quickly. - Scale down moderately. Removing a replica the instant load dips risks thrashing — shedding capacity you need again seconds later. A
scaleDownWindowof several minutes rides out normal traffic ripples. - Scale to zero reluctantly. Dropping the last replica is the hardest decision to reverse: nothing is left to answer the next request until the endpoint has scaled back up.
scaleToZeroWindowis intentionally the longest — wait until the endpoint is genuinely idle before giving up the warm floor.
The three windows let you weigh each of those independently.
#Choosing a metric
Scales on in-flight requests per replica. This is the right default for almost all chat and completion traffic: it tracks the thing users actually feel (are requests waiting?) and is robust across prompt/response shapes. threshold is the target number of simultaneous requests each replica should carry — e.g. 75.
#replicas or autoscaling — never both
An endpoint is either statically sized or autoscaled — never both.
- A configuration with
replicas(a fixed count) pins the endpoint at that size. - A configuration with an
autoscalingblock scales on load.
In the console, a configuration with both — or neither — is rejected as a 400 validation error. You can't reach either state from the form itself: Scaling mode is a two-option choice, Autoscaling or Fixed replicas, with exactly one always selected — "Select one — the fields below follow your choice." The 400 is the console's backstop behind that control, not something the form lets you produce.
The HTTP API doesn't apply the same rule to "neither", so don't carry this one across. Sending both is a 400 there too, on create and on update alike. Sending neither is not: on create the selected flavor's configured scaling default is applied and the endpoint is created, and only on update is it a 400 — a rescale call that asks for no rescale.
#Editing a live endpoint
Autoscaling settings are not fixed at deploy time. You can adjust the envelope (minReplicas/maxReplicas), the metric, the threshold, and all three scale windows on a running endpoint without redeploying — the change is applied to the live endpoint as an update, and its replicas then follow the new settings (for example, lowering maxReplicas below the current count scales the endpoint down to fit).
You can also flip between modes: switching a running endpoint to a fixed replicas count disables autoscaling (autoscaling stops and the endpoint is pinned to that count), and sending an autoscaling block re-enables it. As with create, replicas and autoscaling remain mutually exclusive on an edit.
Two fields cannot be changed one at a time on an update, which matters if you are calling the API rather than using the console:
metricandthresholdride together. Send one without the other and the update is rejected — a threshold means nothing without the metric it applies to, and vice versa.minReplicasandmaxReplicasride together. Send one without the other and the update is rejected, so adjusting only the ceiling means restating the floor alongside it. Note that an explicitminReplicas: 0counts as sending it; omitting the key entirely does not.
The console always submits complete pairs, so this only bites hand-built requests.
Start conservative, watch the endpoint under real traffic, then raise maxReplicas or lower threshold to scale out sooner — all as live edits. Model, GPU, and flavor are fixed at deploy time; scaling is not.
#Configuration examples
These are the same shapes the console's Endpoint manifest panel assembles as you configure scaling — exactly what its Copy config button gives you.
{
"name": "scale-to-zero-endpoint",
"modelId": "Omniva/glm-5.3",
"flavorName": "single-gpu-standard",
"autoscaling": {
"minReplicas": 0,
"maxReplicas": 3,
"metric": "concurrency",
"threshold": 75,
"scaleUpWindow": 30,
"scaleDownWindow": 900,
"scaleToZeroWindow": 1800
}
}For a statically sized endpoint, drop the autoscaling block and send a fixed count instead:
{
"name": "pinned-endpoint",
"modelId": "Omniva/glm-5.3",
"flavorName": "single-gpu-standard",
"replicas": 2
}You never write metric queries or scaling policies by hand. You set the envelope, the target metric and threshold, and how long demand must persist in each direction before the platform acts — the scale-up, scale-down, and scale-to-zero windows. Turning those settings into replica changes is the platform's job.
Once an endpoint is live at the scale you've configured, you call it exactly like any other model — see Call a dedicated endpoint for the Model ID and request examples.