Dedicated Endpoints

Reserved capacity, isolated serving.

Dedicated Endpoints put the tuned Token Factory runtime on capacity that's yours alone — predictable throughput for production workloads, isolation for the work that needs it, and the same OpenAI-compatible interface as serverless. Available now for catalog models; custom weights are coming.

Two ways to go dedicated

Optimized open models, reserved
Run models like Kimi, GLM and DeepSeek on capacity reserved for you: predictable throughput and isolation, with the tuned runtime included.
Best for: Steady production traffic and latency budgets that need reserved headroom.
Contact Sales about this path
Your custom models
Coming soon
Bring your own or fine-tuned weights. We profile and tune the serving path for your model, then run it on dedicated capacity as a managed endpoint.
Best for: Fine-tuned or proprietary weights that need managed, isolated serving.
Contact Sales about this path

Serverless or dedicated

Same runtime, same interface — the difference is whose capacity your tokens run on.

ServerlessDedicated
CapacityShared, on-demandReserved — yours alone
ServingMulti-tenant serviceIsolated serving
Best forVariable traffic, getting startedSteady volume, isolation needs
InterfaceOpenAI-compatible APIThe same API — no code changes
Getting startedRequest Early AccessScope a reservation with sales

What reserved gets you

Built for production traffic.

Dedicated capacity is for the workloads that outgrow on-demand — steady high volume, strict isolation requirements, or latency budgets that need headroom held in reserve.

What you get:

  • Predictable throughput, held in reserve
  • Isolated serving for sensitive workloads
  • The same OpenAI-compatible interface
  • Capacity planned with our team

How to get started

Dedicated Endpoints are sales-assisted — the path starts with a conversation.

01
Talk to us
Tell us about your workload: models, traffic shape, isolation and latency needs.
02
Scope the capacity
We plan the reservation with you — model mix, throughput target, and rollout timing.
03
Go live
We provision the reservation and you serve on the same OpenAI-compatible API.

Questions, answered honestly

How is this different from serverless?
Serverless runs on shared, on-demand capacity — you pay per token and never think about infrastructure. Dedicated Endpoints reserve capacity that is yours alone: predictable throughput and isolated serving for the workloads that need it.
Will I need to change any code?
No. Dedicated Endpoints speak the same OpenAI-compatible API as serverless — the endpoint changes, your requests don't.
Can I bring my own model today?
Not yet. Reserved capacity for models like Kimi, GLM and DeepSeek is available now. Dedicated Endpoints for your own or fine-tuned weights are coming; talk to us now and we will scope that path with you ahead of it.
What happens when I contact sales now?
A scoping conversation: your models, traffic shape, isolation and latency needs, and rollout timing. No commitment implied — we size the reservation with you before anything is agreed.

Interested in either path?

Tell us which path fits and we will scope the capacity with you. Reserved capacity for catalog models is available now; custom models are coming, and early conversations shape that rollout.