Reserved capacity, isolated serving.
Dedicated Endpoints put the tuned Token Factory runtime on capacity that's yours alone — predictable throughput for production workloads, isolation for the work that needs it, and the same OpenAI-compatible interface as serverless. Available now for catalog models; custom weights are coming.
Two ways to go dedicated
Serverless or dedicated
Same runtime, same interface — the difference is whose capacity your tokens run on.
| Serverless | Dedicated | |
|---|---|---|
| Capacity | Shared, on-demand | Reserved — yours alone |
| Serving | Multi-tenant service | Isolated serving |
| Best for | Variable traffic, getting started | Steady volume, isolation needs |
| Interface | OpenAI-compatible API | The same API — no code changes |
| Getting started | Request Early Access | Scope a reservation with sales |
What reserved gets you
Built for production traffic.
Dedicated capacity is for the workloads that outgrow on-demand — steady high volume, strict isolation requirements, or latency budgets that need headroom held in reserve.
What you get:
- Predictable throughput, held in reserve
- Isolated serving for sensitive workloads
- The same OpenAI-compatible interface
- Capacity planned with our team
How to get started
Dedicated Endpoints are sales-assisted — the path starts with a conversation.
Questions, answered honestly
Interested in either path?
Tell us which path fits and we will scope the capacity with you. Reserved capacity for catalog models is available now; custom models are coming, and early conversations shape that rollout.