Early Access

Managed inference, tuned for every token.

Token Factory is an Early Access managed inference service for optimized open models. Start with serverless, on-demand access to high-performance models engineered for low-latency, high-throughput serving.

Request Early AccessAlready have access? Sign in

Early Access is sales-assisted. We'll scope your workload, provision access, and help you get started.

How it works

How Early Access Works

01
Request access
Tell us about your workload, model needs, traffic profile, and latency goals.
02
We provision your access
We match you to a supported optimized model and set up access.
03
Start serving requests
Use Token Factory through a managed endpoint, with support from our team during Early Access.
Available now

Serverless Token Factory

Run optimized open models through a managed, on-demand inference endpoint. No infrastructure to reserve, no runtime tuning to manage.

Built for teams that want:

  • Fast access to optimized model serving
  • Low-latency, high-throughput inference
  • OpenAI-compatible API access
  • Managed runtime performance without operating the stack themselves
Models

Optimized Open Models

Start with one of the supported models available through Token Factory Early Access.

Kimi K2.6
Long-horizon, autonomous coding and multi-step agents.
Available in Early Access
Kimi K3
Frontier reasoning across a million-token context, for repo-scale coding and debugging.
Available in Early Access
GLM-5.2
Agentic software engineering and coding, with selectable reasoning effort.
Available in Early Access
DeepSeek V4 Pro
Deep coding and mathematical reasoning at flagship scale.
Available in Early Access
Runtime

The fastest inference takes more than a GPU.

Renting compute is the easy part. Token Factory optimizes the runtime layer — kernels, quantization, scheduling, batching, cache reuse, and backend selection — to improve serving performance for supported models.

Tuned Runtime
Custom kernels, quantization, and model-specific profiling.
Low-Latency Serving
Speculative decoding, prefix caching, KV-cache reuse, and continuous batching.
Multi-Backend Routing
Serve workloads across optimized backends such as vLLM, SGLang, and TensorRT-LLM.
Coming soon

Dedicated Endpoints

Reserve capacity for predictable throughput and isolated serving — for supported open models and your own custom models alike.

Dedicated Endpoints for Optimized Open Models
Reserve capacity for predictable throughput, isolated serving, and OpenAI-compatible API access.
Dedicated Endpoints for Custom Models
Bring your own custom or fine-tuned model and work with Token Factory to optimize it for managed inference.
Interested in either?Contact Sales
Early Access

Start with Token Factory Early Access.

Get serverless, on-demand access to optimized open models today. Dedicated Endpoints with Optimized Open Models and Custom Models are coming soon.

Request Early AccessAlready have access? Sign in