Early Access

Managed inference, tuned for every token.

One platform for serverless inference and dedicated endpoints. Built for speed, scale, and efficiency.

Request Early AccessAlready have access? Sign in

Early Access is sales-assisted. We'll scope your workload, provision access, and help you get started.

How it works

How Early Access Works

01
Request access
Tell us about your workload, model needs, traffic profile, and latency goals.
02
We provision your access
We match you to a supported optimized model and set up access.
03
Start serving requests
Use Token Factory through a managed endpoint, with support from our team during Early Access.
Available now

Serverless Token Factory

Run optimized open models through a managed, on-demand inference endpoint. No infrastructure to reserve, no runtime tuning to manage.

Built for teams that want:

  • Fast access to optimized model serving
  • Low-latency, high-throughput inference
  • OpenAI-compatible API access
  • Managed runtime performance without operating the stack themselves
Models

Optimized Open Models

Start with one of the supported models available through Token Factory Early Access.

GLM 5.3
Multi-step agent execution across terminal work and tool calls.
Available in Early Access
DeepSeek V4 Pro
Deep coding and mathematical reasoning with high accuracy.
Available in Early Access
Runtime

Same GPUs. Faster tokens.

Renting compute is the easy part. Token Factory optimizes the runtime layer — kernels, quantization, scheduling, batching, cache reuse, and backend selection — to improve serving performance for supported models.

Tuned Runtime
Custom kernels, quantization, and model-specific profiling.
Low-Latency Serving
Speculative decoding, prefix caching, KV-cache reuse, and continuous batching.
Multi-Backend Routing
Serve workloads across optimized backends such as vLLM, SGLang, and TensorRT-LLM.
Dedicated Endpoints

Capacity that is yours alone.

Reserve capacity for predictable throughput and isolated serving — available now for supported open models, with custom models coming.

Dedicated Endpoints for Optimized Open Models
Reserve capacity for predictable throughput, isolated serving, and OpenAI-compatible API access.
Dedicated Endpoints for Custom Models
Coming soon
Bring your own custom or fine-tuned model and work with Token Factory to optimize it for managed inference.
Early Access

Start with Token Factory Early Access.

Get serverless, on-demand access to optimized open models — or reserve dedicated capacity for them on the same tuned runtime.

Request Early AccessAlready have access? Sign in