Models

A catalog that earns its place.

Token Factory serves a deliberately small set of open models — each one profiled, tuned, and validated end to end before it reaches the catalog.

Request Early AccessAlready have access? Sign in

Additional models

Z.ai

GLM 5.2

Parameters
753B
Context
1.0M tokens
View model

DeepSeek

DeepSeek V4 Pro

Parameters
1.6T
Context
1.0M tokens
View model

DeepSeek

DeepSeek V4 Flash

Parameters
291B
Context
1.0M tokens
View model

OpenAI

gpt-oss-120b

Parameters
117B
Context
131K tokens
View model

Have another open-weight model in mind? Talk to us about deploying it on reserved capacity.

How a model earns the catalog

Every entry clears the same gate. No drive-by additions, no checkbox integrations.

01
Evaluate

We evaluate the model against real workloads — coding, agents, reasoning — before committing to serve it.

02
Tune

The runtime is tuned to the model: kernels, quantization, batching, cache strategy, backend fit.

03
Serve and hold

It ships behind an OpenAI-compatible endpoint, and its runtime profile is re-validated as engines evolve.

Not sure which one fits?

Tell us the workload — latency budget, traffic shape, context length — and we will match it to a model and a deployment mode, and say plainly what we can commit to serving.