A catalog that earns its place.
Token Factory serves a deliberately small set of open models — each one profiled, tuned, and validated end to end before it reaches the catalog.
Featured models
A selection from the Token Factory catalog.
Models marked Serverless are offered through our per-token API.
Additional models
Have another open-weight model in mind? Talk to us about deploying it on reserved capacity.
How a model earns the catalog
Every entry clears the same gate. No drive-by additions, no checkbox integrations.
We evaluate the model against real workloads — coding, agents, reasoning — before committing to serve it.
The runtime is tuned to the model: kernels, quantization, batching, cache strategy, backend fit.
It ships behind an OpenAI-compatible endpoint, and its runtime profile is re-validated as engines evolve.
Not sure which one fits?
Tell us the workload — latency budget, traffic shape, context length — and we will match it to a model and a deployment mode, and say plainly what we can commit to serving.