Pay per token, or reserve the node.
Serverless inference is metered per token, per model, with nothing to reserve up front. Dedicated Endpoints reserve single-tenant GPU nodes by the hour.
Serverless rates
Every model is priced individually, and each table below states the unit it is priced in. Rates are the same for every workspace.
| Model | Input | Output |
|---|---|---|
| GLM 5.3 Flash | $0.15 | $0.50 |
| GLM 5.3 | $1.40 | $4.40 |
Every model listed here is sold per token; models served only on dedicated capacity are on the models page. For anything else, email [email protected] about that model, or about one that isn't listed.
Serverless inference is available to prepaid credit customers only. Contact us to get set up.
How tokens are counted, quotas, and the cost formula — in the docsDedicated Endpoint rates
Single-tenant GPU capacity reserved for your workloads. Priced per GPU-hour, sold by the node, with a 1 month minimum commitment.
| Hardware | Per GPU | Per node |
|---|---|---|
| NVIDIA HGX B200 | $9.00 | $72.00 |
| NVIDIA HGX B300 | $12.00 | $96.00 |
Dedicated Endpoints require a minimum 1 month commitment. Volume and longer commitments qualify for discounted rates — contact sales.
What a Dedicated Endpoint gets you — isolation, throughput, custom modelsStart with usage-based serving.
Request Early Access and we'll scope your workload — you pay for what you process, from the first token.