Z.ai
GLM 5.3
Terminal workflows, parallel tool calls and long-context sessions.
Sales-assisted Early Access
How you can run it
On Token Factory
- Serverless
- Available
- Dedicated endpoint
- Available
- Modalities
- Text
- Context window
- 1.0M tokens
The model
- Developed by
- Z.ai
- Family
- GLM
- Parameters
- 753B
- Licence
- Custom
Per-token rate for this modelDedicated ratesUpstream model card
The API is OpenAI-compatible. See the documentation
Before you request access
- What happens after I request access?
- We scope the workload with you, confirm what we can serve, set up your account and provision access.
- What does a dedicated endpoint commit me to?
- Capacity is allocated in whole 8-GPU nodes and carries a minimum term. We confirm sizing and term before anything is reserved.
- How are model changes and deprecations handled?
- Catalog changes are communicated before they take effect. Ask us for the current policy on this model.
What the call covers
Workload fit for GLM 5.3, which deployment mode suits it, what we can commit to serving and on what capacity, commercial terms, and onboarding.