Zhipu Z.ai full lineup · text generation · multimodal understanding · image/video/voice · vector retrieval — Sources: BigModel model overview · Pricing (synced 2026-08-30)
Direct recommendation by task type
Understand these patterns and instantly grasp any model name
| Element | Meaning | Example |
|---|---|---|
| GLM prefix | Zhipu general-purpose LLM family | GLM-5.3, GLM-4.7 |
| Version number (4→5→5.3) | Higher number = newer and stronger | GLM-4 → GLM-5 → GLM-5.3 |
| -Flash | Lightweight and fast, budget-friendly | GLM-5.3-Flash, GLM-4.7-Flash |
| -FlashX | Enhanced high-speed Flash | GLM-4.7-FlashX |
| -Turbo | Speed-optimized version | GLM-5-Turbo, GLM-5V-Turbo |
| -Air | Lightweight, balanced | GLM-4.5-Air |
| -AirX | Ultra-fast, low-latency | GLM-4.5-AirX |
| -Long | Built for ultra-long context | GLM-4-Long |
| -V | Vision understanding model | GLM-5V-Turbo, GLM-4.6V |
| -V-Thinking | Vision reasoning model | GLM-4.1V-Thinking-FlashX |
| GLM-Realtime | Realtime audio-video conversation | GLM-Realtime |
| GLM-ASR | Automatic speech recognition | GLM-ASR-2512 |
| GLM-TTS | Text to speech | GLM-TTS |
| CogView | Image generation | CogView-4 |
| CogVideoX | Video generation | CogVideoX-3 |
| Vidu | Video generation (partner) | Vidu Q1 |
| Embedding | Text embedding | Embedding-3 |
| -YYYYMMDD | Date snapshot | GLM-4-Flash-250414 |
Zhipu's latest flagship generation, recommended for complex reasoning, coding, agents and multimodal tasks
| Model ID | Tier | Price | Modality | Reasoning | Speed | Context | Input | Output | Notes |
|---|---|---|---|---|---|---|---|---|---|
| GLM-5.3NEW | Flagship | Best | Standard | 1M | 1M | 128K | Latest flagship for complex software engineering and long-horizon agent tasks; coding experience 50% better than the previous generation | ||
| GLM-5.3-FlashNEW | Balanced | Deep | Fast | 1M | 1M | 128K | Native multimodal budget model; natively understands images/videos/files; list price ¥0.8/¥2.8, limited-time 50% off (¥0.4/¥1.4) | ||
| GLM-5.2 | Flagship | Best | Standard | 1M | 1M | 128K | Stable execution of complex long-horizon tasks; significantly improved coding | ||
| GLM-5.1 | Flagship | Deep | Standard | 200K | 200K | 128K | Coding on par with Claude Opus 4.6; autonomous long-horizon tasks for up to 8 hours; ¥8/¥28 when input ≥32K | ||
| GLM-5 | Flagship | Deep | Standard | 200K | 200K | 128K | Coding on par with Claude Opus 4.5; excels at agentic long-horizon planning and execution; ¥6/¥22 when input ≥32K | ||
| GLM-5V-Turbo | Flagship | Deep | Fast | 200K | 200K | 128K | Multimodal coding base, balancing vision understanding, reasoning and code generation; ¥7/¥26 when input ≥32K | ||
| GLM-5-Turbo | Balanced | Deep | Fast | 200K | 200K | 128K | Optimized for long-task core needs; good continuity on complex long tasks; ¥7/¥26 when input ≥32K |
Text models focused on natural-language understanding and generation, covering chat, writing, reasoning and code
| Model ID | Tier | Price | Modality | Reasoning | Speed | Context | Input | Output | Notes |
|---|---|---|---|---|---|---|---|---|---|
| GLM-4.7 | Balanced | Standard | Standard | 200K | 200K | 128K | Upgraded general chat, reasoning and agent capabilities; ¥2/¥8 when input <32K and output <0.2K | ||
| GLM-4.7-FlashX | Budget | Basic | Fastest | 200K | 200K | 128K | Lightweight high-speed Flash; fast reasoning; ideal for high-concurrency calls; ¥0.5 in / ¥3 out | ||
| GLM-4.7-Flash | Budget | Basic | Fastest | 200K | 200K | 128K | Free text model carrying over the general capabilities of the GLM-4.7 base | ||
| GLM-4.6 | Balanced | — | Standard | Standard | 200K | 200K | 128K | Excels at advanced coding, complex reasoning and tool calling; pay-as-you-go pricing not public; deployable via private instances | |
| GLM-4.5-Air | Balanced | Standard | Standard | 128K | 128K | 96K | Lightweight model; stable on reasoning, coding and agent tasks | ||
| GLM-4.5-AirX | Budget | Basic | Fast | 128K | 128K | 96K | Ultra-fast version for low-latency, high-response business scenarios | ||
| GLM-4-Long | Balanced | Standard | Standard | 1M | 1M | 4K | Designed for ultra-long text and memory-intensive tasks | ||
| GLM-4-FlashX-250414 | Budget | Basic | Fastest | 128K | 128K | 16K | Enhanced high-speed Flash; ideal for high-concurrency calls | ||
| GLM-4.5-Flash | Budget | Basic | Fastest | 128K | 128K | 96K | Free text model; handles up to 128K context | ||
| GLM-4-Flash-250414 | Budget | Basic | Fastest | 128K | 128K | 16K | Free text model | ||
| GLM-4-Plus | Flagship | Deep | Standard | 128K | 128K | 128K | High-intelligence flagship; fully self-developed 4th-gen base model; supports advanced agent capabilities | ||
| GLM-4-Air-250414 | Balanced | Standard | Standard | 128K | 128K | 128K | GLM-4-Air date snapshot 250414; cost-effective text model | ||
| GLM-4-AirX | Balanced | Standard | Fast | 8K | 8K | 8K | Ultra-fast reasoning version for low-latency, high-response scenarios | ||
| GLM-4-Assistant | Balanced | Standard | Standard | 128K | 128K | 128K | Text model optimized for agent applications | ||
| GLM-Z1-Air | Balanced | Deep | Standard | 128K | 128K | 128K | Cost-effective reasoning model | ||
| GLM-Z1-AirX | Balanced | Deep | Fast | 32K | 32K | 32K | Ultra-fast reasoning model, balancing depth and response speed | ||
| GLM-Z1-FlashX | Budget | Standard | Fastest | 128K | 128K | 16K | High-speed, low-cost reasoning model for high-concurrency reasoning | ||
| GLM-Z1-Flash | Budget | Standard | Fastest | 128K | 128K | 16K | Free reasoning model | ||
| GLM-4-9B | Balanced | Basic | Standard | 128K | 128K | 128K | Open-source 9B text model | ||
| ChatGLM3-6B | Balanced | Basic | Standard | 8K | 8K | 8K | Open-source chat model (3rd gen) |
Image and video understanding, vision reasoning and mobile agent models
| Model ID | Tier | Price | Modality | Reasoning | Speed | Context | Input | Output | Notes |
|---|---|---|---|---|---|---|---|---|---|
| GLM-4.6V | Balanced | Standard | Standard | 128K | 128K | 32K | Native tool calling; stable frontend code replication; ¥1/¥3 when input <32K, ¥2/¥6 for [32,128) | ||
| GLM-4.5V | Balanced | Standard | Standard | 64K | 64K | 32K | Vision understanding model; ¥2/¥6 when input <32K | ||
| GLM-4V-Plus-0111 | Balanced | Standard | Standard | 16K | 16K | — | Vision reasoning model, ¥4/1M tokens | ||
| AutoGLM-Phone | Flagship | — | Standard | Standard | 20K | 20K | 2048 | Mobile AI assistant framework; completes app operations via natural language | |
| GLM-4.1V-Thinking-FlashX | Balanced | Deep | Fast | 64K | 64K | 16K | Excels at complex-scene understanding and multi-step analysis; ideal for high-concurrency vision reasoning | ||
| GLM-OCR | Balanced | ¥0.2 / 1M tokens | Basic | Fast | 32K | 32K | — | Lightweight image-text parsing model balancing high-accuracy, high-efficiency document understanding; input: single image ≤10MB / PDF ≤50MB / 100 pages | |
| GLM-4.6V-FlashX | Budget | Basic | Fastest | 128K | 128K | 32K | High-speed vision reasoning model; ¥0.15/¥1.5 when input <32K, ¥0.3/¥3 for [32,128) | ||
| GLM-4.6V-Flash | Budget | Basic | Fastest | 128K | 128K | 32K | Free model with vision reasoning | ||
| GLM-4.1V-Thinking-Flash | Budget | Standard | Fastest | 64K | 64K | 16K | Free model with vision reasoning | ||
| GLM-4V-Flash | Budget | Fast | Fastest | 16K | 16K | 1K | Free model with image understanding |
Text-to-image and image editing models, billed per request
| Model ID | Tier | Price | Modality | Notes |
|---|---|---|---|---|
| GLM-Image | Flagship | ¥0.1 / request | Flagship image generation model; stronger complex-instruction following and knowledge-dense generation; outstanding text rendering; multi-resolution | |
| CogView-4 | Balanced | ¥0.06 / request | General image generation model; high quality with richer detail; suits diverse creative scenarios; multi-resolution | |
| CogView-3-Flash | Budget | Free | Free model for lightweight image creation; multi-resolution |
Text-to-video, image-to-video and first/last-frame generation models, billed per request
| Model ID | Tier | Price | Modality | Notes |
|---|---|---|---|---|
| CogVideoX-3 | Flagship | ¥1 / request | Flagship video model; stronger instruction following and physics simulation; improved realism and 3D scenes; supports first/last-frame generation | |
| ViduQ1-Text | Balanced | ¥2.5 / request | Vidu Q1 text-to-video, 1080p | |
| ViduQ1-Image | Balanced | ¥2.5 / request | Vidu Q1 image-to-video, 1080p | |
| ViduQ1-Start-End | Balanced | ¥2.5 / request | Vidu Q1 first/last-frame generation, 1080p | |
| Vidu2-Image | Balanced | ¥1.25 / request | Vidu 2 image-to-video, 720p | |
| Vidu2-Start-End | Balanced | ¥1.25 / request | Vidu 2 first/last-frame generation, 720p | |
| Vidu2-Reference | Balanced | ¥2.5 / request | Vidu 2 reference-to-video, 720p | |
| CogVideoX-2 | Balanced | ¥0.5 / request | General video generation model; multi-resolution | |
| CogVideoX-Flash | Budget | Free | Free video generation model; accepts image and text input |
Speech recognition, speech synthesis, realtime voice/video conversation and voice cloning models, with varying billing units
| Model ID | Tier | Price | Modality | Notes |
|---|---|---|---|---|
| GLM-Realtime-Flash | Flagship | Audio ¥0.18/min · Video ¥1.2/min | Realtime audio-video model, low-latency version; accepts video, audio and text multimodal input | |
| GLM-Realtime-Air | Balanced | Audio ¥0.3/min · Video ¥2.1/min | Realtime audio-video model, lightweight balanced version; accepts video, audio and text multimodal input | |
| GLM-4-Voice | Balanced | ¥80 / 1M tokens | Realtime voice chat model; supports text and audio; audio token unit price, not directly comparable to text tiers | |
| GLM-TTS-Clone | Balanced | ¥6 / request | Voice cloning model; 3 seconds of audio is enough to clone a similar voice | |
| GLM-TTS | Balanced | ¥2 / 10K chars | Speech synthesis model; ultra-humanlike voice with emotion; offers streaming and non-streaming APIs | |
| GLM-ASR-2512 | Balanced | Input ¥16/1M tokens (~¥0.0002/sec), free output | High-accuracy speech recognition; low character error rate; custom vocabulary; covers major languages and dialects |
Specialized models for vector retrieval, reranking, roleplay, emotional support and code completion
| Model ID | Tier | Price | Modality | Context | Output | Notes |
|---|---|---|---|---|---|---|
| Embedding-3 | Balanced | 8K | — | 3rd-gen text embedding model (V3) for semantic retrieval, clustering, topic modeling and classification | ||
| Embedding-2 | Balanced | 8K | — | 2nd-gen text embedding model (V2) | ||
| Rerank | Balanced | 4K | — | Text reranking model, ¥0.8/1M input tokens | ||
| CodeGeeX-4 | Balanced | 128K | 32K | Code completion model | ||
| CharGLM-4 | Balanced | 8K | 4K | Humanlike chat model | ||
| Emohaa | Balanced | 8K | 4K | Emotional support model |
Older models no longer recommended or under maintenance; for compatibility reference only
| Model ID | Tier | Price | Modality | Reasoning | Speed | Context | Input | Output | Notes |
|---|---|---|---|---|---|---|---|---|---|
| GLM-4-0520 | Balanced | Standard | Standard | 128K | 128K | 128K | GLM-4 0520 legacy version | ||
| GLM-4V-Plus | Balanced | Standard | Standard | 8K | 8K | 8K | GLM-4V Plus legacy version | ||
| GLM-4V | Balanced | Standard | Standard | 8K | 8K | 8K | GLM-4V legacy version | ||
| GLM-4 | Balanced | Standard | Standard | 128K | 128K | 128K | GLM-4 base legacy version | ||
| GLM-4-Air | Balanced | Standard | Standard | 128K | 128K | 128K | GLM-4-Air legacy version | ||
| GLM-4-Flash | Budget | Basic | Fastest | 128K | 128K | 16K | GLM-4-Flash legacy version, free | ||
| GLM-4.5 | Balanced | — | Standard | Standard | 128K | 128K | 128K | Pay-as-you-go API retired; available via fine-tuning and private deployment | |
| CogView-3 | Balanced | Basic | Standard | — | — | — | Pay-as-you-go API retired; private deployment only; historical price ¥0.06/request |
Find the primary and alternative model for each typical scenario
| Scenario | Recommended | Alternatives | Key capability |
|---|---|---|---|
| Complex coding / software engineering | GLM-5.3 | GLM-5.2 | Reasoning Long context Agent |
| Deep reasoning / math derivation | GLM-5.3 | GLM-5.1 | Best reasoning 128K output |
| Long-document processing | GLM-4-Long | GLM-5.3 | 1M context Low cost |
| Image / video understanding | GLM-5.3-Flash | GLM-4.6V | Native multimodal Video understanding |
| Image generation | GLM-Image | CogView-4 | Text rendering Multi-resolution |
| Video generation | CogVideoX-3 | ViduQ1-Text | First/last frame Physics simulation |
| Realtime voice chat | GLM-Realtime-Flash | GLM-4-Voice | Realtime Audio-video |
| Speech recognition | GLM-ASR-2512 | — | High accuracy Dialects |
| Semantic search / RAG | Embedding-3 | Embedding-2 | 8K context Semantic retrieval |
| Cheapest | GLM-4.7-Flash | GLM-4-Flash-250414 | Free 200K context |