All Qwen models analyzed · hybrid thinking · omni-modal · visual generation · full speech pipeline · open-source deployment — Sources: Alibaba Cloud Bailian official docs Model catalog · Model pricing · Classification docs (synced 2026-08-30)
Direct recommendation by task type
Master these patterns and instantly know any model name
| Suffix | Meaning | Example |
|---|---|---|
| qwen3.X version number | Higher number = newer and stronger | qwen3 → qwen3.5 → qwen3.8 |
| -max | Flagship, strongest capability | qwen3.8-max |
| -prime | Prime mode, faster, stronger and pricier | qwen3.8-max-prime |
| -plus | Balanced enhanced | qwen3.7-plus |
| -flash | Lightweight and fast, cheapest | qwen3.7-flash |
| -turbo | Speed-optimized (legacy) | qwen-turbo |
| qwq- | Deep thinking only (thinking mode) | qwq-plus |
| -coder | Coding | qwen3-coder-plus |
| -long | Ultra-long text only | qwen-long |
| -vl | Vision-language model | qwen3-vl-plus |
| -omni | Omni-modal (text/image/audio/video) | qwen3.5-omni-plus |
| -ocr | Document text extraction | qwen3.5-ocr |
| -asr / -tts | Speech recognition / synthesis | qwen-audio-3.0-asr-flash |
| -realtime | WebSocket realtime | qwen3.5-omni-plus-realtime |
| -filetrans / -streaming | File transcription / streaming realtime | qwen-audio-3.0-asr-flash-filetrans |
| -livetranslate | Realtime speech translation | qwen3.5-livetranslate-flash |
| -instruct / -vc / -vd | Instruct / voice clone / voice design | qwen3-tts-instruct-flash |
| -embedding / -rerank | Embedding / reranking | text-embedding-v4, qwen3-rerank |
| wan2.X- | Wanxiang visual generation (image/video) | wan2.7-image, wan2.6-t2v |
| -t2i/-t2v/-i2v/-r2v/-s2v | T2I/T2V/I2V/reference-to-video/audio-to-video | wan2.6-t2i, wan2.6-r2v |
| -latest / -YYYYMMDD | Latest snapshot / dated snapshot | qwen3.7-plus-2026-05-26 |
| cosyvoice- / fun-asr- | Speech synthesis / recognition series | cosyvoice-v3.5-plus, fun-asr |
| z-image- / happyhorse- | Fast image / third-party video generation | z-image-turbo, happyhorse-1.1-t2v |
Qwen's core text models; the whole family supports hybrid thinking mode (the enable_thinking switch) and 1M context; vision input is covered in the "Vision Understanding" section
| Model ID | Tier | Price | Modality | Thinking | Speed | Context | Max output | Notes |
|---|---|---|---|---|---|---|---|---|
| qwen3.8-max-primeNEW | Flagship | Best | Slow | 1M | 64k | Prime mode, currently the strongest, ¥24/¥72 | ||
| qwen3.8-maxNEW | Flagship | Best | Slow | 1M | 64k | Latest flagship, hybrid thinking on by default, ¥12/¥36 | ||
| qwen3.7-max | Flagship | Best | Slow | 1M | 64k | Previous-gen flagship, 50% off for a limited time | ||
| qwen3.7-plusRec | Balanced | Deep | Standard | 1M | 64k | Flagship balanced choice, ¥2/¥8 (20% off, limited time), built-in tools + web | ||
| qwen3.8-27bNEW | Balanced | Deep | Standard | 1M | 64k | Dense 27B (1M context), ¥3/¥12 | ||
| qwen3.6-plus | Balanced | Deep | Standard | 1M | 64k | Previous-gen balanced | ||
| qwen3.5-flash | Budget | Basic | Fast | 1M | 64k | Previous-gen lightweight, from ¥0.2 | ||
| qwen3.5-plus | Balanced | Standard | Standard | 1M | 64k | Mature and stable | ||
| qwen3.8-flashNEW | Budget | Standard | Fast | 1M | 64k | Latest lightweight, hybrid thinking on by default, ¥0.8/¥2.7 | ||
| qwen3.7-flashRec | Budget | Basic | Fast | 1M | 64k | Cheapest, ¥0.2/¥0.8 | ||
| qwen3.6-flash | Budget | Standard | Fast | 1M | 64k | Previous-gen lightweight | ||
| qwen3-coder-plus | Flagship | Deep | Standard | 1M | 64k | Coding-specialized, agentic coding optimized, ¥4/¥16 | ||
| qwen-long | Budget | Basic | Standard | 10M | 8k | Ultra-long text (tens of millions of chars), ¥0.5/¥2 | ||
| qwen-mt-plus | Balanced | Basic | Fast | — | — | Machine translation, multilingual | ||
| qwen3-max | Balanced | Deep | Slow | 128K | 32k | Qwen3-generation flagship (legacy), from ¥2.5 |
Thinking-only mode (always reasons before answering; cannot be disabled); for a controllable switch, use the hybrid thinking mode of text generation models
| Model ID | Tier | Price | Modality | Thinking | Speed | Context | Max output | Notes |
|---|---|---|---|---|---|---|---|---|
| qwq-plusRec | Balanced | Best | Slow | 128K | 32k | Thinking-only mode, for logic/math reasoning, ¥1.6/¥4 |
Understands text, audio, images and video at once; outputs text and speech; for realtime voice chat see the "Realtime voice chat" row
| Model ID | Tier | Price | Input | Output | API | Thinking | Features | Notes |
|---|---|---|---|---|---|---|---|---|
| qwen3.5-omni-plus | Flagship | HTTP | Fast | FC+web | Omni-modal flagship, 3h audio / 1h video | |||
| qwen3.5-omni-plus-realtime | Flagship | WebSocket | Fast | FC+web | Realtime audio/video chat | |||
| qwen3.5-omni-flash | Balanced | HTTP | Fast | FC+web | Lightweight omni-modal, from ¥2.2 | |||
| qwen3.5-omni-flash-realtime | Balanced | WebSocket | Fast | FC+web | Realtime | |||
| qwen3-omni-flash | Budget | HTTP | Standard | FC | Supports thinking mode; single input limited to 150s | |||
| qwen3-omni-flash-realtime | Budget | WebSocket | Fast | — | Realtime; no FC/web/thinking | |||
| qwen-audio-3.0-realtime-plusNEW | Flagship | WebSocket | Fast | Semantic VAD+FC | S2S voice chat flagship; meaningless backchannels won't interrupt | |||
| qwen-audio-3.0-realtime-flashNEW | Budget | WebSocket | Fast | Semantic VAD+FC | Cost-sensitive voice chat | |||
| qwen-omni-turbo | Budget | HTTP | Fast | — | Legacy omni-modal; migrate to qwen3.5-omni |
~3s-latency simultaneous interpretation, works out of the box; file mode translates audio/video files
| Model ID | Tier | Price | Input | API | Languages | Notes |
|---|---|---|---|---|---|---|
| qwen3.5-livetranslate-flash-realtimeNEWRec | Balanced | Audio ¥40/1M | WebSocket | 60 | Top choice for realtime interpreting; 29 speech + 31 text outputs | |
| qwen3.5-livetranslate-flash | Balanced | Audio ¥40/1M | HTTP | 60 | Audio/video file translation | |
| qwen3-livetranslate-flash-realtime | Budget | Audio ¥64/1M | WebSocket | 18 | Previous-gen realtime, incl. 5 Chinese dialects | |
| qwen3-livetranslate-flash | Budget | Audio ¥10/1M | HTTP | 18 | File mode, video-context aware |
Image analysis, video understanding, OCR document extraction. Note: all qwen3.8/3.7/3.6/3.5 models in the Text Generation section also support vision input (1M context, up to 2h video); the VL series adds vision-specialized enhancements
| Model ID | Tier | Price | Modality | Thinking | Speed | Context | Max output | Features | Notes |
|---|---|---|---|---|---|---|---|---|---|
| qwen3-vl-plusRec | Flagship | Deep | Standard | 256K | 64k | 1h video | Vision flagship, from ¥1/¥10 | ||
| qwen3-vl-flash | Budget | Standard | Fast | 256K | 64k | 1h video | Lightweight vision | ||
| qwen3.5-ocr | Balanced | Basic | Standard | — | — | Docs/tables/handwriting | OCR-specialized, optimized text extraction accuracy | ||
| qwen-vl-ocr | Budget | Fast | Standard | — | — | — | Legacy OCR; migrate to qwen3.5-ocr | ||
| qvq-max | Balanced | Deep | Slow | — | — | Visual reasoning | Legacy visual reasoning; capability merged into Qwen3-VL | ||
| qvq-plus | Budget | Standard | Slow | — | — | Visual reasoning | Legacy visual reasoning | ||
| qwen-vl-max | Balanced | Basic | Standard | — | — | — | Legacy VL | ||
| qwen-vl-plus | Budget | Fast | Fast | — | — | — | Legacy VL |
Text-to-image and image editing; the qwen-image-3.0 family supports agent prompt rewriting and Chinese text rendering
| Model ID | Tier | Price | Text-to-image | Editing | Max output | Max resolution | Notes |
|---|---|---|---|---|---|---|---|
| qwen-image-3.0-proNEWRec | Flagship | Yes | Yes | 6 | 2048x2048 | Image 3.0 flagship, agent prompt rewriting, small-text/multilingual rendering | |
| qwen-image-3.0NEW | Balanced | Yes | Yes | 6 | 2048x2048 | Same as above, faster generation | |
| wan2.7-image-proNEW | Flagship | Yes | Yes | 4 (12 consecutive) | 4096x4096 | Brand color palette, multi-image reference (9 images), character consistency | |
| wan2.7-imageNEW | Balanced | Yes | Yes | 4 (12 consecutive) | 2048x2048 | Same as above, faster generation | |
| z-image-turbo | Budget | Yes | No | 1 | 2048x2048 | 10x faster, ~1/5 the price, photorealistic portraits/product shots | |
| qwen-image-2.0-pro | Balanced | Yes | Yes | 6 | 2048x2048 | Previous-gen pro | |
| qwen-image-2.0 | Balanced | Yes | Yes | 6 | 2048x2048 | Previous-gen standard | |
| qwen-image-max | Balanced | Yes | No | 1 | 1664x928 | Legacy flagship | |
| qwen-image-plus | Budget | Yes | No | 1 | 1664x928 | Legacy balanced | |
| qwen-image-edit-max | Balanced | No | Yes | 6 | 2048x2048 | Image editing only | |
| qwen-image-edit-plus | Budget | No | Yes | 6 | 2048x2048 | Lightweight image editing | |
| wan2.6-t2i | Balanced | Yes | Yes | 4 | 1440x1440 | Wanxiang text-to-image | |
| wan2.6-image | Balanced | Yes | Yes | 4 | 1440x1440 | Wanxiang image generation & editing |
Billed by output video seconds (input free); some models tiered by resolution
| Model ID | Tier | Type | Billing | Resolution | Notes |
|---|---|---|---|---|---|
| wan3.0-video-primeNEWRec | Flagship | Text-to-video | ¥0.45-1.8/sec | 1080P | Latest Wanxiang video flagship |
| wan3.0-videoNEW | Flagship | Text-to-video | ¥0.3-1.2/sec | 1080P | 30% off for a limited time |
| wan2.7-t2vNEW | Flagship | Text-to-video | ¥0.6-1/sec | 1080P | |
| wan2.7-i2vNEW | Flagship | Image-to-video (with audio) | ¥0.6-1/sec | 1080P | |
| wan2.7-r2vNEW | Flagship | Reference-to-video (with audio) | ¥0.6-1/sec | 1080P | |
| happyhorse-1.1-t2vNEWRec | Flagship | Text-to-video | ¥0.45-1.2/sec | 1080P | 60% off for a limited time, latest third-party video generation |
| happyhorse-1.1-i2vNEW | Flagship | Image-to-video | ¥0.45-1.2/sec | 1080P | Image-driven video |
| happyhorse-1.1-r2vNEW | Flagship | Reference-to-video | ¥0.45-1.2/sec | 1080P | Reference image/character driven |
| happyhorse-1.0-video-editNEW | Flagship | Video editing | ¥0.9-1.6/sec | 1080P | 20% off for a limited time; input + output both billed |
| wan2.6-t2v | Flagship | Text-to-video | Per-second billing | 1080P | Wanxiang text-to-video |
| wan2.6-i2v | Flagship | Image-to-video | Per-second billing | 1080P | Wanxiang image-to-video |
| wan2.6-r2v | Flagship | Reference-to-video | Per-second billing | 1080P | Wanxiang reference-to-video |
| wan2.2-t2v-plus | Balanced | Text-to-video | Per-second billing | 1080P | Previous-gen pro |
| wan2.2-i2v-plus | Balanced | Image-to-video | Per-second billing | 1080P | Previous-gen pro |
| wan2.2-kf2v-flash | Balanced | First/last-frame video | Per-second billing | 720P | Generated by interpolating first + last frames |
| wan2.2-s2v | Balanced | Audio-to-video | Per-second billing | 720P | Audio-driven digital human/animation |
| wan2.2-animate-move | Balanced | Motion transfer | Per-second billing | 720P | Transfers video motion to an image character |
| wan2.2-animate-mix | Balanced | Character swap | Per-second billing | 720P | Character replacement/swap in video |
| wanx2.1-t2v-turbo | Budget | Text-to-video | Per-second billing | 720P | Economy |
| wanx2.1-i2v-turbo | Budget | Image-to-video | Per-second billing | 720P | Economy |
| wanx2.1-kf2v-plus | Budget | First/last-frame video | Per-second billing | 720P | Economy |
| wanx2.1-vace-plus | Budget | Video editing | Per-second billing | 720P | Wanxiang video editing |
Two paths: realtime (WebSocket streaming) and non-realtime (HTTP file transcription); the Qwen-Audio-3.0 family supports hotwords and Prompt context injection
| Model ID | Tier | Mode | API | Accuracy boost | Speaker diarization | Duration/Size | Notes |
|---|---|---|---|---|---|---|---|
| qwen-audio-3.0-asr-flash-streamingNEWRec | Flagship | Realtime | WebSocket | Hotwords+Prompt | — | Unlimited | Realtime recognition flagship, multilingual incl. dialects |
| qwen-audio-3.0-asr-flash-filetransNEWRec | Flagship | Non-realtime | HTTP | Hotwords+Prompt | Yes | 12h/2GB | File transcription flagship, speaker diarization |
| qwen-audio-3.0-asr-flashNEW | Balanced | Non-realtime | HTTP | Hotwords+Prompt | — | 5 min/2GB | Fast short-audio transcription |
| fun-asr-realtime | Balanced | Realtime | WebSocket | Hotwords | — | Unlimited | FunASR realtime |
| fun-asr-mtl-realtime | Balanced | Realtime | WebSocket | Hotwords | — | Unlimited | FunASR multilingual realtime |
| fun-asr | Balanced | Non-realtime | HTTP | Hotwords | Yes | 12h/2GB | FunASR file version, speaker diarization |
| fun-asr-mtl | Balanced | Non-realtime | HTTP | Hotwords | Yes | 12h/2GB | FunASR multilingual |
| qwen3-asr-flash-realtime | Balanced | Realtime | WebSocket | — | — | Unlimited | Supports emotion recognition |
| qwen3-asr-flash-filetrans | Balanced | Non-realtime | HTTP | — | — | 12h/2GB | Supports emotion recognition |
| qwen3-asr-flash | Budget | Non-realtime | HTTP (OpenAI-compatible) | — | — | 5 min/10MB | Supports emotion recognition |
| paraformer-realtime-v2 | Budget | Realtime | WebSocket | — | — | — | Legacy; migrate to Fun-ASR / Qwen-ASR |
| paraformer-v2 | Budget | Non-realtime | HTTP | — | Yes | — | Legacy |
| sensevoice-v1Deprecated | Budget | Non-realtime | HTTP | — | — | — | Retiring soon; migrate ASAP |
| gummy-realtime-v1Deprecated | Budget | Realtime | WebSocket | — | — | — | Retiring soon; migrate ASAP |
| gummy-chat-v1Deprecated | Budget | Realtime | WebSocket | — | — | 1 min | Retiring soon; migrate ASAP |
Standard synthesis / voice cloning (audio sample cloning) / voice design (created from text description); instruct mode tunes tone, emotion and speaking rate in natural language
| Model ID | Tier | Price | Series | API | Voice clone | Voice design | Instruct | Notes |
|---|---|---|---|---|---|---|---|---|
| qwen-audio-3.0-tts-plusNEWRec | Flagship | Qwen-Audio-TTS | WS / HTTP | Yes | Yes | Yes | Latest TTS flagship, full capability | |
| qwen-audio-3.0-tts-flashNEW | Balanced | Qwen-Audio-TTS | WS / HTTP | Yes | Yes | Yes | Lightweight | |
| cosyvoice-v3.5-plus | Flagship | CosyVoice | WS / HTTP | Yes | Yes | Yes | Supports SSML and LaTeX formula reading | |
| cosyvoice-v3.5-flash | Balanced | CosyVoice | WS / HTTP | Yes | Yes | Yes | Lightweight | |
| cosyvoice-v3-flash | Budget | CosyVoice | WS / HTTP | Yes | Yes | Yes | Lightweight, ¥1/10K chars | |
| cosyvoice-v3-plus | Balanced | CosyVoice | WS / HTTP | Yes | Yes | — | Previous-gen pro | |
| MiniMax/speech-2.8-hd | Balanced | MiniMax | HTTP | Yes | — | — | Third-party high-fidelity voices | |
| qwen3-tts-flash | Budget | Qwen3-TTS | HTTP | — | — | — | Economy standard synthesis | |
| qwen3-tts-instruct-flash | Budget | Qwen3-TTS | HTTP | — | — | Yes | Instructs tone/emotion/style | |
| qwen3-tts-vc | Balanced | Qwen3-TTS-VC | WS / HTTP | Yes | — | — | Voice cloning only | |
| qwen3-tts-vd | Balanced | Qwen3-TTS-VD | WS / HTTP | — | Yes | — | Voice design only | |
| qwen-tts | Budget | Qwen-TTS (Legacy) | HTTP / WS | — | — | — | Legacy token-billed version; migrate to Qwen3-TTS | |
| qwen-voice-enrollment | Balanced | ¥0.01/voice | Voice management | — | — | — | — | Voice registration and management (prerequisite for voice cloning) |
| qwen-voice-design | Balanced | ¥0.2/voice | Voice design service | — | — | — | — | Creates new voices from text descriptions, billed per voice |
Semantic search, RAG retrieval and accuracy; rerank is recommended after Embedding retrieval to re-rank Top-N results
| Model ID | Tier | Price | Type | Dimensions | Max tokens | Notes |
|---|---|---|---|---|---|---|
| text-embedding-v4Rec | Flagship | Text embedding | 64~2048 | 8,192 | Latest text embedding, v3-compatible dimensions, MTEB leader | |
| qwen3.7-text-embeddingNEW | Flagship | Text embedding | 64~2048 | 8,192 | New-gen text embedding (Qwen3.7), ¥0.5 | |
| text-embedding-v3 | Budget | Text embedding | 512~1024 | 8,192 | Migrate existing v3 indexes | |
| qwen3-vl-embedding | Balanced | Multimodal embedding | 256~2560 | 32,000 | Image-text hybrid retrieval (fused + independent vectors) | |
| qwen2.5-vl-embedding | Balanced | Multimodal embedding | 512~2048 | 32,000 | Image-text hybrid retrieval (fused vectors only) | |
| tongyi-embedding-vision-plus | Balanced | Multimodal embedding | 64~1152 | 1,024 | Cross-modal search (independent vectors) | |
| tongyi-embedding-vision-flash | Budget | Multimodal embedding | 64~768 | 1,024 | Low-cost cross-modal | |
| qwen3-rerankRec | Flagship | Rerank | — | 4,000/passage | 100+ languages, up to 500 docs | |
| qwen3-vl-rerank | Balanced | Multimodal rerank | — | 8,000/passage | Text/image/video hybrid ranking | |
| gte-rerank-v2 | Budget | Rerank | — | 4,000/passage | Text semantic retrieval |
Vertical-use models for translation, document parsing, deep research and text analysis
| Model ID | Tier | Price | Use | Notes |
|---|---|---|---|---|
| qwen-mt-plus | Balanced | Machine translation | Multilingual translation, with terminology injection | |
| qwen-doc-turbo | Balanced | Document parsing | Document content extraction and structuring | |
| qwen-deep-research | Flagship | Deep research | Web search + multi-step reasoning + long reports, ¥54/¥163 | |
| tongyi-xiaomi-analysis-pro | Balanced | Conversation analysis | Tongyi Xiaomi Pro, ¥1.0/¥2.7 | |
| tongyi-xiaomi-analysis-flash | Budget | Conversation analysis | Tongyi Xiaomi Lite, ¥0.2/¥0.4 | |
| qwen-math-plus | Flagship | Math reasoning | Math/contest problem solving, ¥4/¥12 | |
| qwen-math-turbo | Balanced | Math reasoning | Lightweight math reasoning, ¥2/¥6 | |
| qwen-mt-flash | Budget | Machine translation | Lightweight, ¥0.7/¥1.95 | |
| qwen-mt-lite | Budget | Machine translation | Economy, ¥0.6/¥1.6 | |
| qwen-mt-turbo | Budget | Machine translation | Speed-optimized, ¥0.7/¥1.95 | |
| qwen-mt-image-2.0 | Balanced | Image translation | Translates text in images into multiple languages, ¥0.004/image | |
| fun-music-v1 | Budget | ¥0.002/sec | Music generation | Generates music from lyrics/prompts, billed per second |
Open weights, self-deployable, no API call limits, only hardware costs; Bailian also offers hosted calling of the open-source versions
| Model ID | Tier | Price | Thinking | Context | Notes |
|---|---|---|---|---|---|
| qwen3.8-2.4t-a95bNEWOSS | Flagship | Self-hosted | Best | — | Latest open-source flagship, hybrid thinking on by default |
| qwen3.6-35b-a3bOSS | Balanced | Self-hosted | Deep | 256K | Open-source balanced (MoE 35B-A3B) |
| qwen3.6-27bOSS | Balanced | Deep | — | Open-source dense 27B (Qwen3.6-gen), ¥3/¥18 | |
| qwen3.5-397b-a17bOSS | Flagship | Self-hosted | Deep | 32K | Open-source flagship (MoE 397B-A17B) |
| qwen3.5-122b-a10bOSS | Balanced | Self-hosted | Standard | 32K | Open-source mid-size (MoE 122B-A10B) |
| qwen3.5-35b-a3bOSS | Budget | Self-hosted | Standard | 32K | Open-source small (MoE 35B-A3B) |
| qwen3.5-27bOSS | Budget | Self-hosted | Standard | 32K | Open-source dense 27B |
| qwen3-next-80b-a3b-instructOSS | Flagship | Fast | 256K | Qwen3-Next new-gen open-source MoE 80B-A3B, ¥1/¥4 | |
| qwen3-next-80b-a3b-thinkingOSS | Flagship | Deep | 256K | Qwen3-Next thinking edition (thinking-only), ¥1/¥10 | |
| qwen3-coder-nextOSS | Flagship | Fast | 256K | New open-source coding model, ¥1/¥4 | |
| qwen3-235b-a22b-instruct-2507OSS | Flagship | Fast | 256K | Qwen3 open-source flagship MoE 235B-A22B | |
| qwen3-235b-a22b-thinking-2507OSS | Flagship | Deep | 256K | Qwen3 open-source thinking flagship, ¥2/¥20 | |
| qwen3-vl-235b-a22b-instructOSS | Flagship | Fast | 256K | Open-source vision flagship MoE 235B-A22B | |
| qwen3-vl-235b-a22b-thinkingOSS | Flagship | Deep | 256K | Open-source vision thinking flagship | |
| qwen3-omni-30b-a3b-captionerOSS | Balanced | Fast | — | Open-source omni-modal video captioning/description |
Legacy/retiring info for the models covered on this page. Migrate as recommended and don't use them in new projects. Source: Alibaba Cloud Bailian official docs (synced 2026-08-30)
| Model ID | Lifecycle | Replacement | Migration advice |
|---|---|---|---|
| sensevoice-v1Deprecated | Deprecated | →fun-asr | Retiring soon; migrate to Fun-ASR / Qwen-ASR |
| gummy-realtime-v1Deprecated | Deprecated | →fun-asr-realtime | Retiring soon; migrate ASAP |
| gummy-chat-v1Deprecated | Deprecated | →fun-asr-realtime | Retiring soon; migrate ASAP |
| qwen-omni-turbo | Legacy | →qwen3.5-omni-plus | Legacy omni-modal; capabilities fully upgraded |
| qwen-tts | Legacy | →qwen3-tts-flash / qwen-audio-3.0-tts-plus | Legacy token-billed TTS |
| paraformer-realtime-v2 | Legacy | →fun-asr-realtime / qwen3-asr-flash-realtime | Legacy ASR, migration advised |
| paraformer-v2 | Legacy | →fun-asr / qwen3-asr-flash-filetrans | Legacy ASR |
| qwen-vl-ocr | Legacy | →qwen3.5-ocr | Legacy OCR |
| qvq-max | Legacy | →qwen3-vl-plus | Visual reasoning merged into Qwen3-VL |
| qvq-plus | Legacy | →qwen3-vl-flash | Legacy visual reasoning |
| qwen-vl-max | Legacy | →qwen3-vl-plus | Legacy VL |
| qwen-vl-plus | Legacy | →qwen3-vl-flash | Legacy VL |
Quickly match the best model to your needs
| Scenario | Recommended | Alternatives | Key capability |
|---|---|---|---|
| Chat / Writing | qwen3.7-plus | qwen3.6-plus | Deep + 1M context |
| Deep thinking / reasoning | qwen3.8-max | qwq-plus | Best Hybrid thinking |
| Top flagship | qwen3.8-max-prime | qwen3.8-max | ¥24/¥72 Prime mode |
| Ultra-low cost | qwen3.7-flash | qwen3.8-flash | ¥0.2/¥0.8 |
| Agentic coding workflows | qwen3-coder-plus | qwen3.7-plus | + 1M context |
| Long-document processingTens of millions of chars | qwen-long | qwen3.7-plus | 10M context |
| Image / video understanding | qwen3-vl-plus | qwen3.7-plus | 2h video · 1M context |
| Omni-modal understanding (image+audio+video) | qwen3.5-omni-plus | qwen3-omni-flash | |
| Realtime voice chat | qwen-audio-3.0-realtime-plus | qwen3.5-omni-plus-realtime | Semantic VAD + WebSocket |
| Speech transcription / dictation | qwen-audio-3.0-asr-flash-filetrans | fun-asr | 12h + speaker diarization |
| Speech synthesis | qwen-audio-3.0-tts-plus | cosyvoice-v3.5-plus | Clone / Design / Instruct |
| Realtime speech translation | qwen3.5-livetranslate-flash-realtime | qwen3-livetranslate-flash-realtime | 60 languages · 3s latency |
| High-quality image generation | qwen-image-3.0-pro | wan2.7-image-pro | 2048x2048 / 4096x4096 |
| Video generation | happyhorse-1.1-t2v | wan2.6-t2v | 1080P · per-second billing |
| Semantic search / RAG | text-embedding-v4 | qwen3-vl-embedding | 2048-dim + qwen3-rerank |
| Self-hosted deployment | qwen3.8-2.4t-a95b | qwen3.6-35b-a3b | Open weights |