Rank #4 / 9
Qwen3.8-Max
Qwen3.8-Max is Alibaba's Qwen model (proprietary). It is ranked 4 of 9 models on the Visual Intelligence Index.
- Alibaba
- Proprietary
- Qwen
- Video input
- Structured output
- Tool calling
- Streaming
- Batch API
Comparison Summary
Qwen3.8-Max ranks 4 of 9 on the perception index. Immediately ahead is Claude Fable 5.1. Immediately behind is Kimi K3. List price is $2.00 per million input tokens and $6.00 per million output. 1 of 5 priced models cost less to prompt. The context window is 1M. Every score is from our own evaluation runs; list prices are as published by the provider. Strongest capability in this set is Causal Reasoning (78%). Weakest is Action / Event Understanding (72%).
- Visual Intelligence
- 73.8%
- Availability
- Proprietary
- Input / 1M tokens
- $2.00
- Output / 1M tokens
- $6.00
- Context Window
- 1M
- Parameters
- 2T
All benchmark scores
| Benchmark | Metric | Score | Coverage | Setting |
|---|---|---|---|---|
| Video-MME v2 | Accuracy, no subtitles | 60.2% | 98.5% (3200) | 1 fps native video, ~256k visual-token ceiling (168 videos thinned) |
| LVBench | Test accuracy | 58.9% | 70.8% (1549) | 1 fps native video, ~256k visual-token ceiling (all videos thinned to ~860 frames) |
| Perception Test | Overall accuracy | 83.9% | Full | 1 fps native video |
| NExT-QA | Accuracy (hard split) | 87.6% | 98.6% (4996) | 1 fps native video |
| Q-Bench Video | Overall accuracy | 69.0% | 99.8% (868) | 1 fps native video |
| EgoSchema | Accuracy (fullset) | 85.2% | 99.4% (500) | 1 fps native video |
| UCF101-AD | Accuracy | 56.5% | 100.0% (4224) | 1 fps native video |
Benchmark Scores
Compare reported model scores across each available benchmark or capability index.
9 of 9 models
Display
Capabilities Index Scores
Compare reported model scores across each available benchmark or capability index.
9 of 9 models
Display
Usability
3.5/ 5
What it's like to build against this model, scored out of five from the developer-facing capabilities in the dataset.
- Reachable1.0
There is an endpoint you can call without hosting anything.
- First-party API — supported
- Third-party API — supported
- Portable0.0
You can run it yourself, and are not tied to one vendor.
- Open weights — not supported
- Self-hostable — not supported
- Multimodal input0.5
It takes the footage directly, rather than frames you extracted.
- Video ingestion — supported
- Audio — not supported
- Programmable1.0
Output you can parse, and tools it can call on its own.
- Structured output — supported
- Tool calling — supported
- Operable1.0
Usable interactively and in bulk, not only one call at a time.
- Streaming — supported
- Batch API — supported
Other Alibaba models
Other models from the Alibaba family.
| Date | Model | Visual Intelligence |
|---|---|---|
| Jan 28, 2026 | Qwen3.5-35B-A3B | 65.7% |
| Jul 22, 2026 | Qwen3.8-27B | 68.5% |
| Aug 3, 2026 | Qwen3.8-Max | 73.8% |