
Robust evaluation of real-world vision tasks.
Vision Index measures frontier vision models against benchmarks and evaluations that match the complexity of real-world video tasks.
2026/09/10Introducing Vision Index
The Visual Intelligence Index
The Visual Intelligence Index combines different benchmarks to evaluate vision model capabilities in understanding video.
Visual Intelligence Index
Visual Intelligence Index v0 is a weighted average of 6 capability indexes. Updated September 22, 2026.
9 of 9 models
Display
Capabilities Index Scores
Each index is a cross-benchmark weighted score of benchmark questions classified by capability.
9 of 9 models
Display
| Model | Video-MME v2 | LVBench | Perception Test | NExT-QA | Q-Bench Video | EgoSchema | UCF101-AD | Average |
|---|---|---|---|---|---|---|---|---|
| GPT-6 Astra | 73.4% | 81.6% | 92.5% | 90.7% | 68.3% | 83.4% | 61.4% | 78.8% |
| Gemini 3.8 Flash | 71.3% | 86.4% | 85.5% | 88.0% | 69.3% | 80.8% | 63.5% | 77.8% |
| Claude Fable 5.1 | 66.2% (missing coverage) | 75.0% | 77.3% | 86.8% | 71.0% | 80.4% (missing coverage) | 60.6% | 73.9% |
| Qwen3.8-Max | 60.2% (missing coverage) | 58.9% (missing coverage) | 83.9% | 87.6% (missing coverage) | 69.0% (missing coverage) | 85.2% (missing coverage) | 56.5% (missing coverage) | 71.6% |
| Kimi K3 | 55.2% | 66.8% | 82.4% | 86.3% | 67.5% | 78.6% | 39.4% | 68.0% |
| Qwen3.8-27B | 50.5% | 64.6% | 72.3% | 82.2% | 67.9% | 79.1% | 45.5% | 66.0% |
| Qwen3.5-35B-A3B | 42.5% | 64.4% | 66.7% | 83.4% | 66.9% | 77.2% | 41.3% | 63.2% |
| Nemotron 3 Nano Omni 30B-A3B | 35.0% | 54.2% | 70.3% | 81.2% | 64.9% | 65.8% | 31.0% | 57.5% |
| Molmo 2 8B | 27.5% | 48.6% | 67.4% | 89.2% | 62.9% | 58.2% | 19.5% | 53.3% |
Benchmark scores report multiple choice accuracy. * score has missing coverage
Vision Benchmarks
Video-MME v2
Full-spectrum video
A benchmark that scores video understanding through grouped questions, factoring in answer consistency and reasoning coherence.
LVBench
Long-context reasoning
A benchmark tailored for comprehensive long video understanding on hour-plus videos across sports, documentaries, events, TV shows, and more.
Perception Test
Spatial perception
A benchmark by DeepMind for low- and mid-level visual perception, covering memory, physics, and semantics across video, audio, and text.



