Skip to content
Vision Models

Vision Models

Comparison of vision-language models on the Visual Intelligence Index, a weighted average across 7 benchmarks we track, alongside list price per million tokens, the long-context rate where a provider charges one, context window, and provider.

RankModelVisual IntelligencePrice in/out $/1MLong context in/out $/1MContext windowProviderOpen model page
01GPT-6 Astra82.2%$10.00 / $50.00$20.00 / $75.001.05MOpenAI
02Gemini 3.8 Flash80.6%$0.75 / $3.75—1MGoogle
03Claude Fable 5.176.2%$10.00 / $50.00—1MAnthropic
04Qwen3.8-Max73.8%$2.00 / $6.00—1MAlibaba
05Kimi K371.9%$3.00 / $15.00—1.05MMoonshot
06Qwen3.8-27B68.5%——256KAlibaba
07Qwen3.5-35B-A3B65.7%——256KAlibaba
08Nemotron 3 Nano Omni 30B-A3B60.3%——262KNVIDIA
09Molmo 2 8B57.6%——37KAi2

Why these models

Leading models with published evidence of video-understanding performance in their last few releases. We are actively adding more models in. If you would like to see a particular model included, please contact us at visionindex@ondeckai.com.