
Vision Models
Comparison of vision-language models on the Visual Intelligence Index, a weighted average across 7 benchmarks we track, alongside list price per million tokens, the long-context rate where a provider charges one, context window, and provider.
Why these models
Leading models with published evidence of video-understanding performance in their last few releases. We are actively adding more models in. If you would like to see a particular model included, please contact us at visionindex@ondeckai.com.