Skip to content
Vision Benchmarks

Vision Benchmarks

Handpicked open benchmarks for video understanding. These benchmarks are used to evaluate the performance of video understanding models on a variety of capabilities, including long context, temporal reasoning, perception, retrieval, and more.

Video-MME v2

Perception, Spatial Reasoning, Temporal Reasoning, Action / Event Understanding, Causal Reasoning, Information Retrieval

A benchmark that scores video understanding through grouped questions, factoring in answer consistency and reasoning coherence.

LVBench

Perception, Spatial Reasoning, Temporal Reasoning, Action / Event Understanding, Causal Reasoning, Information Retrieval

A benchmark tailored for comprehensive long video understanding on hour-plus videos across sports, documentaries, events, TV shows, and more.

Perception Test

Perception, Spatial Reasoning, Temporal Reasoning, Action / Event Understanding, Causal Reasoning, Information Retrieval

A benchmark by DeepMind for low- and mid-level visual perception, covering memory, physics, and semantics across video, audio, and text.

NExT-QA

Perception, Spatial Reasoning, Temporal Reasoning, Action / Event Understanding, Causal Reasoning, Information Retrieval

A video question-answering benchmark for causal and temporal reasoning about object interactions in everyday activities.

Q-Bench Video

Perception, Spatial Reasoning, Temporal Reasoning, Action / Event Understanding, Causal Reasoning, Information Retrieval

A benchmark specifically designed to evaluate LMMs' proficiency in discerning video quality.

EgoSchema

Perception, Spatial Reasoning, Temporal Reasoning, Action / Event Understanding, Causal Reasoning, Information Retrieval

A benchmark for video understanding using temporal certificate sets to measure the actual reasoning length a task requires.

UCF101-AD

Action / Event Understanding

A benchmark built to evaluate MLLM "capacity for denial," a measurement of model agreeableness and the tendency to hallucinate actions occurring.

Why these benchmarks

We choose these benchmarks as a starting point for visual intelligence based on their size, coverage, reputation, and relevant citations in recent video understanding research. If you'd like to see a benchmark added, please contact us at visionindex@ondeckai.com.

Real-world task benchmarks

Coming soon

Public benchmarks saturate, get trained on, use gameable inputs, and are built around narrow query types (VQA). A high score in these benchmarks does not signal the capability to do real-world tasks. It is also highly biased to single step, decoding-friendly tasks.

Addressing this requires designing a whole new dimension for vision evals: tasks based on real work.

Describe corrosion and damage on the subsea pipeline for infrastructure inspections

Watch ROV survey video of a subsea pipeline and fill in the inspection report: corrosion level on the anodes, freespan and vibration, any UXO, or hazards that may compromise the pipeline. Whether it warrants a scheduled repair.

Build highlight reels for newsrooms from archival & live footage

Given a story brief, search across archival & the incoming feed for every shot that could relate to the segment, return results a producer can build the highlight reel from.

Identify anomalous maritime behaviour consistent with smuggling or illegal fishing

Watch the commercial vessel traffic (not the pleasure craft) for suspicious vessels and anomalous behaviour, such as repetitive rendezvous or behaviour consistent with illegal fishing.

Triage security alerts from CCTV footage

Watch a night of camera feeds from around the facility, and rank the events worth a guard’s attention from the four hundred that aren’t, describing each the way an incident log would.

Flagging dangerous driving in dashcam footage

Review dashcam footage and flag the moments a claims adjuster would call dangerous or aggressive driving, such as tailgating, driver on their phone, or running red lights.

Monitor offshore wind turbines for bird strikes

Detect and count every bird crossing the frame, and then identify the species across buoy cameras & turbine mounted cameras.