Rank #2 / 9
Gemini 3.8 Flash
Gemini 3.8 Flash is Google's Gemini model (proprietary). It is ranked 2 of 9 models on the Visual Intelligence Index.
- Proprietary
- Gemini
- Video input
- Audio
- Structured output
- Tool calling
- Streaming
- Batch API
Comparison Summary
Gemini 3.8 Flash ranks 2 of 9 on the perception index. Immediately ahead is GPT-6 Astra. Immediately behind is Claude Fable 5.1. List price is $0.75 per million input tokens and $3.75 per million output. 0 of 5 priced models cost less to prompt. The context window is 1M. Every score is from our own evaluation runs; list prices are as published by the provider. Strongest capability in this set is Spatial Reasoning (83%). Weakest is Action / Event Understanding (78%).
- Visual Intelligence
- 80.6%
- Availability
- Proprietary
- Input / 1M tokens
- $0.75
- Output / 1M tokens
- $3.75
- Context Window
- 1M
- Parameters
- —
All benchmark scores
| Benchmark | Metric | Score | Coverage | Setting |
|---|---|---|---|---|
| Video-MME v2 | Accuracy, no subtitles | 71.3% | Full | 1 fps |
| LVBench | Test accuracy | 86.4% | Full | 1 fps, 5889-frame cap (18 videos thinned) |
| Perception Test | Overall accuracy | 85.5% | Full | 1 fps |
| NExT-QA | Accuracy (hard split) | 88.0% | Full | 1 fps |
| Q-Bench Video | Overall accuracy | 69.3% | Full | 1 fps |
| EgoSchema | Accuracy (fullset) | 80.8% | Full | 1 fps |
| UCF101-AD | Accuracy | 63.5% | Full | 1 fps |
Benchmark Scores
Compare reported model scores across each available benchmark or capability index.
9 of 9 models
Display
Capabilities Index Scores
Compare reported model scores across each available benchmark or capability index.
9 of 9 models
Display
Usability
4.0/ 5
What it's like to build against this model, scored out of five from the developer-facing capabilities in the dataset.
- Reachable1.0
There is an endpoint you can call without hosting anything.
- First-party API — supported
- Third-party API — supported
- Portable0.0
You can run it yourself, and are not tied to one vendor.
- Open weights — not supported
- Self-hostable — not supported
- Multimodal input1.0
It takes the footage directly, rather than frames you extracted.
- Video ingestion — supported
- Audio — supported
- Programmable1.0
Output you can parse, and tools it can call on its own.
- Structured output — supported
- Tool calling — supported
- Operable1.0
Usable interactively and in bulk, not only one call at a time.
- Streaming — supported
- Batch API — supported
Other Google models
Other models from the Google family.
| Date | Model | Visual Intelligence |
|---|---|---|
| Sep 2, 2026 | Gemini 3.8 Flash | 80.6% |