Skip to content

All Models

Rank #2 / 9

Gemini 3.8 Flash

Gemini 3.8 Flash is Google's Gemini model (proprietary). It is ranked 2 of 9 models on the Visual Intelligence Index.

  • Google
  • Proprietary
  • Gemini
  • Video input
  • Audio
  • Structured output
  • Tool calling
  • Streaming
  • Batch API
Model Card
77.8%Visual Intelligence Index · Rank #2 of 9

Capability profile

Comparison Summary

Gemini 3.8 Flash ranks 2 of 9 on the perception index. Immediately ahead is GPT-6 Astra. Immediately behind is Claude Fable 5.1. List price is $0.75 per million input tokens and $3.75 per million output. 0 of 5 priced models cost less to prompt. The context window is 1M. Every score is from our own evaluation runs; list prices are as published by the provider. Strongest capability in this set is Spatial Reasoning (83%). Weakest is Action / Event Understanding (78%).

Visual Intelligence
80.6%
Availability
Proprietary
Input / 1M tokens
$0.75
Output / 1M tokens
$3.75
Context Window
1M
Parameters
—

All benchmark scores

BenchmarkMetricScoreCoverageSetting
Video-MME v2Accuracy, no subtitles71.3%Full1 fps
LVBenchTest accuracy86.4%Full1 fps, 5889-frame cap (18 videos thinned)
Perception TestOverall accuracy85.5%Full1 fps
NExT-QAAccuracy (hard split)88.0%Full1 fps
Q-Bench VideoOverall accuracy69.3%Full1 fps
EgoSchemaAccuracy (fullset)80.8%Full1 fps
UCF101-ADAccuracy63.5%Full1 fps

Benchmark Scores

Compare reported model scores across each available benchmark or capability index.

9 of 9 models
Filter by model access

Display

Capabilities Index Scores

Compare reported model scores across each available benchmark or capability index.

9 of 9 models
Filter by model access

Display

Usability

4.0/ 5

What it's like to build against this model, scored out of five from the developer-facing capabilities in the dataset.

Reachable1.0

There is an endpoint you can call without hosting anything.

  • First-party API — supported
  • Third-party API — supported
Portable0.0

You can run it yourself, and are not tied to one vendor.

  • Open weights — not supported
  • Self-hostable — not supported
Multimodal input1.0

It takes the footage directly, rather than frames you extracted.

  • Video ingestion — supported
  • Audio — supported
Programmable1.0

Output you can parse, and tools it can call on its own.

  • Structured output — supported
  • Tool calling — supported
Operable1.0

Usable interactively and in bulk, not only one call at a time.

  • Streaming — supported
  • Batch API — supported

Other Google models

Other models from the Google family.

DateModelVisual Intelligence
Sep 2, 2026Gemini 3.8 Flash80.6%