Rank #5 / 9
Kimi K3
Kimi K3 is Moonshot's Kimi model (open weights). It is ranked 5 of 9 models on the Visual Intelligence Index.
- Moonshot
- Open weights
- Kimi
- Video input
- Structured output
- Tool calling
- Streaming
- Self-hostable
Comparison Summary
Kimi K3 ranks 5 of 9 on the perception index. Immediately ahead is Qwen3.8-Max. Immediately behind is Qwen3.8-27B. List price is $3.00 per million input tokens and $15.00 per million output. 2 of 5 priced models cost less to prompt. The context window is 1.05M. Every score is from our own evaluation runs; list prices are as published by the provider. Strongest capability in this set is Causal Reasoning (75%). Weakest is Action / Event Understanding (67%).
- Visual Intelligence
- 71.9%
- Availability
- Open weights
- Input / 1M tokens
- $3.00
- Output / 1M tokens
- $15.00
- Context Window
- 1.05M
- Parameters
- 3T
All benchmark scores
| Benchmark | Metric | Score | Coverage | Setting |
|---|---|---|---|---|
| Video-MME v2 | Accuracy, no subtitles | 55.2% | Full | 1 fps |
| LVBench | Test accuracy | 66.8% | Full | 1 fps |
| Perception Test | Overall accuracy | 82.4% | Full | 1 fps |
| NExT-QA | Accuracy (hard split) | 86.3% | Full | 1 fps |
| Q-Bench Video | Overall accuracy | 67.5% | Full | 1 fps |
| EgoSchema | Accuracy (fullset) | 78.6% | Full | 1 fps |
| UCF101-AD | Accuracy | 39.4% | Full | 1 fps |
Benchmark Scores
Compare reported model scores across each available benchmark or capability index.
9 of 9 models
Display
Capabilities Index Scores
Compare reported model scores across each available benchmark or capability index.
9 of 9 models
Display
Usability
4.0/ 5
What it's like to build against this model, scored out of five from the developer-facing capabilities in the dataset.
- Reachable1.0
There is an endpoint you can call without hosting anything.
- First-party API — supported
- Third-party API — supported
- Portable1.0
You can run it yourself, and are not tied to one vendor.
- Open weights — supported
- Self-hostable — supported
- Multimodal input0.5
It takes the footage directly, rather than frames you extracted.
- Video ingestion — supported
- Audio — not supported
- Programmable1.0
Output you can parse, and tools it can call on its own.
- Structured output — supported
- Tool calling — supported
- Operable0.5
Usable interactively and in bulk, not only one call at a time.
- Streaming — supported
- Batch API — not supported
Other Moonshot models
Other models from the Moonshot family.
| Date | Model | Visual Intelligence |
|---|---|---|
| Jul 16, 2026 | Kimi K3 | 71.9% |