Skip to content

All Models

Rank #5 / 9

Kimi K3

Kimi K3 is Moonshot's Kimi model (open weights). It is ranked 5 of 9 models on the Visual Intelligence Index.

  • Moonshot
  • Open weights
  • Kimi
  • Video input
  • Structured output
  • Tool calling
  • Streaming
  • Self-hostable
Model Card
68.0%Visual Intelligence Index · Rank #5 of 9

Capability profile

Comparison Summary

Kimi K3 ranks 5 of 9 on the perception index. Immediately ahead is Qwen3.8-Max. Immediately behind is Qwen3.8-27B. List price is $3.00 per million input tokens and $15.00 per million output. 2 of 5 priced models cost less to prompt. The context window is 1.05M. Every score is from our own evaluation runs; list prices are as published by the provider. Strongest capability in this set is Causal Reasoning (75%). Weakest is Action / Event Understanding (67%).

Visual Intelligence
71.9%
Availability
Open weights
Input / 1M tokens
$3.00
Output / 1M tokens
$15.00
Context Window
1.05M
Parameters
3T

All benchmark scores

BenchmarkMetricScoreCoverageSetting
Video-MME v2Accuracy, no subtitles55.2%Full1 fps
LVBenchTest accuracy66.8%Full1 fps
Perception TestOverall accuracy82.4%Full1 fps
NExT-QAAccuracy (hard split)86.3%Full1 fps
Q-Bench VideoOverall accuracy67.5%Full1 fps
EgoSchemaAccuracy (fullset)78.6%Full1 fps
UCF101-ADAccuracy39.4%Full1 fps

Benchmark Scores

Compare reported model scores across each available benchmark or capability index.

9 of 9 models
Filter by model access

Display

Capabilities Index Scores

Compare reported model scores across each available benchmark or capability index.

9 of 9 models
Filter by model access

Display

Usability

4.0/ 5

What it's like to build against this model, scored out of five from the developer-facing capabilities in the dataset.

Reachable1.0

There is an endpoint you can call without hosting anything.

  • First-party API — supported
  • Third-party API — supported
Portable1.0

You can run it yourself, and are not tied to one vendor.

  • Open weights — supported
  • Self-hostable — supported
Multimodal input0.5

It takes the footage directly, rather than frames you extracted.

  • Video ingestion — supported
  • Audio — not supported
Programmable1.0

Output you can parse, and tools it can call on its own.

  • Structured output — supported
  • Tool calling — supported
Operable0.5

Usable interactively and in bulk, not only one call at a time.

  • Streaming — supported
  • Batch API — not supported

Other Moonshot models

Other models from the Moonshot family.

DateModelVisual Intelligence
Jul 16, 2026Kimi K371.9%