Best video-understanding models
Models are ranked by Video-MME understanding score with current API pricing, context, and host availability. This page measures video understanding, not video-generation quality or cost-per-second output.
| # | Model | Video-MME | Input $/MTok | Output $/MTok | Hosts | Context |
|---|---|---|---|---|---|---|
| 1 | Qwen2.5 VL 72B Instruct | 73.5 | $0.25 | $0.75 | 5 | 32k |
| 2 | GPT-4o (2024-11-20) | 71.9 | $2.50 | $10.00 | 2 | 128k |
| 3 | GPT-4o (2024-08-06) | 71.9 | $2.50 | $10.00 | 4 | 128k |
| 4 | GPT-4o-mini (2024-07-18) | 64.8 | $0.15 | $0.60 | 2 | 128k |
| 5 | Qwen VL Max | 51.3 | $0.52 | $2.08 | 1 | 131k |