Best coding models for design

This page uses the WebDev Arena benchmark as a frontend/design signal, then adds current token price, context, and host availability. It is not a general visual-design quality ranking: the benchmark source is an external dataset and remains non-commercial until its upstream licence is confirmed.

#ModelWebDev scoreInput $/MTokOutput $/MTokHostsContext
1Claude Opus 51690.6$5.00$25.00101.0M
2Kimi K31673.7$2.55$12.75291.0M
3Qwen3.8 Max1668.9$2.00$6.0021.0M
4Claude Fable 51626.3$10.00$50.00101.0M
5GLM 5.31598.6$1.40$4.4021.0M
6Qwen3.8 27B1594.8$0.35$2.5520262k
7GLM 5.21593.3$0.50$2.0066203k
8Claude Opus 4.71566.8$5.00$25.00101.0M
9Grok 4.51556.4$2.00$6.002500k
10Claude Opus 4.61556.3$5.00$25.00101.0M
11Claude Opus 4.81552.2$5.00$25.00101.0M
12Muse Spark 1.11539.0$1.25$4.2521.0M
13Gemini 3.6 Flash1527.8$0.38$1.884
14Claude Sonnet 4.61522.5$3.00$15.00101.0M
15Qwen3.7 Max1516.7$1.48$4.4221.0M
16GLM 5.11508.9$0.91$2.8644203k
17Kimi K2.61508.7$0.55$2.5248256k
18Gemini 3.5 Flash1506.4$0.75$4.504
19MiniMax M31487.6$0.23$0.9628262k
20Qwen3.6 Max Preview1478.9$1.03$6.162262k
21MiMo-V2.5-Pro1475.7$0.30$0.61141.1M
22Kimi K2.7 Code1472.7$0.67$3.4035262k
23DeepSeek V4 Pro 04231463.6$0.79$1.74411.0M
24Qwen3.6 Plus1459.7$0.33$1.9521.0M
25gpt-5.5 (<272K context length)1457.5$5.00$30.006