Alternatives to Grok 4.20

These models are alternatives by measurable coding evidence: they share the current closed-model classificationand are closest to Grok 4.20's SWE-bench score. Similarity here does not mean identical capabilities, licence terms, or output quality.

The target model

SWE-bench
1373.6%
Cheapest input
$1.25
Cheapest output
$2.50
Hosts
2
AlternativeSWE-benchInput $/MTokOutput $/MTokHostsContextTradeoff
Gemma 4 31B1363.6%$0.08$0.3433262kcheaper, lower coding score
Laguna M.11347.3%$0.20$0.401262kcheaper, lower coding score
Inkling Small1401.6%$0.45$1.206524khigher score, no higher input price
GLM 4.61340.2%$0.39$1.7513198kcheaper, lower coding score
Inkling1407.6%$0.95$4.056524khigher score, no higher input price
GPT-5.1-Codex1336.3%$1.25$10.003400kclosest measured match
MiMo-V2-Flash1330.2%$0.09$0.294262kcheaper, lower coding score
MiMo-V2-Pro1433.5%$1.00$3.0011.0Mhigher score, no higher input price
MiMo-V2.51437.8%$0.10$0.24131.1Mhigher score, no higher input price
Laguna XS.21302.0%$0.10$0.201262kcheaper, lower coding score
MiniMax M21297.0%$0.26$1.007197kcheaper, lower coding score
MiMo-V2.5-Pro1475.7%$0.30$0.61141.1Mhigher score, no higher input price

How alternatives are selected

Alternatives must have a benchmark score and the same current open/closed classification as the target. They are sorted by score distance, then by price/performance. The table compares hosted API economics; it does not establish licence compatibility or self-hosting rights.

Check the model licence, provider Terms, context limits, tool support, and output quality before switching a production workload.