Alternatives to gpt-5.5 (<272K context length)

These models are alternatives by measurable coding evidence: they share the current closed-model classificationand are closest to gpt-5.5 (<272K context length)'s SWE-bench score. Similarity here does not mean identical capabilities, licence terms, or output quality.

The target model

SWE-bench
82.6%
Cheapest input
$5.00
Cheapest output
$30.00
Hosts
6
AlternativeSWE-benchInput $/MTokOutput $/MTokHostsContextTradeoff
Muse Spark 1.182.0%$1.25$4.2521.0Mlower input price
Claude Opus 4.783.5%$5.00$25.00101.0Mclosest measured match
Gemini 3.7 Flash80.8%$0.19$0.944lower input price
Gemini 3.6 Flash79.6%$0.38$1.884lower input price
Claude Sonnet 579.6%$2.00$10.0010lower input price
Gemini 3.5 Flash78.8%$0.75$4.504lower input price
Gemini 3.1 Pro Preview78.8%$1.00$6.0041.0Mlower input price
Claude Opus 4.678.7%$5.00$25.00101.0Mclosest measured match
GLM 5.278.7%$0.50$2.0066203klower input price
Muse Spark 1.286.6%$1.25$4.2521.0Mlower input price
Grok 4.586.6%$2.00$6.002500klower input price
gpt-5.4 (<272K context length)78.2%$2.50$15.006lower input price

How alternatives are selected

Alternatives must have a benchmark score and the same current open/closed classification as the target. They are sorted by score distance, then by price/performance. The table compares hosted API economics; it does not establish licence compatibility or self-hosting rights.

Check the model licence, provider Terms, context limits, tool support, and output quality before switching a production workload.