โ† Back to Rankings

๐Ÿ“ Scoring Methodology

Each model is scored on a weighted 4-criteria matrix. Only open-weight/open-source models (MIT, Apache 2.0, Gemma, Llama licenses).

Arena ELO
35%
Chatbot Arena+ overall ranking
Coding Performance
30%
SWE-bench, HumanEval, LiveCodeBench
Reasoning
15%
GPQA Diamond, MMLU-Pro
Value
20%
Price per token, license, efficiency

Sources: Chatbot Arena+ ยท Onyx LLM ยท LLM Stats ยท BenchLM

#1
โ˜…โ˜…โ˜…โ˜…โ˜…
9.4/10
Best for: Coding, reasoning, overall capability 1.6T MoE MIT 1M context
DeepSeek-V4-Pro is the top-ranked open-source model at 1467 Arena ELO, the highest among open-weight models. Leads in coding with 80.6% SWE-bench Verified and 93.5 LiveCodeBench. Strong reasoning at GPQA Diamond 90.1%. MIT license, 1M token context, and competitive API pricing at $1.74/$0.87 per 1M tokens.
Arena ELO: 1467 openlm.ai SWE-bench: 80.6% Onyx LLM GPQA Diamond: 90.1% Price: $1.74 / $0.87 per 1M
Get Config โ†’ Dev Agent Toolkit $19
#2
โ˜…โ˜…โ˜…โ˜…โ˜…
8.9/10
Best for: Knowledge & MMLU-Pro leader 397B (17B active) Apache 2.0 Alibaba
Qwen 3.5-397B-A17B by Alibaba scores 1450 Arena ELO with the highest MMLU-Pro among open models at 87.8%. Its efficient MoE architecture uses only 17B active parameters out of 397B total, which is remarkably efficient. Apache 2.0 license. Strong general knowledge and reasoning capabilities. Linear attention mechanism for long-context performance.
Arena ELO: 1450 openlm.ai MMLU-Pro: 87.8% Arena+ Active params: 17B (of 397B) Price: $0.39 / $2.45 per 1M
Get Config โ†’ Workflow Playbook $27
#3
โ˜…โ˜…โ˜…โ˜…โ˜…
8.8/10
Best for: Performance per parameter 31B dense Apache 2.0 Google
Gemma 4 31B by Google punches far above its weight class. At just 31B dense parameters, it scores 1449 Arena ELO, competitive with models 10x its size. Apache 2.0 license, 262K context. Remarkable efficiency: 85.2 MMLU-Pro. Available at $0.13/$0.38 per 1M tokens. The best value dense model in 2026.
Arena ELO: 1449 openlm.ai Params: 31B (dense) MMLU-Pro: 85.2% Price: $0.13 / $0.38 per 1M
Get Config โ†’ Dev Agent Toolkit $19
#4
โ˜…โ˜…โ˜…โ˜…โ˜†
8.4/10
Best for: European AI leader, multilingual 675B MoE Apache 2.0 Mistral AI
Mistral Large 3 is Mistral AI's flagship open-weight model at 1428 Arena ELO. 675B MoE architecture with Apache 2.0 license and 256K context. Strong on multilingual tasks and reasoning (81.0 MMLU-Pro). European AI leader with a focus on developer-friendly APIs. Available through La Platforme and multiple cloud providers.
Arena ELO: 1428 openlm.ai Params: 675B MoE MMLU-Pro: 81.0% License: Apache 2.0
Get Config โ†’ Workflow Playbook $27
#5
โ˜…โ˜…โ˜…โ˜…โ˜†
8.3/10
Best for: Budget coding & reasoning 284B MoE MIT $0.14/M input
DeepSeek-V4-Flash is the distilled version of V4-Pro at 1445 Arena ELO, only 22 points behind the Pro variant. Costs just $0.14/$0.28 per 1M tokens, roughly 10x cheaper than V4-Pro while maintaining 98% of its benchmark performance. 79.0% SWE-bench Verified, 91.6 LiveCodeBench. MIT license, 1M context. Insane value.
Arena ELO: 1445 openlm.ai SWE-bench: 79.0% Onyx LLM Params: 284B MoE Price: $0.14 / $0.28 per 1M
Get Config โ†’ Dev Agent Toolkit $19
#6
โ˜…โ˜…โ˜…โ˜…โ˜†
7.8/10
Best for: Meta ecosystem & long context 400B (17B active) Llama 4 license 1M context
Llama 4 Maverick is Meta's flagship open-weight model at 1292 Arena ELO. 400B total parameters with 17B active (128 experts). Despite lower raw benchmarks, Llama models have the widest self-hosting ecosystem and community support. 1M token context, 80.9 MMLU-Pro. Strong choice if you're in the Meta/PyTorch ecosystem.
Arena ELO: 1292 openlm.ai Params: 400B (17B active, 128E) MMLU-Pro: 80.9% Context: 1M tokens
Get Config โ†’ Workflow Playbook $27
#7
โ˜…โ˜…โ˜…โ˜…โ˜†
7.5/10
Best for: Self-hosted free model 27B dense Gemma license Free self-host
Gemma 3 27B by Google is a solid mid-tier open model at 1356 Arena ELO. HumanEval 87.8% coding pass rate, impressive for a 27B model. Completely free for self-hosting under the Gemma license. Good entry point for teams wanting to run their own model without API costs. Superseded by Gemma 4 for higher performance.
Arena ELO: 1356 openlm.ai HumanEval: 87.8% Params: 27B dense Price: Free (self-hosted)
Get Config โ†’ Dev Agent Toolkit $19
#8
โ˜…โ˜…โ˜…โ˜…โ˜†
7.2/10
Best for: Small on-device models 14B dense MIT Free self-host
Phi-4 by Microsoft is the best small open-source model at just 14B parameters. Scores 1222 Arena ELO with 71.4 MMLU-Pro, impressive density. MIT license, free to self-host and modify. Runs on consumer hardware. Great for on-device applications, fine-tuning experiments, and cost-sensitive deployments where model size matters.
Arena ELO: 1222 openlm.ai MMLU-Pro: 71.4% Params: 14B dense License: MIT
Get Config โ†’ Workflow Playbook $27

Want Optimized Configs for These Models?

Our config packs include optimized prompts and system instructions for DeepSeek, Qwen, Gemma, Mistral, and more. Ready to use in any agent.