Open Source AI Benchmark

Open Source AI Benchmark — Model Comparison 2026

An independent comparison of open source AI models by benchmark scores, parameters, context window, license, and real EU pricing. Covers frontier models like Kimi K3, Qwen3.8, GLM-5.2, DeepSeek V4, and Llama 3.3 — and the models you can actually call through a sovereign EU API.

17

All Models

12

On Frontière AI

10

EU sovereign

5

Reasoning models

Benchmark scores at a glance

MMLU
GPQA
Sovereign (Frontière AI)
Not on Frontière AI
Kimi 2.8T
89.5
Qwen 2.4T
89.2
DeepSeek ~1.6T
88.1
GLM-5.2 (Zhipu/Z.ai)
87.6
NVIDIA ~692B
87.2
Qwen3.5 397B (multimodal)
86.4
Qwen3 235B (instruct)
84.7
DeepSeek V4 Flash 0731 (1M context)
83.2
Llama 3.3 70B
79.5
Qwen3.6 27B (multimodal)
78.9
Qwen3 Coder 30B
75.3
Qwen3 32B
76.8
Mistral Small 3.2 24B
74.2
Qwen3.5 9B (fast, budget)
69.4
OpenAI 120B
80.1
0255075100

Capability vs. price

How open source models compare by intelligence (Arena Elo) and cost per 100k input tokens on sovereign EU infrastructure. Bigger bubbles = more parameters. Sovereign models on Frontière AI are highlighted.

13601245113010159000.000.060.120.170.230.29Price per 100k input tokens (€)Arena EloGLM-5.2 (Zhipu/Z.ai)GLM-5.2 (Zhipu/Z.ai) Arena Elo: 1295 Price / 100k: €0.2520 Params: 753BQwen3.5 397B (multimodal)Qwen3.5 397B (multimodal) Arena Elo: 1260 Price / 100k: €0.0840 Params: 403BQwen3 235B (instruct)Qwen3 235B (instruct) Arena Elo: 1230 Price / 100k: €0.1050 Params: 235BDeepSeek V4 Flash 0731 (1M context)DeepSeek V4 Flash 0731 (1M context) Arena Elo: 1210 Price / 100k: €0.0560 Params: 158BLlama 3.3 70BLlama 3.3 70B Arena Elo: 1180 Price / 100k: €0.0898 Params: 70BQwen3.6 27B (multimodal)Qwen3.6 27B (multimodal) Arena Elo: 1165 Price / 100k: €0.0574 Params: 28BQwen3 Coder 30BQwen3 Coder 30B Arena Elo: 1120 Price / 100k: €0.0084 Params: 31BQwen3 32BQwen3 32B Arena Elo: 1140 Price / 100k: €0.0112 Params: 33BMistral Small 3.2 24BMistral Small 3.2 24B Arena Elo: 1100 Price / 100k: €0.0126 Params: 24BQwen3.5 9B (fast, budget)Qwen3.5 9B (fast, budget) Arena Elo: 1050 Price / 100k: €0.0140 Params: 10BSovereign (Frontière AI)Not on Frontière AI

Models sorted by Arena Elo by default. Click column headers to sort by MMLU, GPQA, LiveCodeBench, parameters, or price. All scores come from public model cards, technical reports, and LMSYS Chatbot Arena.

ModelLabParamsContextLicenseMMLUGPQALiveCodeArena EloPrice / 100kJurisdictionVerified
🧠
Kimi (2.8T)
Kimi2.8T/~120B256KMIT89.574.258.31342N/A
🧠
Qwen (2.4T)
Qwen2.4T/~95B256KApache-2.089.272.856.11328N/A
🧠
DeepSeek (~1.6T)
DeepSeek~1.6T/~100B128KMIT88.171.555.41310N/A
🧠
GLM-5.2 (Zhipu/Z.ai)On Frontière AI
Zhipu / Z.ai753B128KMIT87.669.852.71295€0.252EU sovereign6/6
NVIDIA (~692B)
NVIDIA~692B/~60B128KApache-2.087.268.548.91278N/A
Qwen403B/17B128KApache-2.086.466.246.31260€0.084EU sovereign6/6
Qwen3 235B (instruct)On Frontière AI
Qwen235B/22B128KApache-2.084.762.142.51230€0.105EU sovereign0/6
DeepSeek158B/13B1MMIT83.258.940.11210€0.056EU sovereign6/6
Llama 3.3 70BOn Frontière AI
Meta70B128KLlama Community79.554.338.21180€0.0898EU sovereign5/6
Qwen3.6 27B (multimodal)On Frontière AI
Qwen27.8B128KApache-2.078.952.135.71165€0.0574EU sovereign6/6
Qwen3 Coder 30BOn Frontière AI
Alibaba30.5B/3B128KApache-2.075.348.644.21120€0.0084EU sovereign6/6
Qwen3 32BOn Frontière AI
Alibaba32.8B4KApache-2.076.850.433.11140€0.0112EU sovereign6/6
Mistral Small 3.2 24BOn Frontière AI
Mistral AI24B32KApache-2.074.246.830.51100€0.0126EU sovereign6/6
Qwen9.65B128KApache-2.069.440.225.81050€0.014EU sovereign6/6
OpenAI (120B)
OpenAI120B/~20B128KApache-2.080.155.739.41190N/A
BAAI (768M)On Frontière AI
BAAI768M8KApache-2.0N/A
Qwen (8.2B)On Frontière AI
Qwen8.2B4KApache-2.0N/A

Benchmark methodology

Scores are sourced from official model cards and technical reports. Where multiple reports exist, we use the highest independently verified figure. All benchmarks follow their official evaluation protocol.

MMLU (5-shot)

MMLU (5-shot) — 57 subjects covering STEM, humanities, law, medicine, and more. Measures general knowledge and reasoning across domains.

GPQA Diamond

GPQA Diamond — Graduate-level science Q&A in physics, chemistry, and biology. Only the hardest subset with expert agreement, considered the gold standard for scientific reasoning.

LiveCodeBench v5

LiveCodeBench v5 — Real-world coding tasks from live competitive programming platforms. Measures practical code generation ability.

LMSYS Chatbot Arena

LMSYS Chatbot Arena Elo — Crowdsourced blind A/B comparisons. The most trusted real-world performance signal, updated continuously. Higher Elo = better user-preferred responses.

Sources: Hugging Face model cards, official technical reports, LMSYS Chatbot Arena leaderboard.

How to read benchmark scores

Benchmark scores measure specific capabilities under controlled conditions. They do not predict real-world performance for your specific use case. A model that scores lower on MMLU may still be better for your workload due to context length, tool calling support, language coverage, or pricing. Use benchmarks as one input among many — and test models yourself with your actual prompts.

Frequently asked questions

What is the best open source AI model in 2026?
Kimi K3 and Qwen3.8 2.4T lead in raw intelligence (Arena Elo 1340+), followed closely by DeepSeek V4 Pro and GLM-5.2. For practical use, GLM-5.2 is the strongest model available through a sovereign EU API — scoring 87.6% on MMLU and 69.8% on GPQA Diamond, served through Scaleway starting at €0.25 per 100k input tokens.
Can I self-host open source AI models?
Yes — all models listed here have downloadable weights. Requirements range from a single GPU for small models (Qwen3.5 9B, ~20 GB VRAM) to multi-GPU setups for frontier models (Kimi K3 requires 8×H100 minimum). Alternatively, Frontière AI provides instant API access to sovereign-hosted models starting at €10 prepaid.
Are open source AI models GDPR-compliant?
The model itself is never the issue — it's where it runs. A model hosted on OVHcloud or Scaleway in France stays under EU jurisdiction. The same model hosted on an US cloud provider is subject to the CLOUD Act. Frontière AI labels every model as 'EU sovereign' or 'fast access' so you know before you call.
Why do benchmark scores differ between sources?
Different benchmark providers use different prompts, temperature settings, and evaluation splits. We source scores from official model cards and technical reports to ensure consistency. Even then, a model reported at 87.6% MMLU by its creator might score 86% in an independent evaluation — both are legitimate measurements of the same model.
What's the difference between reasoning and non-reasoning models?
Reasoning models (marked with a thinking indicator) perform chain-of-thought internally before answering. They typically score higher on benchmarks involving math, code, and complex reasoning — but may be slower and more expensive due to extra tokens spent thinking. Non-reasoning models are faster and cheaper for straightforward tasks.
How does Frontière AI price open source models?
Frontière AI applies a flat 40% margin on top of infrastructure costs. Sovereign providers (OVHcloud, Scaleway) are the cheapest — starting at €0.0045 per 100k input tokens for Qwen3.5 9B. Frontier models like GLM-5.2 start at €0.25 per 100k input tokens. No subscription, purely prepaid.

Try these models yourself

Create an account and call any model in the catalog in minutes.

Create an account