Open Source AI Benchmark
An independent comparison of open source AI models by benchmark scores, parameters, context window, license, and real EU pricing. Covers frontier models like Kimi K3, Qwen3.8, GLM-5.2, DeepSeek V4, and Llama 3.3 — and the models you can actually call through a sovereign EU API.
17
All Models
12
On Frontière AI
10
EU sovereign
5
Reasoning models
How open source models compare by intelligence (Arena Elo) and cost per 100k input tokens on sovereign EU infrastructure. Bigger bubbles = more parameters. Sovereign models on Frontière AI are highlighted.
Models sorted by Arena Elo by default. Click column headers to sort by MMLU, GPQA, LiveCodeBench, parameters, or price. All scores come from public model cards, technical reports, and LMSYS Chatbot Arena.
| Model | Lab | Params | Context | License | MMLU | GPQA | LiveCode | Arena Elo | Price / 100k | Jurisdiction | Verified |
|---|---|---|---|---|---|---|---|---|---|---|---|
🧠 Kimi (2.8T) | Kimi | 2.8T/~120B | 256K | MIT | 89.5 | 74.2 | 58.3 | 1342 | N/A | — | — |
🧠 Qwen (2.4T) | Qwen | 2.4T/~95B | 256K | Apache-2.0 | 89.2 | 72.8 | 56.1 | 1328 | N/A | — | — |
🧠 DeepSeek (~1.6T) | DeepSeek | ~1.6T/~100B | 128K | MIT | 88.1 | 71.5 | 55.4 | 1310 | N/A | — | — |
🧠 GLM-5.2 (Zhipu/Z.ai)On Frontière AI | Zhipu / Z.ai | 753B | 128K | MIT | 87.6 | 69.8 | 52.7 | 1295 | €0.252 | EU sovereign | 6/6 |
NVIDIA (~692B) | NVIDIA | ~692B/~60B | 128K | Apache-2.0 | 87.2 | 68.5 | 48.9 | 1278 | N/A | — | — |
Qwen3.5 397B (multimodal)On Frontière AI | Qwen | 403B/17B | 128K | Apache-2.0 | 86.4 | 66.2 | 46.3 | 1260 | €0.084 | EU sovereign | 6/6 |
Qwen3 235B (instruct)On Frontière AI | Qwen | 235B/22B | 128K | Apache-2.0 | 84.7 | 62.1 | 42.5 | 1230 | €0.105 | EU sovereign | 0/6 |
🧠 DeepSeek V4 Flash 0731 (1M context)On Frontière AI | DeepSeek | 158B/13B | 1M | MIT | 83.2 | 58.9 | 40.1 | 1210 | €0.056 | EU sovereign | 6/6 |
Llama 3.3 70BOn Frontière AI | Meta | 70B | 128K | Llama Community | 79.5 | 54.3 | 38.2 | 1180 | €0.0898 | EU sovereign | 5/6 |
Qwen3.6 27B (multimodal)On Frontière AI | Qwen | 27.8B | 128K | Apache-2.0 | 78.9 | 52.1 | 35.7 | 1165 | €0.0574 | EU sovereign | 6/6 |
Qwen3 Coder 30BOn Frontière AI | Alibaba | 30.5B/3B | 128K | Apache-2.0 | 75.3 | 48.6 | 44.2 | 1120 | €0.0084 | EU sovereign | 6/6 |
Qwen3 32BOn Frontière AI | Alibaba | 32.8B | 4K | Apache-2.0 | 76.8 | 50.4 | 33.1 | 1140 | €0.0112 | EU sovereign | 6/6 |
Mistral Small 3.2 24BOn Frontière AI | Mistral AI | 24B | 32K | Apache-2.0 | 74.2 | 46.8 | 30.5 | 1100 | €0.0126 | EU sovereign | 6/6 |
Qwen3.5 9B (fast, budget)On Frontière AI | Qwen | 9.65B | 128K | Apache-2.0 | 69.4 | 40.2 | 25.8 | 1050 | €0.014 | EU sovereign | 6/6 |
OpenAI (120B) | OpenAI | 120B/~20B | 128K | Apache-2.0 | 80.1 | 55.7 | 39.4 | 1190 | N/A | — | — |
BAAI (768M)On Frontière AI | BAAI | 768M | 8K | Apache-2.0 | — | — | — | — | N/A | — | — |
Qwen (8.2B)On Frontière AI | Qwen | 8.2B | 4K | Apache-2.0 | — | — | — | — | N/A | — | — |
Scores are sourced from official model cards and technical reports. Where multiple reports exist, we use the highest independently verified figure. All benchmarks follow their official evaluation protocol.
MMLU (5-shot)
MMLU (5-shot) — 57 subjects covering STEM, humanities, law, medicine, and more. Measures general knowledge and reasoning across domains.
GPQA Diamond
GPQA Diamond — Graduate-level science Q&A in physics, chemistry, and biology. Only the hardest subset with expert agreement, considered the gold standard for scientific reasoning.
LiveCodeBench v5
LiveCodeBench v5 — Real-world coding tasks from live competitive programming platforms. Measures practical code generation ability.
LMSYS Chatbot Arena
LMSYS Chatbot Arena Elo — Crowdsourced blind A/B comparisons. The most trusted real-world performance signal, updated continuously. Higher Elo = better user-preferred responses.
Sources: Hugging Face model cards, official technical reports, LMSYS Chatbot Arena leaderboard.
Benchmark scores measure specific capabilities under controlled conditions. They do not predict real-world performance for your specific use case. A model that scores lower on MMLU may still be better for your workload due to context length, tool calling support, language coverage, or pricing. Use benchmarks as one input among many — and test models yourself with your actual prompts.
Create an account and call any model in the catalog in minutes.