Open-weight models have closed most of the gap with closed frontier labs in 2026, but "open-source" doesn't tell you which license you're under, how much VRAM you need, or whether you can run it on infrastructure that isn't subject to US jurisdiction. This is a working list of the models we actually route traffic to, what they're good at, and — where it applies — what it costs per million tokens on EU sovereign infrastructure.
If you need one takeaway: GLM-5.2 remains the strongest open-weight generalist served on EU sovereign infrastructure and the one we'd default to for most business use cases; Qwen3.5 397B is the only model in this list running on two separate EU sovereign clouds at once, which matters if provider redundancy is a real requirement for model deployment; and Mistral Small 3.2 is worth a specific look if the fact that the lab itself is European, not just the hosting, is part of your evaluation. Below is the full picture, model by model, with real pricing and licensing — not marketing claims. For most generative AI workloads, the ai inference cost per token is the number that matters at scale.
GLM-5.2 and GLM-5.3 — the all-rounders
Full spec sheet, licence and price →
Built by Zhipu AI / Z.ai, GLM-5.2 is a 753-billion-parameter mixture-of-experts model released under the MIT license — about as permissive as open-source licensing gets, with no restrictions on commercial use or fine-tuning. It's widely regarded as the strongest fully open-weight generalist model available as of mid-2026, competitive with closed frontier models on long-horizon coding and reasoning benchmarks. On Frontière AI, it's served through Scaleway's Generative APIs at €2.52 per million input tokens and €7.70 per million output tokens — a premium price for a premium model, and the one we'd point most teams to first if budget isn't the binding constraint.
Zhipu has since shipped the GLM-5.3 line: GLM-5.3 and the lighter GLM-5.3 Flash, both MIT-licensed and multimodal, with about 1.3M tokens of context. They extend the family's long-horizon coding and reasoning profile, but today they are routed as fast access through non-EU-controlled providers — so GLM-5.2 on Scaleway remains the sovereign-lane GLM choice. GLM-5.3 spec sheet → GLM-5.3 Flash spec sheet →
The Qwen family — from Qwen3 to Qwen3.8
Full spec sheet, licence and price →
Alibaba's Qwen family is the deepest open-weight lineup available, and Frontière AI's catalog spans five of its Apache-licensed models across three price tiers — plus the newer proprietary Qwen3.8 line below:
- Qwen3.5 397B (403B total parameters, 17B active — a mixture-of-experts design) is the flagship: multimodal, Apache-2.0 licensed, and — unusually — available on both OVHcloud and Scaleway, the only model in our catalog with that redundancy. €0.84 / €5.04 per million tokens (input/output) on Scaleway.
- Qwen3 235B is the previous generation's flagship, still fully capable and noticeably cheaper: €1.05 / €3.15 per million tokens.
- Qwen3.6 27B is a newer, smaller multimodal model — Apache-2.0, and the one Qwen model in our catalog exclusive to OVHcloud rather than Scaleway.
- Qwen3 32B is a dense (non-MoE) mid-tier workhorse, well-priced and stable since its mid-2025 release.
- Qwen3 Coder 30B is code-specialized — the model to reach for if the workload is CI pipelines, code review, or developer tooling rather than general chat.
- Qwen3.5 9B is the budget tier: small, fast, and cheap enough to run at high volume for simple classification or lightweight chat.
- Qwen3.8 Max is the newest flagship of the family: about 2.4T total parameters with 95B active, a 1M-token context, and strong reasoning — but its weights carry Qwen's own proprietary license, so check the terms before commercial redistribution. It reaches the catalog as fast access only.
- Qwen3.8 27B is the surprise of the line: its weights turned out to be Apache-2.0 (verified on its Hugging Face repository, unlike the rest of the 3.8 family — Qwen3.8 Flash stays proprietary-licensed as far as we could verify, and its repo is not public). Better still, it just crossed to EU sovereign infrastructure: OVHcloud picked it up on 2026-09-08, it is a reasoning model with a 256k context, and it costs €0.56 / €3.78 per million tokens (input/output) there.
The OVHcloud-served Qwen models are priced in USD there and converted to EUR for billing. Of the newer Qwen3.8 entries, only the 27B is routed through a sovereign provider today — the Max stays on fast-access lanes.
Mistral Small 3.2 — the European option
Full spec sheet, licence and price →
Mistral AI is a Paris-based lab, which makes Mistral Small 3.2 (24B parameters, Apache-2.0) the one model in this list where the sovereignty argument extends to the model's origin, not just where it's hosted. It supports an unusually wide range of languages — 26 are listed on its model card, more than any other model in our catalog — which matters if your business genuinely operates across multiple European markets rather than English-first with translation layered on top.
gpt-oss 120B/20B — OpenAI's own open weights
Full spec sheet, licence and price →
OpenAI released two open-weight models, gpt-oss-120b and gpt-oss-20b, in August 2025 under Apache-2.0. They're not the newest models on this list, but they're inexpensive, well-documented, and carry brand recognition that matters in procurement conversations where "which company built this" is part of the internal sign-off. Both are available on OVHcloud, at €0.112 / €0.574 per million tokens for the 120B and €0.056 / €0.224 for the 20B — the cheapest entry point in the catalog.
DeepSeek V4, Kimi K3 and R1 — the frontier tier
Full spec sheet, licence and price →
The frontier tier has moved quickly since this guide was first published. DeepSeek V4 Flash 0731 (158B total parameters, 13B active, MIT, 1M-token context) has crossed the line: it is now served on Scaleway, making it an EU sovereign option rather than a fast-access compromise — the first V4 variant with that status. DeepSeek V4 Pro 0813 (345B total, 44B active, MIT) and DeepSeek R1 (671B MoE, 39B active, MIT) remain on the fast-access tier, with smaller R1 distills (Qwen 32B and 14B) as lightweight reasoning options. Kimi K3, from Moonshot AI, is still the largest open-weight model available anywhere at 2.8 trillion parameters, with native vision support and a 1M-token context window; a dedicated EU sovereign variant of K3 is planned on Frontière AI's own infrastructure but is not callable yet.
The pace of these releases is the practical lesson: a model can move from fast access to sovereign within a single month once a managed sovereign cloud picks it up. Check the label on the model page before you call — sovereignty is per model and per provider, not a property of the family.
Beyond the original families
Several capable open-weight families sit outside the GLM/Qwen/DeepSeek/Mistral grouping of the first edition of this guide. Llama 3.3 70B from Meta is served on OVHcloud, making it one of the few large generalists available on EU sovereign infrastructure from a non-Chinese lab. Meta's newer Muse Spark 1.1 and 1.2 reasoning models (1M-token context, multimodal) are reachable only as fast access today. Gemma 3 27B from Google is Apache-2.0 licensed and inexpensive, also fast access today. Newer public releases such as Mistral Small 4 119B, Llama 4 Maverick and Gemma 4 exist but are not yet in the Frontière AI catalog, so no pricing or hosting claim is made for them here.
Licenses at a glance
| Model | License | Notably |
|---|---|---|
| GLM-5.2 | MIT | No restrictions on commercial use |
| Qwen3.5 397B / Qwen3 235B / Qwen3.6 27B / Qwen3 32B / Qwen3 Coder 30B / Qwen3.5 9B | Apache-2.0 | Permissive, patent grant included |
| Mistral Small 3.2 | Apache-2.0 | 26 languages documented |
| gpt-oss 120B / 20B | Apache-2.0 | OpenAI's first open-weight release since GPT-2 |
| DeepSeek V4 (Pro/Flash) | MIT | 1M-token context |
| Qwen2.5-VL 72B | Qwen's own license ("other") | Not Apache/MIT — check terms before commercial redistribution |
| GLM-5.3 / GLM-5.3 Flash | MIT | Successors to GLM-5.2, served as fast access today |
| Qwen3.8 Max / Flash | Qwen proprietary license | Not Apache/MIT — check terms before commercial redistribution |
| Qwen3.8 27B | Apache-2.0 | Verified on the Hugging Face repo (2026-09-08); served on OVHcloud at €0.56 / €3.78 per million tokens (input/output) |
| DeepSeek R1 (incl. distills) | MIT | Reasoning lineage, fast access today |
| Gemma 3 27B | Apache-2.0 | Google's open-weight line, fast access today |
FAQ
Which of these models should I use if I just want the best general-purpose option?
GLM-5.2 for capability, if budget allows it. Qwen3.5 397B if you specifically want a model available redundantly on two separate EU sovereign providers. Qwen3 32B if you want a solid, inexpensive default for everyday workloads.
Are DeepSeek V4 and Kimi K3 available today through Frontière AI?
They are. DeepSeek V4 Flash 0731 is now served on Scaleway and labeled EU sovereign. Kimi K3 remains fast access while a dedicated EU sovereign variant is prepared on our own infrastructure. DeepSeek V4 Pro 0813 and DeepSeek R1 are fast access today.
Which models in this guide are EU sovereign today?
Everything hosted on OVHcloud, Scaleway or our own EU servers: GLM-5.2, Qwen3.5 397B, Qwen3 235B, Qwen3.6 27B, Qwen3 32B, Qwen3 Coder 30B, Qwen3.5 9B, Qwen2.5-VL 72B, Mistral Small 3.2, gpt-oss 120B/20B, Llama 3.3 70B and DeepSeek V4 Flash 0731. GLM-5.3, Qwen3.8 Max and Flash, DeepSeek V4 Pro, DeepSeek R1, Muse Spark and Gemma 3 27B are fast access today — Qwen3.8 27B joined the sovereign list on 2026-09-08.
Why does the same model sometimes cost different amounts depending on the provider?
OVHcloud bills in USD, Scaleway in EUR, and each negotiates its own hosting economics — so the same open-weight model can carry different prices depending on which cloud serves it. Where a model is available on both (currently only Qwen3.5 397B), we route through the EUR-native provider to avoid introducing currency conversion into billing.
Is Apache-2.0 or MIT better for a commercial product?
Both are highly permissive and allow commercial use, modification, and redistribution with minimal obligations (attribution, typically). Apache-2.0 additionally includes an explicit patent grant, which some legal teams prefer for that reason. Neither imposes a copyleft requirement to open-source your own product.
What does 'mixture-of-experts' (MoE) mean for models like GLM-5.2 or Qwen3.5 397B?
It means the model has many more total parameters than it actually uses per token — Qwen3.5 397B has 403B total parameters but only activates about 17B per forward pass. That keeps inference cost and latency closer to a much smaller dense model while retaining the capacity of a much larger one.
Ready to try it?
Create an account and call any model in the catalog in minutes.