Sovereignty & compliance

Sovereign AI: build a cluster or use an API?

Published 2026-10-02 · 10 min read

There are two ways to get sovereign AI: own the hardware, or route to models run by EU-controlled providers through an API. The first is now documented in detail. In March 2026, Palantir and NVIDIA published a reference architecture for an on-premises "sovereign AI operating system", and its smallest qualified configuration is 16 servers: 4 GPU servers carrying 32 NVIDIA B300 GPUs, plus 12 CPU servers to run the platform, wired through three separate networks. Building gives you the strongest custody of your data there is, and it pays off when you need an air gap, custom weights or saturating volume. For most organisations the question is narrower: which companies in the serving chain sit under which jurisdiction. An API whose model providers are EU-controlled answers it from the first call, without a data centre.

Sovereignty has three axes, not one

"Sovereign AI" is used for very different promises, so start by separating them. Three questions decide how much control you actually have over an AI workload:

  • Custody — who physically holds your prompts, documents and outputs while they are processed, and who could hand them over.
  • Jurisdiction — under which country's law the companies holding that data operate. This is what the CLOUD Act turns on: a provider subject to US jurisdiction can be compelled to disclose data in its possession, custody or control, wherever the servers are. Hosting location and legal exposure are two different questions.
  • Dependency — whose software, hardware, updates and support you need to keep running. A system you fully control can still stop improving when a vendor stops shipping.

Owning a cluster and calling an API score differently on each axis. Neither route wins on all three, which is why the decision deserves more than a slogan.

What the reference architecture specifies

The clearest public description of the "build" route is the Palantir Sovereign AI Operating System Reference Architecture, published by Palantir and NVIDIA in March 2026. It describes a complete stack, qualified in lab testing, to run Palantir's software suite on NVIDIA hardware in a customer's own data centre: NVIDIA HGX B300 servers with eight GPUs each, NVIDIA Spectrum-X Ethernet networking, and a set of CPU servers that run the platform, all sold and deployed as one turnkey solution. The document states that the architecture is designed for air-gapped and on-premises deployments, with full autonomy after the initial installation.

It comes in three qualified sizes. The GPU servers are ordered in "scalable units" of four; the CPU "platform nodes" are fixed per size.

SizePlatform (CPU) serversGPUsIntended for (per the document)
Small3 control + 9 worker32 to 128 (1 to 4 units)Departmental pilots, specialised workloads
Medium3 control + 21 worker128 to 320 (4 to 10 units)Enterprise-wide, multi-tenant production
Large3 control + 45 worker320 to 640 (10 to 20 units)Organisation-wide, mission-critical, multi-workload

To be clear about where we stand: Frontière AI has no connection with Palantir or NVIDIA. We cite their public document because it is the most precise answer available to a question our customers ask us — what would it take to do this ourselves?

What “small” actually means

The entry point labelled "Small" is a serious installation. At its minimum, one scalable unit, it is:

  • 4 GPU servers, each with eight B300 GPUs carrying 288 GB of high-bandwidth memory apiece — 2.3 TB of GPU memory per server, 32 GPUs in total;
  • 12 CPU servers — three for the Kubernetes control plane, nine workers for the platform's data services — each specified with two 32-core processors, 1 TB of RAM and ten 7.68 TB NVMe drives;
  • three physical networks: a compute fabric between GPUs at up to 800 Gb/s per GPU, a converged network for storage and user traffic, and an isolated management network for the servers' baseboard controllers. The document standardises on Ethernet and does not support InfiniBand.

The ratio is the instructive part. Twelve of the sixteen servers never run a model. They run orchestration, deployment, monitoring, storage and access control — the machinery every AI service needs around the GPUs. A managed provider builds that machinery once and spreads its cost across all its customers. When you build, it is yours alone, sized before your first request.

The document also deploys it in a fixed order — hardware procurement, cluster bootstrap, hardened runtime, platform services, activation with your data, and only then model deployment. Each step is a project with its own owner. Rack power is a constraint too: the document allows splitting GPU servers across racks when the data centre cannot deliver the power density.

What the document does not contain is a price. Hardware in this class is quoted deal by deal, and we are not going to invent a figure. If you want the method for comparing a machine you rent or buy against per-token pricing, our self-hosting cost breakdown walks through it with dated numbers.

Why the minimum is not oversized

Thirty-two GPUs can look like excess for a first project. The model sizes say otherwise. The reference architecture notes that models beyond roughly 120 billion parameters may need to be split across several GPUs, and that hosting foundation models calls for more than eight GPUs across several servers.

Apply that to models people actually want to run. At one byte per parameter (FP8), the weights alone take roughly one gigabyte per billion parameters:

  • Mistral Large 4, the largest model on our EU-sovereign tier at 1.0 trillion parameters, needs about 1,000 GB for its weights — at least 4 B300s before any memory is left for the context of concurrent users.
  • Kimi K3, at 2.8 trillion parameters, needs about 2.8 TB — more than the 2.3 TB of a full eight-GPU server. It does not fit on one machine at all; it is a multi-server deployment from day one.

Lower-precision quantisation shrinks these figures, at a quality cost you have to measure on your own tasks. Either way, frontier open-weight models are sized for clusters, and the minimum configuration reflects that rather than inflating it. (For what open weights do and do not give you, see what open-weight AI means.)

What building buys you

Owning the hardware is the strongest form of custody available. On an air-gapped site that you operate, no outside provider holds your data, and no outside provider can be ordered to produce it. That matters for classified work, for some defence and health contexts, and for organisations whose risk analysis rules out any third party on the data path.

It also buys capabilities a shared API does not offer:

  • Custom weights — fine-tuned models, or open-weight models no provider serves;
  • Training and fine-tuning, which the architecture explicitly sizes for;
  • Guaranteed capacity — the GPUs are yours, so nobody else's traffic queues in front of yours;
  • No per-token bill — once the hardware is paid for, marginal usage is power and people.

What building leaves open

Custody is not the whole picture. In this particular design, the platform software comes from Palantir, the GPUs and networking from NVIDIA, and the reference servers from Dell — three US companies — and support is delivered through their respective contracts. On an air-gapped site that dependency does not put your data within anyone's reach, but your roadmap, licences, security updates and next hardware generation do depend on vendors you do not control. That is the dependency axis, and it is distinct from the jurisdiction question. Treat them separately in a risk analysis.

Then there are the costs no specification sheet lists: data-centre space with enough power and cooling for GPU racks, the engineers who run a Kubernetes platform and a GPU fabric, on-call coverage, and the redeployment work every time a better open-weight model ships. Our cost article breaks down why utilisation, not hardware price, usually decides whether owning pays off.

The API route, and what to check

The other route is to send requests to models run by companies under EU jurisdiction. You give up custody — the provider processes your data — but you can choose who that provider is. Sovereignty then depends entirely on the chain, so check every link:

  1. Who operates the servers that run the model, and who controls that company? An EU region of a US provider is still a US provider.
  2. Who operates the gateway or platform in between, and does it store prompts?
  3. What does each link retain, for how long, and under what exceptions?
  4. Is the label applied model by model? A catalog that mixes providers has to say which model runs where.

Here is our chain, stated plainly. Models labelled EU sovereign — 15 of the 24 live models today — run at OVHcloud and Scaleway, French companies, in data centres in France, which puts the model provider outside the CLOUD Act's direct reach. Neither of these AI services is SecNumCloud-qualified. Our gateway runs in the Paris region of Vercel, a US company, in transit only: Frontière AI does not store your prompts, and our privacy policy states each provider's retention. The other 9 live models run at US companies; our catalog and the API label them "fast access (US)", never sovereign.

That is less custody than a cluster in your basement, and we say so. What you get in exchange is the jurisdictional answer, a standard OpenAI-compatible API, and a curated catalog of open models, from the first call.

Side by side

Own cluster (reference architecture)EU-sovereign API (Frontière AI)
Smallest commitment16 servers, 32 GPUs, three networksA €10 prepaid top-up, no subscription
Custody of dataYours, air gap possibleThe model provider processes it; Frontière AI does not store prompts
Jurisdiction of the operatorYouFrench companies for sovereign models; US companies for models labelled fast access
Vendor dependencySoftware, hardware and support vendors (US in this design)Model providers and our gateway (hosted by a US company, in transit)
ModelsWhatever you deploy, custom weights includedA curated catalog of open models, no custom weights
Fine-tuningYesNo
Time to first requestProcurement, then a six-step deploymentMinutes: create a key, change the base URL
Cost shapeFixed, paid whether busy or idlePer token, nothing when idle

How to choose

Build when at least one of these is true:

  • you need an air gap, or your risk analysis excludes every third party from the data path;
  • you run custom or fine-tuned weights, or you train;
  • you have sustained, saturating volume and a team that already runs infrastructure at this scale;
  • the model you need is not served by any provider you can accept.

Otherwise, start with an API whose chain you have checked, and measure. Real token volumes are the input every build decision needs, and most teams have none before they start. You can move later: an OpenAI-compatible API means the switch is a base URL and a key, not a rewrite. If you would rather run the routing layer yourself on top of providers you choose, our review of open-source self-hosted gateways covers that middle path.

We apply the same rule to ourselves. For brand-new frontier models that no sovereign managed catalog serves yet, our plan is to rent dedicated EU servers — none is serving traffic yet: the first one is planned, not live. Kimi K3, whose sovereign variant is planned for exactly that lane and not live yet, is the textbook case: too large for one machine, too new for sovereign managed catalogs. We take on a machine when demand justifies it, not before. If you are preparing for the AI Act as well, our AI Act readiness page separates what the infrastructure can document from what remains your responsibility.

Is the future of cloud on-prem?

The reference architecture closes on one line: "The future of cloud is on-prem." For organisations that can absorb its entry point — even the size it qualifies for small enterprises starts at 16 servers — that may well be right: at that scale, owning the full stack is a reasonable answer to all three sovereignty axes at once.

For a ten-person company, a regional hospital group or a public agency's first AI project, the first step towards sovereign AI is a choice of supplier, not a data centre. Choose providers under EU jurisdiction, check the chain link by link, keep your code portable, and decide about hardware once you have the volumes to justify it.

FAQ

What hardware does an on-premises sovereign AI stack need?

Per the reference architecture Palantir and NVIDIA published in March 2026, the smallest qualified configuration is 16 servers: 4 HGX B300 GPU servers (32 GPUs, 288 GB of memory each) plus 12 CPU servers that run the platform — 3 for the control plane, 9 workers — connected by three separate networks. The largest qualified size reaches 640 GPUs and 48 platform servers. The document gives no price.

Is an on-premises AI cluster automatically sovereign?

It gives you the strongest custody of your data: on an air-gapped site you operate, no outside provider holds it. It does not remove dependency on the vendors of your software, hardware and support — in the Palantir and NVIDIA reference design, three US companies. Custody, jurisdiction and dependency are separate axes, and a risk analysis should treat each one.

Can an API be sovereign?

For the jurisdiction axis, yes, if every company that processes your data is under EU control. Check who operates the model servers and who controls that company, whether the gateway in between stores prompts, and what each link retains. On Frontière AI, 15 live models run at OVHcloud and Scaleway in France and are labelled EU sovereign; our gateway runs in Vercel's Paris region, in transit only; models served by US companies are labelled fast access, never sovereign.

How many GPUs does a frontier open-weight model need?

At one byte per parameter (FP8), roughly one gigabyte per billion parameters for the weights alone. Mistral Large 4 (1.0 trillion) needs about 1,000 GB, at least 4 NVIDIA B300s with 288 GB each before any room for context. Kimi K3 (2.8 trillion) needs about 2.8 TB, more than one eight-GPU server holds, so it requires several servers. Quantisation reduces this at a quality cost you have to measure.

When does building your own AI cluster make sense?

When you need an air gap, run custom or fine-tuned weights, train models, or have sustained saturating volume and a team that already runs infrastructure at this scale. Otherwise, starting with an EU-sovereign API costs a €10 prepaid top-up, with no subscription, gives you real usage figures, and keeps the option open: an OpenAI-compatible API means switching later is a change of base URL and key.

Ready to try it?

Create an account and call any model in the catalog in minutes.

Create an account