For air-gapped & sovereign deployments
Your data stays put. The model comes to it.
Most AI tools assume they can reach the internet. On an air-gapped network they can’t, and a residency rule may forbid it. We build AI that runs entirely on infrastructure you control, and we’ll tell you which workloads don’t need that. We learned how in hospitals.
The exposure
Every default in the AI stack assumes your data can leave.
- 01
Your staff already use AI tools you never approved.
In IBM’s 2025 study, one in five breached organizations reported a breach involving shadow AI, and heavy users paid about US$670,000 more per breach. A ban fails if staff have nothing approved to use.a
- 02
- 03
The model you validated retires on someone else’s schedule.
Microsoft retires a hosted model version 18 months after launch, with no extensions. Google’s long-term Gemini models get at least 12 months. Each retirement means validating your workflows again. Hold the weights and you choose when the model changes.d
- 04
An air gap breaks every install script.
Installing the vLLM inference server pulls in 196 Python packages, 27 of them from NVIDIA. Behind an air gap, someone mirrors, scans, approves, and carries across each one, at every update.e
- 05
The spectrum
Sovereignty is a spectrum.An air gap is the far end of it.
Most teams we work with need two: one for everyday work, and a stricter one for the workloads that carry real risk.
Ordered by isolation, least to most
Tier 1 of 4
Hosted API
A frontier model under an enterprise contract.
- Guards against
- Retention, and training on your data.
- Trade-off
- Their jurisdiction and their retirement schedule.
Right for
Most everyday work
Tier 2 of 4
Sovereign cloud
Dedicated capacity in your region, sometimes run by a local operator.
- Guards against
- Data leaving the region.
- Trade-off
- Still networked, and you don’t hold the weights.
Right for
In-region mandates
Tier 3 of 4
Self-hosted
Open-weight models on your hardware, with outbound traffic blocked.
- Guards against
- Data leaving, forced upgrades, and foreign legal process.
- Trade-off
- You run the GPUs, and open models still trail the best closed ones.i
Right for
Sensitive data, validated workflows
Tier 4 of 4
Air-gapped
No network path at all. Updates cross on approved media.
- Guards against
- Everything above, plus attacks over the network.
- Trade-off
- Every update becomes a procedure.
Right for
Plant floors, isolated labs, trade secrets
We will tell you which tier each workload needs, including when the answer is the cloud.
What transfers
We learned this in hospitals, where the data can’t leave either.
- Federated training across 71 sites on six continents3One model trained across countries or partners, each dataset kept where the law requires.
- Inference tuned for limited clinical hardware4jA useful model on GPUs you can actually buy and run on-site.
- Reproducible builds and maintained conda-forge packagesEvery dependency pinned and mirrored, so the stack rebuilds offline.
What we do for sovereign deployments
Six services, run inside your perimeter.
We don’t sell GPUs, models, or licenses. We build on what you own and hand it to your team.
01
Readiness & architecture sprint
Every workload placed on a tier, with hardware sized and a build-or-buy call.
02
Data engineering inside the perimeter
Pipelines and search indexes built where the data lives, never copied out.
03
Model selection & evaluation
Open-weight models tuned on your data and tested on cases you write.
04
Air-gapped deployment & MLOps
Mirrored registries, a written transfer procedure, and monitoring that sends nothing out.
05
Governance & documentation
Data-flow maps and validation files, written for your assessor.
06
Fractional AI leadership
A senior AI lead a few days a month, testing every vendor’s “sovereign” claim.
How we engage
Start with one workload, on one tier.
Each engagement has its own scope, fee, and exit criteria.
Readiness sprint
Two to four weeks at a fixed fee, covering every workload you’re weighing.
A costed plan your security office has already read.
Private assistant pilot
One team and one document set, on an open-weight model running on your hardware.
A visible win, and an offline stack you keep.
Stalled-pilot rescue
For a pilot stuck at security review. We close the findings and rebuild what can’t pass.
The pilot back in front of the reviewers.
We build and document the controls. Your security office and your assessor sign off on them.
Bring the model to the data.
Start with a readiness sprint or a one-workload pilot on your own hardware.
Sources & attribution
- 1. Led by our founder across NIH/NCI-funded programs at the University of Pennsylvania and Indiana University.
- 2. Vaiyu client engagements: pre-training optimization with model accuracy maintained or improved.
- 3. Pati, S. et al. “Federated learning enables big data for rare cancer boundary detection.” Nature Communications 13 (2022). doi:10.1038/s41467-022-33407-5; 71 sites across 6 continents, the largest real-world federated learning study to date.
- 4. Founder track record at Indiana University: inference latency reduced by up to 70%, compute requirements by 10–50%, in clinical research environments.
- a. IBM and Ponemon Institute, Cost of a Data Breach Report 2025: 600 breached organizations; 20% reported a shadow-AI breach; average cost USD 4.74M with high shadow-AI use, USD 4.07M with low or none.
- b. CLOUD Act, 18 U.S.C. § 2713: data in a provider’s “possession, custody, or control,” wherever located.
- c. French Senate inquiry on public procurement, Microsoft France hearing, 10 June 2025: “Non, je ne peux pas le garantir, mais […] cela ne s’est encore jamais produit.”
- d. Microsoft Learn, Foundry Models lifecycle policy, July 2026: retirement 18 months after launch (12 for some partner models), not extendable. Google Cloud, model lifecycle, September 2026: long-term Gemini models available at least 12 months.
- e. Measured by Vaiyu, 23 September 2026:
uv pip compilefor vllm 0.30.0, Python 3.12, x86-64 Linux. - f. DORA, Regulation (EU) 2022/2554, applicable 17 January 2025, Article 29. It does not require data localization.
- g. NERC Project 2023-09 white paper, December 2025: current CIP rules are “effectively prohibiting” all but basic cloud use.
- h. Reserve Bank of India circular, 6 April 2018: payment data “stored in a system only in India.”
- i. Stanford HAI, AI Index Report 2026: the top closed model led the top open model by 3.3% on the Arena leaderboard in March 2026.
- j. Thakur, S., Pati, S. et al. Computers in Biology and Medicine 196 (2025). doi:10.1016/j.compbiomed.2025.110615.