Skip to content
Vaiyu Solutions, home

For air-gapped & sovereign deployments

Your data stays put. The model comes to it.

Most AI tools assume they can reach the internet. On an air-gapped network they can’t, and a residency rule may forbid it. We build AI that runs entirely on infrastructure you control, and we’ll tell you which workloads don’t need that. We learned how in hospitals.

15+

years operationalizing AI, prototype to production

$9M+1

federally funded AI R&D led

up to50%2

training cost cut for clients

713

sites in one federated learning study

up to90%4

inference latency removed

The exposure

Every default in the AI stack assumes your data can leave.

  • 01

    Your staff already use AI tools you never approved.

    In IBM’s 2025 study, one in five breached organizations reported a breach involving shadow AI, and heavy users paid about US$670,000 more per breach. A ban fails if staff have nothing approved to use.a

  • 02

    Storing data in a region doesn’t decide who can reach it.

    Under the CLOUD Act, a US provider must hand over data it controls, wherever it sits. In 2025, Microsoft France told the French Senate under oath that it couldn’t rule this out for French public-sector data, though it had never happened.bc

  • 03

    The model you validated retires on someone else’s schedule.

    Microsoft retires a hosted model version 18 months after launch, with no extensions. Google’s long-term Gemini models get at least 12 months. Each retirement means validating your workflows again. Hold the weights and you choose when the model changes.d

  • 04

    An air gap breaks every install script.

    Installing the vLLM inference server pulls in 196 Python packages, 27 of them from NVIDIA. Behind an air gap, someone mirrors, scans, approves, and carries across each one, at every update.e

  • 05

    The rules keep asking where the data is.

    DORA makes EU financial firms manage their reliance on cloud providers. NERC rules keep US bulk-power systems largely off the public cloud. India requires payment data to be stored in India.fgh

The spectrum

Sovereignty is a spectrum.An air gap is the far end of it.

Most teams we work with need two: one for everyday work, and a stricter one for the workloads that carry real risk.

Ordered by isolation, least to most

  1. Tier 1 of 4

    Hosted API

    A frontier model under an enterprise contract.

    Guards against
    Retention, and training on your data.
    Trade-off
    Their jurisdiction and their retirement schedule.

    Right for

    Most everyday work

  2. Tier 2 of 4

    Sovereign cloud

    Dedicated capacity in your region, sometimes run by a local operator.

    Guards against
    Data leaving the region.
    Trade-off
    Still networked, and you don’t hold the weights.

    Right for

    In-region mandates

  3. Tier 3 of 4

    Self-hosted

    Open-weight models on your hardware, with outbound traffic blocked.

    Guards against
    Data leaving, forced upgrades, and foreign legal process.
    Trade-off
    You run the GPUs, and open models still trail the best closed ones.i

    Right for

    Sensitive data, validated workflows

  4. Tier 4 of 4

    Air-gapped

    No network path at all. Updates cross on approved media.

    Guards against
    Everything above, plus attacks over the network.
    Trade-off
    Every update becomes a procedure.

    Right for

    Plant floors, isolated labs, trade secrets

We will tell you which tier each workload needs, including when the answer is the cloud.

What transfers

We learned this in hospitals, where the data can’t leave either.

  • Federated training across 71 sites on six continents3One model trained across countries or partners, each dataset kept where the law requires.
  • Inference tuned for limited clinical hardware4jA useful model on GPUs you can actually buy and run on-site.
  • Reproducible builds and maintained conda-forge packagesEvery dependency pinned and mirrored, so the stack rebuilds offline.

What we do for sovereign deployments

Six services, run inside your perimeter.

We don’t sell GPUs, models, or licenses. We build on what you own and hand it to your team.

01

Readiness & architecture sprint

Every workload placed on a tier, with hardware sized and a build-or-buy call.

02

Data engineering inside the perimeter

Pipelines and search indexes built where the data lives, never copied out.

03

Model selection & evaluation

Open-weight models tuned on your data and tested on cases you write.

04

Air-gapped deployment & MLOps

Mirrored registries, a written transfer procedure, and monitoring that sends nothing out.

05

Governance & documentation

Data-flow maps and validation files, written for your assessor.

06

Fractional AI leadership

A senior AI lead a few days a month, testing every vendor’s “sovereign” claim.

How we engage

Start with one workload, on one tier.

Each engagement has its own scope, fee, and exit criteria.

Readiness sprint

Two to four weeks at a fixed fee, covering every workload you’re weighing.

A costed plan your security office has already read.

Private assistant pilot

One team and one document set, on an open-weight model running on your hardware.

A visible win, and an offline stack you keep.

Stalled-pilot rescue

For a pilot stuck at security review. We close the findings and rebuild what can’t pass.

The pilot back in front of the reviewers.

We build and document the controls. Your security office and your assessor sign off on them.

Bring the model to the data.

Start with a readiness sprint or a one-workload pilot on your own hardware.

Write to ussupport [at] vaiyu [dot] solutions

Sources & attribution

  1. 1. Led by our founder across NIH/NCI-funded programs at the University of Pennsylvania and Indiana University.
  2. 2. Vaiyu client engagements: pre-training optimization with model accuracy maintained or improved.
  3. 3. Pati, S. et al. “Federated learning enables big data for rare cancer boundary detection.” Nature Communications 13 (2022). doi:10.1038/s41467-022-33407-5; 71 sites across 6 continents, the largest real-world federated learning study to date.
  4. 4. Founder track record at Indiana University: inference latency reduced by up to 70%, compute requirements by 10–50%, in clinical research environments.
  5. a. IBM and Ponemon Institute, Cost of a Data Breach Report 2025: 600 breached organizations; 20% reported a shadow-AI breach; average cost USD 4.74M with high shadow-AI use, USD 4.07M with low or none.
  6. b. CLOUD Act, 18 U.S.C. § 2713: data in a provider’s “possession, custody, or control,” wherever located.
  7. c. French Senate inquiry on public procurement, Microsoft France hearing, 10 June 2025: “Non, je ne peux pas le garantir, mais […] cela ne s’est encore jamais produit.”
  8. d. Microsoft Learn, Foundry Models lifecycle policy, July 2026: retirement 18 months after launch (12 for some partner models), not extendable. Google Cloud, model lifecycle, September 2026: long-term Gemini models available at least 12 months.
  9. e. Measured by Vaiyu, 23 September 2026: uv pip compile for vllm 0.30.0, Python 3.12, x86-64 Linux.
  10. f. DORA, Regulation (EU) 2022/2554, applicable 17 January 2025, Article 29. It does not require data localization.
  11. g. NERC Project 2023-09 white paper, December 2025: current CIP rules are “effectively prohibiting” all but basic cloud use.
  12. h. Reserve Bank of India circular, 6 April 2018: payment data “stored in a system only in India.”
  13. i. Stanford HAI, AI Index Report 2026: the top closed model led the top open model by 3.3% on the Arena leaderboard in March 2026.
  14. j. Thakur, S., Pati, S. et al. Computers in Biology and Medicine 196 (2025). doi:10.1016/j.compbiomed.2025.110615.