For frontier AI labs

Count the places your weights rest.
Make it one.

Train and serve on every GPU machine you can access and manage. Weights stay in your storage, stream in encrypted and run in memory.

38attack vectors on model weightsSource: RAND, Securing AI Model Weights

  • Weights stay in your storage
  • About 5 minutes to install the hub
  • A code by e-mail — no password

Where your weights rest

  • A GPU host
  • A copy of your weights at rest

66places your weights rest1place: your storage

Illustrative: assumes each host keeps the model on its own disk to load it, plus checkpoints and backups.

How it helps

More GPUs. One place for your weights.

No weights on host disks

TodayEvery host that trains or serves the model keeps a copy on its disk.

With OmniraWeights stream in encrypted and run in memory; servers keep none.

Retired or seized servers hold 0 bytes of your weights.

Every GPU you can reach

TodayCapacity comes in large blocks, negotiated far ahead of need.

With OmniraRun on every GPU machine you can access and manage.

Marketplace capacity is on our roadmap.

You hold the keys

TodayWhoever runs a host can reach what is on its disk.

With OmniraNothing opens without your own key.

Weights in use on memory-encrypting chips: on our roadmap.

How AI labs start

Weights are sovereign.

Start on the GPU machines you already manage. Weights stay in your own storage from the first job.

Explore

The details, when you want them.

How it works for AI labs

Use more GPUs, anywhere

Run jobs across every GPU machine you can access and manage. Marketplace capacity is on our roadmap.

No weights on server disks

Weights live in your storage, stream in encrypted and run in memory. Retired or seized servers hold none.

You hold the keys

Nothing opens without your own key.

Inference closer to users

Serve models from machines near your users, to cut serving costs.

Protecting weights while they are in use, with memory-encrypting chips, is on our roadmap.

The full story: how traceless compute works

The numbers, with sources

Compute is scarce. Weights are the crown jewels.

38

distinct attack vectors on model weights

Mapped in RAND’s framework for securing frontier AI models.

Source: RAND, Securing AI Model Weights

30%

yearly growth in AI server electricity

Projected in the IEA’s base case. AI servers drive almost half of data center demand growth to 2030.

Source: IEA, Energy demand from AI, 2025

Insiders are the top concern

Labs and government name insider threats and extortion as their leading worry.

Source: RAND, Achieving AI Model Weight Security Level 3

Capacity comes deal by deal

GPU capacity is negotiated in large blocks, far ahead of need.

Security, in short
  • Servers keep no job data at rest: code, data and logs stream in encrypted, and the machine keeps nothing.
  • A 90% smaller attack surface: 9 of the 10 attack elements of a standard server are removed or handled. The one that stays is data in use, which a job needs in memory while it runs.
  • 95.5% fewer ways for data and IP to leak: 10 of 11 paths removed, and network traffic encrypted. Memory-encrypting chips, which protect data in use as well, are on our roadmap.

The scorecard, row by row Each row, with the evidence behind it

Built for how AI runs

Distance only touches the edges of an AI job.

AI training is a batch job, so it tolerates latency.

How it works

  • LoadTraining data and the model stream into GPU memory.
  • TrainDays of steps, all in GPU memory.
  • CheckpointProgress saves to storage every so often.

Where distance shows

  • Start-upThe first load takes longer, once.
  • CheckpointsEach save crosses the network.
  • Training stepsNo effect: they run in memory.

Mitigation

  • Load aheadLoad the next batch of training data while the GPUs work.
  • Save in the backgroundTraining keeps running while it saves.
  • Keep GPUs togetherA job’s GPUs sit in one data center; only storage is far away.

Training depends on how much data moves, not on round-trip time. A run that takes days barely notices milliseconds at start and at each save.

AI inference loads once, then runs in memory.

How it works

  • Load onceModel weights load into GPU memory at start.
  • ServeEach request runs in memory: prompt in, answer out.
  • RetrieveSome requests look up documents in storage.

Where distance shows

  • Start-upLoading the model takes longer, once.
  • The user’s tripPeople feel the trip to the compute, not to storage.
  • Each lookupA remote document search adds one round trip.

Mitigation

  • Serve near usersRun chat and live apps in-country.
  • Index in RAMLoad search indexes into memory once.
  • Batch the workGroup requests and lookups.

AI is moving from chat to agents: an agent gets a task and runs for minutes or hours, like a batch job. Providers already price batch inference at half the cost of live requests.

Batch pricing: OpenAI Batch API (24-hour window) and Anthropic Message Batches (most finish within an hour), both at 50% off.

Questions

What people ask first.

Do weights ever touch a host’s disk?

No. Weights stay in your storage, stream in encrypted and run in memory.

Does distance slow training down?

Barely. Training depends on how much data moves, not on round-trip time: distance touches the first load and each save, not the steps that run in GPU memory.

What about data in use?

Weights in use sit in GPU memory only for the job, never on disk. Protecting them even from someone with full control of the machine, with memory-encrypting chips, is on our roadmap.

More GPUs. Fewer places for weights to leak.

Start on the GPU machines you already manage.

  • Weights stay in your storage
  • About 5 minutes to install the hub
  • A code by e-mail — no password