offload

Contents

offload#

torchsim.offload(devices=('cuda',), *, budget_bytes=_DEFAULT_BUDGET_BYTES, lanes=1)[source]#

Run host-resident volumes on CUDA without holding them there.

Inside the block a workload whose voxels are on the host still runs on devices, in chunks sized so that the device buffers stay within budget_bytes in total. The result comes back on the host.

The budget buys memory with time, and steeply once the chunks get small: a chunk both fills the device less well and pays its own launch and transfer latency, so quartering the budget costs far more than a quarter of the throughput. Give it as much as the card can spare.

lanes is how many chunks may be in flight per device. A lane owns a stream and its own buffers, so a second one lets a chunk’s transfer overlap another chunk’s arithmetic – but it takes its share out of the same budget, and every chunk gets narrower as a result. Narrower chunks lose more than the overlap recovers except when the volume was barely splitting to begin with, so one lane is the default and raising it is worth measuring before believing. A machine that overlaps transfers better than this one would move that balance.

Whatever lanes is set to, it changes only throughput: what makes a volume larger than the card runnable at all is the chunking.

Raises:

ValueError – if devices is empty, names a device that is not CUDA,: or if budget_bytes or lanes is not positive.