offload#
- torchsim.offload(devices=('cuda',), *, budget_bytes=_DEFAULT_BUDGET_BYTES, lanes=1)[source]#
Run host-resident volumes on CUDA without holding them there.
Inside the block a workload whose voxels are on the host still runs on
devices, in chunks sized so that the device buffers stay withinbudget_bytesin total. The result comes back on the host.The budget buys memory with time, and steeply once the chunks get small: a chunk both fills the device less well and pays its own launch and transfer latency, so quartering the budget costs far more than a quarter of the throughput. Give it as much as the card can spare.
lanesis how many chunks may be in flight per device. A lane owns a stream and its own buffers, so a second one lets a chunk’s transfer overlap another chunk’s arithmetic – but it takes its share out of the same budget, and every chunk gets narrower as a result. Narrower chunks lose more than the overlap recovers except when the volume was barely splitting to begin with, so one lane is the default and raising it is worth measuring before believing. A machine that overlaps transfers better than this one would move that balance.Whatever
lanesis set to, it changes only throughput: what makes a volume larger than the card runnable at all is the chunking.- Raises:
ValueError – if
devicesis empty, names a device that is not CUDA,: or ifbudget_bytesorlanesis not positive.