execution#

torchsim.execution(target='auto', *, stream=None, budget_bytes=None, lanes=1, reserve_bytes=_RESERVE_BYTES)[source]#

Choose where work runs inside the block.

target is "auto" to decide per call, "cpu" to insist on the host, or a device or list of devices to insist on those. Deciding weighs the problem against what each card has free right now: work too small to repay a launch stays on the CPU, work that fits goes across in one piece, and work that does not is streamed through in chunks.

stream overrules that last step – False demands the whole volume be resident and lets it fail if it will not fit, True streams even when it would have fit. budget_bytes caps what streaming may hold, and defaults to what the devices report free less reserve_bytes. lanes is passed to offload() for streamed work and rarely wants changing.

Outside a block, work runs wherever its tensors already are.

Raises:

ValueError – for an empty device list, a non-CUDA device, or a: non-positive budget_bytes or lanes.

Examples using torchsim.execution#

Basic Usage

Basic Usage

Dictionary matching

Dictionary matching

MP2RAGE lookup table

MP2RAGE lookup table

T2 mapping by nonlinear least squares

T2 mapping by nonlinear least squares

PERK: kernel ridge regression

PERK: kernel ridge regression