Skip to main content

Running Tasks on GPU

Pass gpu to a Python workload or --gpu to a container deployment. Install your framework and its compatible CUDA dependencies in the runtime image.

Availability

beam machine list reports live serverless availability and on-demand inventory. GPU model names supported by the SDK are not a promise that every model is currently available serverless. Review offers for price, provider, region, and capacity, then use compute pools for reserved hardware.

GPU selection

A priority list lets the scheduler consider alternative types that your workload supports:
Ensure the model fits the smallest GPU in the list and that its libraries support each architecture. gpu_count requests multiple GPUs on a suitable machine. It does not distribute an application automatically; configure your framework’s multi-GPU execution separately. Availability and workspace limits still apply.

Regions and pools

Use a reserved offer’s provider and region selectors when placement matters. Set the workload’s pool to route it to that capacity. Use beam machine list --help and beam pool --help for the available filters and subcommands.