Skip to main content
A cold start includes finding capacity, loading the image, and initializing your app. Here are a few ways to reduce it.

Cache Dependencies and Model Files

Install dependencies in your image and cache model weights in a Volume. Point your model library’s cache path at the mounted directory.

Initialize Once Per Worker

Use on_start for initialization shared by requests handled by that worker. Access its return value through context.on_start_value.
Multiple workers initialize separately, so account for their combined memory use. See startup loaders.

Keep Capacity Warm

Increase keep_warm_seconds or set min_containers to keep containers running between requests. Warm containers are billed; traffic beyond their capacity can still cause cold starts.

Checkpoint Restore

Supported workloads can capture initialized state with checkpoint_enabled=True:
Check lifecycle logs for capture and restore results. Support depends on the runtime and workload; reconnect external services after restoring. For containers, use beam deploy --help-all to find checkpoint and readiness options. For sandboxes, see snapshots.

Measure the Result

Compare cold and warm requests in the dashboard’s task timings. Use logs and metrics to find which startup step takes the longest.