Skip to main content

Configuring CPU and Memory

In addition to choosing a GPU, you can choose the amount of CPU and Memory to allocate:
GPU graphics cards have VRAM and run on servers with RAM.

RAM vs. VRAM

VRAM holds model weights, activations, and GPU working memory. Host RAM is separate and is used by the CPU, model loading, and application processes. Increasing the memory argument changes host RAM, not GPU VRAM. Model fit depends on precision, quantization, context length, batch size, and the serving engine. Measure peak usage for your workload rather than selecting hardware by parameter count alone. Disk capacity is separate again: a streaming download can require substantial disk space without keeping the entire file in RAM. Numeric SDK memory values are MiB; strings such as "2Gi" and "512Mi" are also accepted. CPU values represent cores, including supported fractional values. Configuration through the MCP update_config tool uses millicores for runtime.cpu (1000 for one CPU).

Monitoring Resource Usage

In the web dashboard, you can monitor the amount of CPU, Memory, and GPU memory used for your tasks. On a deployment, click the Metrics button.
On this page, you can see the resource usage over time. The graph will also show the periods when your resource usage exceeded the resource limits set on your app:
See observability for logs and resource timeseries, and GPU acceleration for live hardware availability.