Pack identical GPU pods against accelerator topology, CPU, and RAM; add HA and rollout reserve before checking autoscaler bounds and pricing whole nodes.
Decision this answers: How many whole GPU nodes must be rented, which resource binds, and can the configured autoscaler ceiling schedule the workload?
Use the official surface to confirm native prices, counting rules, limits, and product eligibility.
This offline workbench
Kubernetes documents extended-resource scheduling. This planner makes homogeneous integer bin packing, HA reserve, rollout surge, control-plane fees, and whole-node cost explicit across managed services.
Sources and methodology
Every result-affecting reference is visible here without JavaScript and is retained in the JSON export.
[1] Kubernetes GPU scheduling
Authority
Kubernetes
Native identifier
extended resource nvidia.com/gpu
Unit
integer extended resource
Evidence
exact-derived
As of
2026-08-01
Retrieved
2026-08-01
Freshness
reference
Normalization
GPU requests are treated as indivisible scheduling constraints.