Required warm replicas
2exact-derivedInteger capacity from measured service time.
[1]Portfolio 08 · warm versus elastic
Translate measured service time and concurrency into warm replicas, then price provider granularity, minimum duration, model load, idle retention, redundancy, and provisioned capacity.
Decision this answers: At this deterministic traffic shape, does scale-to-zero save money without violating the entered cold-start tolerance?
Required warm replicas
2exact-derivedInteger capacity from measured service time.
[1]Serverless bill
$112.15exact-derivedRounded billable GPU seconds plus load and idle retention.
[1]Provisioned bill
$981.30exact-derivedWarm replicas held for the full month.
[1]Scenario cold exposure
12.3%model estimateDeterministic gap-model estimate, not a percentile guarantee.
[1]The tables below are the accessible source of truth; bar lengths never carry identity alone.
| Candidate | GPU mode | Granularity | Minimum | Scale to zero | Capacity evidence | Evidence |
|---|---|---|---|---|---|---|
| Cloud Run L4 | No zonal redundancy | 100 ms | Mode-specific | Supported with zero minimum | User measured | published list[1] |
| Azure Container Apps GPU | Serverless GPU | Provider-specific | Provider-specific | Configuration-specific | Required | unknown[2] |
| SageMaker Serverless | Unsupported GPU in this row | N/A | N/A | N/A | Unpriced and last | unavailable[3] |
Use the official surface to confirm native prices, counting rules, limits, and product eligibility.
Cloud Run publishes per-second GPU meters and billing rules. This page adds measured request capacity, model load, traffic gaps, cold exposure, and a provisioned break-even comparison.
Every result-affecting reference is visible here without JavaScript and is retained in the JSON export.
7754-699E-0EBF L4 no zonal redundancyserverless GPUServerless Inference