Portfolio 08 · warm versus elastic

Serverless GPU break-even calculator

Translate measured service time and concurrency into warm replicas, then price provider granularity, minimum duration, model load, idle retention, redundancy, and provisioned capacity.

Decision this answers: At this deterministic traffic shape, does scale-to-zero save money without violating the entered cold-start tolerance?

Scenario

Numeric scenario inputs

Zero can cost zero only when no minimum or fixed meter remains.

measured[1]

Deterministic peak used for warm capacity.

measured[1]

Compatible measured seconds per request.

measured[1]

Measured sustainable concurrent requests.

measured[1]

Headroom kept below saturation.

model estimate[1]

Cloud Run L4 no-zonal-redundancy list snapshot.

published list[1]

The complete default scenario is server rendered. Interactions stay in this browser.

Decision summary

Required warm replicas

2exact-derived

Integer capacity from measured service time.

[1]

Serverless bill

$112.15exact-derived

Rounded billable GPU seconds plus load and idle retention.

[1]

Provisioned bill

$981.30exact-derived

Warm replicas held for the full month.

[1]

Scenario cold exposure

12.3%model estimate

Deterministic gap-model estimate, not a percentile guarantee.

[1]
Direct-labeled result profile

The tables below are the accessible source of truth; bar lengths never carry identity alone.

Serverless GPU capability ledger

Serverless GPU capability ledger
CandidateGPU modeGranularityMinimumScale to zeroCapacity evidenceEvidence
Cloud Run L4No zonal redundancy100 msMode-specificSupported with zero minimumUser measuredpublished list[1]
Azure Container Apps GPUServerless GPUProvider-specificProvider-specificConfiguration-specificRequiredunknown[2]
SageMaker ServerlessUnsupported GPU in this rowN/AN/AN/AUnpriced and lastunavailable[3]

Official surface versus this site

Cloud Run pricing

Use the official surface to confirm native prices, counting rules, limits, and product eligibility.

This offline workbench

Cloud Run publishes per-second GPU meters and billing rules. This page adds measured request capacity, model load, traffic gaps, cold exposure, and a provisioned break-even comparison.

Sources and methodology

Every result-affecting reference is visible here without JavaScript and is retained in the JSON export.

[1] Cloud Run pricing
Authority
Google Cloud
Native identifier
7754-699E-0EBF L4 no zonal redundancy
Region
us-central1
Unit
USD/GPU-second
Evidence
published list
Price kind
list
As of
2026-08-01
Effective from
2026-08-01
Retrieved
2026-08-01
Freshness
fresh
Normalization
The documented $0.0001867/s L4 meter and 100 ms rounding are applied; measured capacity remains user-owned.

Open the authoritative source

[2] Azure Container Apps serverless GPU
Authority
Microsoft
Native identifier
serverless GPU
Unit
GPU-second
Evidence
unknown
As of
2026-08-01
Retrieved
2026-08-01
Freshness
reference
Normalization
Capability remains explicit without assigning an unverified price or performance value.

Open the authoritative source

[3] SageMaker Serverless Inference
Authority
Amazon Web Services
Native identifier
Serverless Inference
Unit
memory-second
Evidence
unavailable
As of
2026-08-01
Retrieved
2026-08-01
Freshness
reference
Normalization
Unsupported GPU/serverless combinations remain visible and cannot win.

Open the authoritative source

Limits and non-claims

  • The opening capacity is backed by a measurement: true.
  • Cold exposure is a deterministic scenario output, not a p95 latency guarantee.
  • Availability, quota, model quality, and provider capacity are not inferred.