Portfolio 02 · production serving

Managed ML endpoint cost calculator

Translate a compatible load-test result into minimum, peak, rollout, and failover replicas across four managed endpoint surfaces.

Decision this answers: What does the endpoint cost before traffic arrives, and can its configured replica ceiling meet the measured peak?

Scenario

Numeric scenario inputs

Peak of the representative traffic schedule.

measured[1]

Compatible measured throughput; zero suppresses the recommendation.

measured[3]

Multiplier applied before integer rounding.

model estimate[2]

The zero-traffic floor.

measured[1]

A lower value than required produces a capacity warning.

measured[1]

Editable native compute or DBU-equivalent rate.

negotiated[4]

The complete default scenario is server rendered. Interactions stay in this browser.

Decision summary

Required peak replicas

8exact-derived

Integer capacity after headroom.

[1]

Zero-traffic floor

$3066.00exact-derived

Minimum replicas for the whole month.

[1]

Rollout surcharge

$134.40exact-derived

Temporary duplicate peak capacity.

[2]

Monthly topology cost

$6951.00exact-derived

Idle, active extra, rollout, and failover terms.

[1][2][3][4]
Direct-labeled result profile

The tables below are the accessible source of truth; bar lengths never carry identity alone.

Provider capability and measurement matrix

Provider capability and measurement matrix
CandidateModeRegionCapacity evidencePriceScaling boundaryEvidence
Vertex AIDedicated endpointus-central1User-measured 32 req/sEditable rateScale-to-zero depends on modemeasured[1]
Azure Machine LearningManaged online endpointeastusUser-measured 32 req/sEditable rateDeployment minimum appliesmeasured[2]
Amazon SageMaker AIReal-time endpointus-east-1User-measured 32 req/sEditable rateServerless is a distinct modemeasured[3]
Databricks Model ServingServerless servingus-east-1User-measured 32 req/sEditable DBU rateInfrastructure included by documented metermeasured[4]

Official surface versus this site

Vertex AI pricing

Use the official surface to confirm native prices, counting rules, limits, and product eligibility.

This offline workbench

Official pages price individual meters. This workbench turns one user-measured sustainable rate into topology-aware capacity and keeps idle, peak, rollout, and failover costs separate.

Sources and methodology

Every result-affecting reference is visible here without JavaScript and is retained in the JSON export.

[1] Vertex AI pricing
Authority
Google Cloud
Native identifier
Vertex AI online prediction
Unit
replica-hour
Evidence
published list
As of
2026-08-01
Retrieved
2026-08-01
Freshness
reference
Normalization
Native endpoint meters and topology are preserved; throughput remains user measured.

Open the authoritative source

[2] Azure ML online endpoints
Authority
Microsoft
Native identifier
managed online endpoint
Unit
deployment replica
Evidence
exact-derived
As of
2026-08-01
Retrieved
2026-08-01
Freshness
reference
Normalization
Topology and deployment behavior are used without inferring capacity.

Open the authoritative source

[3] Amazon SageMaker AI pricing
Authority
Amazon Web Services
Native identifier
real-time inference
Unit
instance-hour
Evidence
negotiated
As of
2026-08-01
Retrieved
2026-08-01
Freshness
reference
Normalization
Serving modes remain distinct and use an editable effective rate.

Open the authoritative source

[4] Databricks serverless DBU consumption
Authority
Databricks
Native identifier
Model Serving
Unit
DBU
Evidence
negotiated
As of
2026-08-01
Retrieved
2026-08-01
Freshness
reference
Normalization
Serverless DBU billing is not given a second infrastructure bill.

Open the authoritative source

Limits and non-claims

  • The opening plan requires 8 replicas and is viable against the configured maximum: true.
  • Capacity requires a compatible measured workload point; instance specifications do not substitute for a load test.
  • Quota, availability, latency guarantees, and negotiated contracts are outside this static plan.