Required peak replicas
8exact-derivedInteger capacity after headroom.
[1]Portfolio 02 · production serving
Translate a compatible load-test result into minimum, peak, rollout, and failover replicas across four managed endpoint surfaces.
Decision this answers: What does the endpoint cost before traffic arrives, and can its configured replica ceiling meet the measured peak?
Required peak replicas
8exact-derivedInteger capacity after headroom.
[1]Zero-traffic floor
$3066.00exact-derivedMinimum replicas for the whole month.
[1]Rollout surcharge
$134.40exact-derivedTemporary duplicate peak capacity.
[2]Monthly topology cost
$6951.00exact-derivedIdle, active extra, rollout, and failover terms.
[1][2][3][4]The tables below are the accessible source of truth; bar lengths never carry identity alone.
| Candidate | Mode | Region | Capacity evidence | Price | Scaling boundary | Evidence |
|---|---|---|---|---|---|---|
| Vertex AI | Dedicated endpoint | us-central1 | User-measured 32 req/s | Editable rate | Scale-to-zero depends on mode | measured[1] |
| Azure Machine Learning | Managed online endpoint | eastus | User-measured 32 req/s | Editable rate | Deployment minimum applies | measured[2] |
| Amazon SageMaker AI | Real-time endpoint | us-east-1 | User-measured 32 req/s | Editable rate | Serverless is a distinct mode | measured[3] |
| Databricks Model Serving | Serverless serving | us-east-1 | User-measured 32 req/s | Editable DBU rate | Infrastructure included by documented meter | measured[4] |
Use the official surface to confirm native prices, counting rules, limits, and product eligibility.
Official pages price individual meters. This workbench turns one user-measured sustainable rate into topology-aware capacity and keeps idle, peak, rollout, and failover costs separate.
Every result-affecting reference is visible here without JavaScript and is retained in the JSON export.
Vertex AI online predictionmanaged online endpointreal-time inferenceModel Serving