Required replicas
9exact-derivedBase measured capacity plus failure and deployment reserve.
[1]Portfolio 15 · measured serving
Select only compatible measured operating points, enforce TTFT, ITL, end-to-end, request, and token units independently, then add failure and rolling-deployment capacity.
Decision this answers: Which measured point satisfies every selected SLO, and how many replicas remain after one failure and a rolling deployment?
Required replicas
9exact-derivedBase measured capacity plus failure and deployment reserve.
[1]Fleet request capacity
64.80measuredSelected measured requests/second times active replicas.
[1]Measured token throughput
33174.00measuredRetains tokens/second rather than converting to requests.
[1]Arrival/capacity ratio
64.8%exact-derivedA bounded consistency indicator, not a tail-latency percentile.
[3]The tables below are the accessible source of truth; bar lengths never carry identity alone.
| Candidate | Workload | TTFT | ITL | End-to-end | Request throughput | Token throughput | Evidence |
|---|---|---|---|---|---|---|---|
| Point A | 2k-in-512-out | 0.45 s | 0.035 s/token | 18.4 s | 5.8 req/s | 2,970 tok/s | measured[1] |
| Point B · selected | 2k-in-512-out | 0.72 s | 0.028 s/token | 15.1 s | 7.2 req/s | 3,686 tok/s | measured[1] |
| Incompatible point | 8k-in-2k-out | 0.30 s | 0.020 s/token | 40 s | 10 req/s | 20,000 tok/s | unavailable[1] |
Use the official surface to confirm native prices, counting rules, limits, and product eligibility.
Benchmark tools measure serving behavior. This page validates workload compatibility, filters points against explicit SLOs, and converts the surviving measured throughput into a resilience-aware replica plan.
Every result-affecting reference is visible here without JavaScript and is retained in the JSON export.
GenAI-Perf metricsserving metricsL = lambda W