Per-rank model state
97.79 GiBexact-derivedFits usable HBM: false.
[2]Portfolio 09 · run envelope
Check per-rank model-state fit, validate topology factors, keep measured throughput beside a clearly labeled theoretical FLOPs envelope, and price checkpoint replay on whole nodes.
Decision this answers: Does the model state fit the selected topology, and what completion time and spend follow from an actual throughput measurement?
Per-rank model state
97.79 GiBexact-derivedFits usable HBM: false.
[2]Measured runtime
21604.94 hmeasuredTokens divided by compatible measured throughput.
[4]Theoretical envelope
55854.30 hupper boundPeak FLOPS path; never labeled a forecast.
[3]Whole-node run cost
$2145439.51model estimateMeasured runtime plus expected replay.
[4]The tables below are the accessible source of truth; bar lengths never carry identity alone.
| Candidate | Memory driver | Compute driver | Runtime authority | Checkpoint treatment | Billing unit | Evidence |
|---|---|---|---|---|---|---|
| Default 8-rank plan | 70B total parameters | 70B active parameters | 18k measured tokens/s | Expected replay separate | Whole node-hour | measured[2][4] |
| Theoretical comparison | Same exact state | 6 × active params × tokens | Peak FLOPS × MFU × efficiency | Not included | Envelope only | upper bound[3][1] |
Use the official surface to confirm native prices, counting rules, limits, and product eligibility.
The primary paper frames training compute. This workbench adds exact state sharding, usable HBM, measured throughput, topology efficiency, checkpoints, interruption replay, and whole-node billing.
Every result-affecting reference is visible here without JavaScript and is retained in the JSON export.
arXiv:2203.15556torch.distributed.fsdpmodel FLOPs utilizationMLPerf Training