77abc95d78
Two more table sections on the Capacity planning tab: **Section 4 — Practical rules of thumb (by context regime)** - Short (S_kv < L*): pack HBM, B ~ 2·B*, high util, cheap tier - Long (S_kv > L*): fewer users/replica, low B, util drops as 1/(1 + S_kv/L*), pricier - Extreme (S_kv >> L*): dedicated pool, heavy CP, disaggregated prefill, very low util without sparse attention Columns: Regime, Batch strategy, Utilization, Cost/token, Deployment. **Section 5 — Sample deployment templates** Same base model, three different sharding recipes routed to by the API gateway based on request context length: - Config_small : CP=1, TP=8, PP=1 → 8 GPUs, up to 32k, B=64 - Config_medium: CP=4, TP=8, PP=1 → 32 GPUs, up to 128k, B=32 - Config_large : CP=32, TP=8, PP=1 → 256 GPUs, up to 1M, B=4 Columns: Tier, CP, TP, PP, Total GPUs/replica, Max context, Typical B, Best for. Playbook section renumbered to #6. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>