analytical-viz: capacity planning — symbol legend + plain-English formulas
Section 1 rewritten so a layman can read it without a glossary:
1. **Symbol legend table** listing every letter that appears in the
three-axis formulas (N, b, N·b, HBM_per_PE, S_kv, kv_bpt,
users_per_replica, TPOT SLO, step_latency, B_at_SLO, n_users,
N_replicas, ⌈ ⌉). Each row: Symbol / Plain English / Value now
for the current sidebar model + chip. For example:
- N = 'Total number of model parameters (attention weights +
FFN weights, across ALL layers)' → 6.98 billion params for
Llama 3 8B
- kv_bpt = 'KV cache bytes per token per user, summed across all
layers (= 2·H_kv·d_head·b·layers)' → 128 KB/token
2. **Three-axis table** widened from (Formula / Meaning / Grows with)
to (In symbols / In plain English / What it says / Value now).
Axis A now shows the actual substituted arithmetic:
⌈ 13.96 GB / 6.0 GB ⌉ = ⌈ 2.33 ⌉ = 3 PEs.
B and C flag 'see calculator' because they depend on user inputs.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
This commit is contained in:
@@ -2751,23 +2751,107 @@ with tab_planning:
|
|||||||
# ── Section 1: three-axis sizing formula + calculator ─────────
|
# ── Section 1: three-axis sizing formula + calculator ─────────
|
||||||
st.markdown("### 1. GPU count = max(A, B, C) × N_replicas")
|
st.markdown("### 1. GPU count = max(A, B, C) × N_replicas")
|
||||||
|
|
||||||
|
# Symbol legend — spell out every letter that appears in the formulas
|
||||||
|
# below with a plain-English meaning + the value for the current
|
||||||
|
# sidebar model + chip. Makes the terse formulas readable at a glance.
|
||||||
|
_n_active_here = total_active_params(model)
|
||||||
|
_hbm_pe_bytes_here = int(_default_machine.pe_hbm_gb * 1e9)
|
||||||
|
_kv_bpt_here = 2 * model.h_kv * model.d_head * model.bytes_per_elem * model.layers
|
||||||
|
_weight_bytes_here = _n_active_here * model.bytes_per_elem
|
||||||
|
|
||||||
|
st.markdown("**Symbol legend** — what each letter in the formulas means")
|
||||||
|
_legend_rows = [
|
||||||
|
{"Symbol": "N",
|
||||||
|
"Plain English": "Total number of model parameters (attention "
|
||||||
|
"weights + FFN weights, across ALL layers)",
|
||||||
|
"Value now": f"{_n_active_here/1e9:.2f} billion params"},
|
||||||
|
{"Symbol": "b",
|
||||||
|
"Plain English": "Bytes per parameter (2 for BF16, 1 for FP8, "
|
||||||
|
"0.5 for INT4)",
|
||||||
|
"Value now": f"{model.bytes_per_elem} bytes (BF16)"},
|
||||||
|
{"Symbol": "N · b",
|
||||||
|
"Plain English": "Total bytes needed to store one full copy of "
|
||||||
|
"the model weights",
|
||||||
|
"Value now": f"{_weight_bytes_here/1e9:.2f} GB"},
|
||||||
|
{"Symbol": "HBM_per_PE",
|
||||||
|
"Plain English": "Amount of high-bandwidth memory on ONE chip "
|
||||||
|
"(GPU / PE)",
|
||||||
|
"Value now": f"{_default_machine.pe_hbm_gb:.1f} GB"},
|
||||||
|
{"Symbol": "S_kv",
|
||||||
|
"Plain English": "Context length in tokens (how much history "
|
||||||
|
"one user's KV cache holds)",
|
||||||
|
"Value now": f"{s_kv:,} tokens (sidebar)"},
|
||||||
|
{"Symbol": "kv_bpt",
|
||||||
|
"Plain English": "KV cache bytes per token per user, summed "
|
||||||
|
"across all layers "
|
||||||
|
"(= 2 · H_kv · d_head · b · layers)",
|
||||||
|
"Value now": f"{_kv_bpt_here:,} B/token "
|
||||||
|
f"({_kv_bpt_here/1024:.0f} KB/token)"},
|
||||||
|
{"Symbol": "users_per_replica",
|
||||||
|
"Plain English": "How many concurrent users one replica serves "
|
||||||
|
"(= per-replica batch size)",
|
||||||
|
"Value now": "(from calculator below)"},
|
||||||
|
{"Symbol": "TPOT SLO",
|
||||||
|
"Plain English": "Latency budget per output token (Time Per "
|
||||||
|
"Output Token — the SLA the operator commits to)",
|
||||||
|
"Value now": "(from calculator below)"},
|
||||||
|
{"Symbol": "step_latency",
|
||||||
|
"Plain English": "Time for one decode forward pass at a given B "
|
||||||
|
"(weight fetch + compute + KV read)",
|
||||||
|
"Value now": "See Chip Roofline tab"},
|
||||||
|
{"Symbol": "B_at_SLO",
|
||||||
|
"Plain English": "Largest batch size that still finishes one "
|
||||||
|
"decode step within the SLO budget",
|
||||||
|
"Value now": "(from calculator below)"},
|
||||||
|
{"Symbol": "n_users",
|
||||||
|
"Plain English": "Total concurrent users you need to serve "
|
||||||
|
"across the whole deployment",
|
||||||
|
"Value now": "(from calculator below)"},
|
||||||
|
{"Symbol": "N_replicas",
|
||||||
|
"Plain English": "Number of complete model copies deployed "
|
||||||
|
"(each is a self-contained instance)",
|
||||||
|
"Value now": "(from calculator below)"},
|
||||||
|
{"Symbol": "⌈ x ⌉",
|
||||||
|
"Plain English": "'Ceiling' — round UP to the next whole "
|
||||||
|
"number (you can't have 3.2 GPUs)",
|
||||||
|
"Value now": "e.g. ⌈ 3.2 ⌉ = 4"},
|
||||||
|
]
|
||||||
|
st.dataframe(pd.DataFrame(_legend_rows), width='stretch',
|
||||||
|
hide_index=True)
|
||||||
|
|
||||||
|
st.markdown("**The three axes** — plain-English formulas")
|
||||||
_axes_rows = [
|
_axes_rows = [
|
||||||
{"Axis": "A. Capacity floor",
|
{"Axis": "A. Capacity floor",
|
||||||
"Formula": "⌈ N·b / HBM_per_PE ⌉",
|
"In symbols": "⌈ N · b / HBM_per_PE ⌉",
|
||||||
"Meaning": ("Min PEs to hold ONE replica's weights alone. "
|
"In plain English": "⌈ (weight bytes) ÷ (HBM per chip) ⌉",
|
||||||
"The bare floor — below this weights don't fit."),
|
"What it says": "You need at least this many chips just to "
|
||||||
"Grows with": "Model size (N)"},
|
"hold ONE copy of the weights. Below this, "
|
||||||
|
"the model literally doesn't fit.",
|
||||||
|
"Value now":
|
||||||
|
f"⌈ {_weight_bytes_here/1e9:.2f} GB / "
|
||||||
|
f"{_default_machine.pe_hbm_gb:.1f} GB ⌉ = "
|
||||||
|
f"⌈ {_weight_bytes_here/_hbm_pe_bytes_here:.2f} ⌉ = "
|
||||||
|
f"**{-(-_weight_bytes_here // _hbm_pe_bytes_here)} PEs**"},
|
||||||
{"Axis": "B. KV headroom",
|
{"Axis": "B. KV headroom",
|
||||||
"Formula": "⌈ (N·b + users_per_replica · S_kv · kv_bpt) / HBM_per_PE ⌉",
|
"In symbols":
|
||||||
"Meaning": ("Min PEs to hold weights + all KV of the users "
|
"⌈ (N·b + users_per_replica · S_kv · kv_bpt) / HBM_per_PE ⌉",
|
||||||
"assigned to this replica."),
|
"In plain English":
|
||||||
"Grows with": "Users × context"},
|
"⌈ (weight bytes + all users' KV cache) ÷ (HBM per chip) ⌉",
|
||||||
|
"What it says": "Weights ALSO need room for every user's KV "
|
||||||
|
"history. As users × context grows, KV eats "
|
||||||
|
"the remaining HBM until you need more chips.",
|
||||||
|
"Value now": "(depends on users + context — see calculator)"},
|
||||||
{"Axis": "C. Throughput SLO",
|
{"Axis": "C. Throughput SLO",
|
||||||
"Formula": ("N_replicas = ⌈ n_users / B_at_SLO ⌉, where "
|
"In symbols":
|
||||||
"B_at_SLO = largest B s.t. step_latency ≤ TPOT SLO"),
|
"N_replicas = ⌈ n_users / B_at_SLO ⌉",
|
||||||
"Meaning": ("Enough replicas so each carries ≤ B_at_SLO users "
|
"In plain English":
|
||||||
"and meets per-token latency SLO."),
|
"Replicas = ⌈ total users ÷ users each replica handles "
|
||||||
"Grows with": "Users, or tighter SLO"},
|
"within the latency budget ⌉",
|
||||||
|
"What it says": "Even if capacity fits, a single replica can "
|
||||||
|
"only serve so many users before per-token "
|
||||||
|
"latency blows the SLO. More users → more "
|
||||||
|
"replicas.",
|
||||||
|
"Value now": "(depends on SLO — see calculator)"},
|
||||||
]
|
]
|
||||||
st.dataframe(pd.DataFrame(_axes_rows), width='stretch',
|
st.dataframe(pd.DataFrame(_axes_rows), width='stretch',
|
||||||
hide_index=True)
|
hide_index=True)
|
||||||
|
|||||||
Reference in New Issue
Block a user