analytical-viz: new 'Parallelism rules' tab

New last tab distilling the parallelism-selection discussion into
tables. Framed up-front as heuristics + hard constraints, not a
production procedure — with public counterexamples (DeepSeek-V3 PP16
no TP; vLLM PP-on-non-NVLink) that break the naive 'start TP at 8'
rule.

Ten sections, all table-first:

  1. The core principle (chattiest comm on fastest links)
  2. Parallelism techniques × comm profile (DP/TP/PP/CP/EP with
     Shards, Pattern, Frequency, Per-link volume as degree ↑,
     Hard ceiling)
  3. Dominant problem → first axis to try (11 rows: weight fit,
     activation memory, single-seq KV, many-seq KV, MoE experts,
     TPOT, aggregate throughput, multi-node, expert layer, global
     batch cap)
  4. When to add / when to stop, per axis (DP/TP/PP/CP/EP/FSDP/SP)
  5. Common misconceptions vs reality (9 rows: KV-head hard
     ceiling, always start TP at 8, EP comm falling, attn-DP
     shards KV for free, DP > CP always, PP inside domain
     worthless, memory is only feasibility, FSDP as substrate,
     pipeline bubble = (P-1)/m)
  6. Inference decision procedure (7 ordered steps)
  7. Inference cases → recommended shape (A–E: fits on 1 GPU /
     within NVLink / multi-node / extreme context / many moderate)
  8. Training baseline recipe (6 ordered steps)
  9. Sanity checks for any candidate config (7 checks)
 10. One-paragraph rule + honest caveat block

All content is st.markdown + st.dataframe — no new pure functions.
Smoke test asserts the tab title + `with tab_prules:` block are both
present in app.py.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
This commit is contained in:
2026-07-30 09:11:08 -07:00
parent 5ce7a5b30b
commit 55653cbb3e
2 changed files with 388 additions and 1 deletions
+377 -1
View File
@@ -555,13 +555,14 @@ if _warnings:
# Auto Suggest is renamed "Parallelism" since it only varies parallelism
# knobs (hardware is held fixed at the sidebar values).
(tab_layout, tab_auto, tab_memory, tab_stages, tab_compare, tab_hw,
tab_roofline, tab_planning) = st.tabs([
tab_roofline, tab_planning, tab_prules) = st.tabs([
"Physical layout",
"Auto Suggest Parallelism",
"Memory breakdown", "Per-stage latency",
"Save & compare", "Auto Hardware",
"Chip roofline & B*",
"Capacity planning",
"Parallelism rules",
])
@@ -3122,3 +3123,378 @@ with tab_planning:
]
st.dataframe(pd.DataFrame(_playbook_rows), width='stretch',
hide_index=True)
# ── TAB 9: Parallelism rules ─────────────────────────────────────
with tab_prules:
st.subheader("Parallelism rules — heuristics, constraints, and honest caveats")
st.info(
"**This is directional guidance, not a production procedure.** "
"Real hyperscalers combine the rules below with measured "
"collectives, workload traces, and per-architecture search "
"(GSPMD, Alpa, hand-tuned Megatron). Public disclosures "
"(DeepSeek-V3 training: PP16 + EP64 + no TP; vLLM's "
"PP-over-TP recommendation on non-NVLink GPUs) show that no "
"universal 'start TP at 8' rule survives contact with real "
"workloads. Use these tables as a starting point, then benchmark."
)
# ── Section 1: core principle ─────────────────────────────────
st.markdown("### 1. The core principle")
st.markdown(
"> **Rank the techniques by how much they communicate per unit "
"of compute, and assign the chattiest ones to the fastest links.**\n\n"
"Memory capacity is the feasibility filter (you fit or you don't). "
"Once you fit, every remaining choice is about minimizing "
"*exposed* communication (comm that couldn't be overlapped with "
"compute)."
)
st.divider()
# ── Section 2: Parallelism technique comparison ──────────────
st.markdown("### 2. Parallelism techniques — comm profile at a glance")
_p_compare = [
{"Technique": "DP (Data Parallelism)",
"Shards": "Nothing (or optimizer state w/ ZeRO)",
"Comm pattern": "AllReduce (gradients)",
"Frequency": "1× per training step",
"Per-link vol as degree ↑": "Flat, but overlappable",
"Hard ceiling": "Critical batch size"},
{"Technique": "TP (Tensor Parallelism)",
"Shards": "Weights + activations + KV by head",
"Comm pattern": "AllReduce",
"Frequency": "2× per layer",
"Per-link vol as degree ↑": "Flat (compute shrinks, comm doesn't)",
"Hard ceiling": "≤ NVLink domain size; efficient ≤ n_kv_heads"},
{"Technique": "PP (Pipeline Parallelism)",
"Shards": "Weights by layer + KV by layer",
"Comm pattern": "Point-to-point (activation hand-off)",
"Frequency": "Per stage boundary",
"Per-link vol as degree ↑": "Negligible per link",
"Hard ceiling": "n_layers"},
{"Technique": "CP (Context Parallelism)",
"Shards": "Activations + KV by position",
"Comm pattern": "Ring P2P (per hop)",
"Frequency": "Per layer, overlappable",
"Per-link vol as degree ↑": "Falls per hop (KV shard shrinks)",
"Hard ceiling": "seq_len / block_size"},
{"Technique": "EP (Expert Parallelism, MoE only)",
"Shards": "Expert weights",
"Comm pattern": "All-to-all (dispatch + combine)",
"Frequency": "2× per MoE layer",
"Per-link vol as degree ↑": "Falls per peer (total flat)",
"Hard ceiling": "n_experts"},
]
st.dataframe(pd.DataFrame(_p_compare), width='stretch',
hide_index=True)
st.caption(
"The TP row is the one that decides most configs. Its "
"AllReduce **volume doesn't shrink** as you add ranks — only "
"compute does. That's why TP has a practical ceiling (~8) "
"even without hitting a hard n_kv_heads limit."
)
st.divider()
# ── Section 3: dominant problem → first axis to try ──────────
st.markdown("### 3. What's hurting → which parallelism to reach for")
_p_problem = [
{"Dominant problem": "Weight memory doesn't fit",
"First to try": "TP (within a fast domain)",
"Second lever": "PP across nodes; FSDP/ZeRO in training",
"Why": "TP shards weights + activations; PP shards by depth"},
{"Dominant problem": "Optimizer / gradient memory (training)",
"First to try": "FSDP / ZeRO-2 or ZeRO-3",
"Second lever": "Increase DP degree",
"Why": "ZeRO shards optimizer state across DP ranks"},
{"Dominant problem": "Activation memory (long sequences)",
"First to try": "CP / sequence parallelism",
"Second lever": "Activation checkpointing",
"Why": "Activations scale with seq_len; CP shards positions"},
{"Dominant problem": "KV cache of ONE long sequence",
"First to try": "CP",
"Second lever": "TP over KV heads; PP by layer",
"Why": "DP can't split one sequence — CP shards positions"},
{"Dominant problem": "KV capacity for MANY sequences",
"First to try": "DP replicas + request scheduling",
"Second lever": "Attention-DP within EP groups",
"Why": "Independent sequences shard by DP for free"},
{"Dominant problem": "MoE expert weight memory",
"First to try": "EP",
"Second lever": "Redundant copies of hot experts",
"Why": "EP places different experts on different ranks"},
{"Dominant problem": "Single-request latency (TPOT)",
"First to try": "TP (up to ~8)",
"Second lever": "CP for very long prefill",
"Why": "TP divides weight-fetch time linearly"},
{"Dominant problem": "Aggregate throughput (many users)",
"First to try": "Outer DP",
"Second lever": "Higher B per replica",
"Why": "DP scales throughput linearly at zero comm cost"},
{"Dominant problem": "Model spans multiple nodes",
"First to try": "TP within node, PP across nodes",
"Second lever": "Consider CP if seq_len is also large",
"Why": "Slow scale-out links favor PP's point-to-point"},
{"Dominant problem": "Expert layer dominates (MoE)",
"First to try": "EP",
"Second lever": "Attention-DP + EP combined",
"Why": "vLLM supports DP-attn + EP-experts topology"},
{"Dominant problem": "Global batch already at critical size (training)",
"First to try": "TP × PP × CP × EP",
"Second lever": "Reduce DP further",
"Why": "Additional DP past critical batch wastes samples"},
]
st.dataframe(pd.DataFrame(_p_problem), width='stretch',
hide_index=True)
st.divider()
# ── Section 4: per-axis select/stop criteria ─────────────────
st.markdown("### 4. When to add / when to stop — per axis")
_p_criteria = [
{"Axis": "DP",
"Add when": "More independent batches/requests are available",
"Stop when": "Global batch exceeds critical batch, or replicas "
"can't fit"},
{"Axis": "TP",
"Add when": "Layers/weights need sharding or TPOT needs improvement",
"Stop when": "Collectives dominate step time OR local GEMMs "
"become too small OR TP > n_kv_heads (efficiency)"},
{"Axis": "PP",
"Add when": "Model depth must be distributed, or TP would cross "
"slow links",
"Stop when": "Bubble fraction > ~10% OR stage imbalance grows"},
{"Axis": "CP",
"Add when": "A single sequence is long enough to need position "
"sharding",
"Stop when": "Ring hops can't be hidden by compute OR local "
"sequence blocks are too small for efficient kernels"},
{"Axis": "EP",
"Add when": "MoE model + expert weights should be distributed",
"Stop when": "Expert GEMMs become too small OR routing "
"imbalance grows OR all-to-all latency dominates"},
{"Axis": "FSDP / ZeRO",
"Add when": "Training-state memory is the dominant problem",
"Stop when": "Parameter all-gather / reduce-scatter starts "
"dominating step time"},
{"Axis": "Sequence parallelism",
"Add when": "TP is used during training (typically bundled)",
"Stop when": "Rarely disabled alone — coupled to TP support"},
]
st.dataframe(pd.DataFrame(_p_criteria), width='stretch',
hide_index=True)
st.divider()
# ── Section 5: naive rules vs nuanced reality ────────────────
st.markdown("### 5. Common misconceptions vs reality")
_p_miscon = [
{"Naive rule": "Memory is only a feasibility filter",
"Reality": "Memory headroom is also a continuous performance "
"variable. More free HBM = larger B, more prefix "
"cache, higher throughput — even after weights fit."},
{"Naive rule": "TP has a hard ceiling at n_kv_heads",
"Reality": "vLLM (and others) support KV-head replication when "
"TP > n_kv_heads. It's an efficiency preference, "
"not a correctness ceiling."},
{"Naive rule": "Always start TP at 8",
"Reality": "Start with the smallest TP that fits memory + "
"meets latency. DeepSeek-V3 inference uses TP=4 "
"(and TP=1 for some MLPs) explicitly to limit "
"collective overhead."},
{"Naive rule": "EP comm falls as EP grows",
"Reality": "Per-peer traffic falls, total outgoing traffic is "
"flat (~N·k·H). Large EP hurts via startup overhead, "
"small local GEMMs, and imbalance."},
{"Naive rule": "Attention-DP shards KV 'for free'",
"Reality": "Only across DIFFERENT sequences. One 1M-token "
"sequence still needs CP/TP/PP to shard — DP can't "
"split it. Attention weights are still replicated "
"across DP ranks."},
{"Naive rule": "DP > CP always if you have many requests",
"Reality": "For MANY short requests, yes. For ONE long "
"sequence with strict prefill SLO, CP wins because "
"DP can't help. Different dimensions solve "
"different problems."},
{"Naive rule": "Avoid PP inside the domain — it's worthless",
"Reality": "vLLM specifically recommends PP over TP on GPUs "
"without NVLink. DeepSeek-V3 training used PP=16 "
"and no TP. PP shines when TP collectives cost more "
"than PP bubbles."},
{"Naive rule": "FSDP/ZeRO-3 is THE training substrate",
"Reality": "DeepSeek-V3 used ZeRO-1 + PP16 + EP64 — no ZeRO-3. "
"FSDP is one option; the right choice depends on "
"the topology, batch size, and model shape."},
{"Naive rule": "Pipeline bubble = (P1)/m",
"Reality": "Actual bubble fraction = (P1)/(m + P1). The "
"(P1)/m form is an approximation valid only when "
"m ≫ P. Interleaved / 1F1B / zero-bubble schedules "
"have their own formulas."},
]
st.dataframe(pd.DataFrame(_p_miscon), width='stretch',
hide_index=True)
st.divider()
# ── Section 6: inference decision procedure ──────────────────
st.markdown("### 6. Inference decision procedure (in order)")
_p_infer = [
{"Step": 1,
"Action": "Feasibility check",
"Criterion": "Compute weights + KV(B, S_kv). If it fits on "
"one GPU, stop — use plain replication (DP only)."},
{"Step": 2,
"Action": "Set TP from TPOT budget",
"Criterion": "TP is the latency knob — divides weight-fetch "
"time. Prefer TP=48 for latency-oriented; TP=12 "
"for throughput-oriented. Try to keep TP ≤ n_kv_heads."},
{"Step": 3,
"Action": "Fill remaining domain with EP (if MoE)",
"Criterion": "EP is the cost knob for MoE. Cap at n_experts. "
"Overlap dispatch/combine with a second micro-batch."},
{"Step": 4,
"Action": "Attention-DP over EP groups",
"Criterion": "Different attention replicas handle different "
"sequence subsets → KV shards by sequence for "
"free, no extra collective."},
{"Step": 5,
"Action": "CP only if attention-DP can't help",
"Criterion": "Trigger: one very long sequence or per-request "
"KV exceeds one GPU. For many short concurrent "
"requests, DP strictly wins."},
{"Step": 6,
"Action": "PP across racks if the model still won't fit",
"Criterion": "Only across slow scale-out boundaries. Inside a "
"fast NVLink domain, PP is usually inferior to TP."},
{"Step": 7,
"Action": "Replicate the whole unit as outer DP",
"Criterion": "The throughput axis. Cap by traffic + critical "
"batch (training) or by aggregate request "
"arrival rate (inference)."},
]
st.dataframe(pd.DataFrame(_p_infer), width='stretch',
hide_index=True)
st.divider()
# ── Section 7: concrete inference cases ──────────────────────
st.markdown("### 7. Inference cases → recommended shape")
_p_cases = [
{"Case": "A. Fits on one GPU (with KV headroom)",
"Recommended": "TP=1, PP=1, CP=1; scale via outer DP",
"Notes": "Benchmark TP=2/4 if TPOT is tight"},
{"Case": "B. Fits within one NVLink node",
"Recommended": "TP = 2/4/8 (smallest that fits + meets SLO)",
"Notes": "Consider (TP=2, PP=2) for very deep models"},
{"Case": "C. Spans multiple nodes",
"Recommended": "TP = GPUs per fast domain, PP = "
"number of nodes; outer DP for throughput",
"Notes": "vLLM's standard multi-node recipe"},
{"Case": "D. One request with extreme context (≫ L*)",
"Recommended": "TP × CP × PP (DP can't split one sequence)",
"Notes": "Long-context tier; heavy CP dominates"},
{"Case": "E. Many moderate-length requests",
"Recommended": "Minimum TP × PP per replica; maximize DP = "
"total_GPUs / (TP × PP)",
"Notes": "Throughput-optimized; the mainline serving case"},
]
st.dataframe(pd.DataFrame(_p_cases), width='stretch',
hide_index=True)
st.divider()
# ── Section 8: training baseline recipe ──────────────────────
st.markdown("### 8. Training baseline recipe (in order)")
_p_train = [
{"Step": 1,
"Action": "Start with DP (or sharded DP / FSDP)",
"Rationale": "Simplest; shards optimizer state; scales at "
"throughput cost only"},
{"Step": 2,
"Action": "Add TP if large layers don't fit or GEMMs stay "
"efficient after sharding",
"Rationale": "TP divides weight memory + activation memory"},
{"Step": 3,
"Action": "Add CP when sequence activations dominate",
"Rationale": "Activation memory ~ seq_len × d; CP shards "
"positions"},
{"Step": 4,
"Action": "Add PP when depth still doesn't fit or scaling "
"across nodes",
"Rationale": "Depth sharding + slow-link tolerance"},
{"Step": 5,
"Action": "Add EP for MoE expert weights",
"Rationale": "Distributes expert weights; preserves tokens "
"per expert"},
{"Step": 6,
"Action": "Use remaining devices as outer DP",
"Rationale": "Global batch = DP × micro-batch × grad-accum; "
"cap by critical batch"},
]
st.dataframe(pd.DataFrame(_p_train), width='stretch',
hide_index=True)
st.caption(
"Total GPUs = **DP × TP × PP × CP × EP**. Global batch = "
"**DP × micro-batch × grad-accum**. The critical batch size "
"(NOT bandwidth) is what forces model parallelism at very "
"large scale — past that, additional DP wastes samples."
)
st.divider()
# ── Section 9: sanity checks for any config ──────────────────
st.markdown("### 9. Sanity checks for any candidate config")
_p_sanity = [
{"Check": "TP stays inside one NVLink / scale-up domain?",
"Pass criterion": "Yes",
"If not": "TP AllReduce crosses slow links → step time balloons"},
{"Check": "TP ≤ n_kv_heads?",
"Pass criterion": "Yes (or KV replication is acceptable)",
"If not": "Wastes KV BW to replication; usually not worth it"},
{"Check": "Any KV or weight state duplicated inside the "
"scale-up domain?",
"Pass criterion": "No",
"If not": "You're wasting aggregate HBM bandwidth"},
{"Check": "Exposed (non-overlapped) comm fraction?",
"Pass criterion": "≤ ~10% of step time",
"If not": "Something is misplaced — swap axes across links"},
{"Check": "Pipeline bubble fraction (P1)/(m+P1)?",
"Pass criterion": "≤ ~10%",
"If not": "Need more microbatches or fewer stages"},
{"Check": "Predicted step time meets TPOT SLO?",
"Pass criterion": "Yes",
"If not": "Reduce B, add DP replicas, or accept slower SLO"},
{"Check": "B sits in the 23× B* sweet spot?",
"Pass criterion": "Yes",
"If not": "Below → underused compute; above → wasted latency"},
]
st.dataframe(pd.DataFrame(_p_sanity), width='stretch',
hide_index=True)
st.divider()
# ── Section 10: honest caveats ────────────────────────────────
st.markdown("### 10. One-paragraph rule + caveats")
st.markdown(
"> **First use the smallest parallelism group that makes the "
"model, training state, activations, and target KV capacity "
"fit. Use TP mainly inside the fastest scale-up domain, PP to "
"divide model depth — especially across slower boundaries — CP "
"to divide an individual long sequence, and EP to distribute "
"MoE experts while preserving enough tokens per expert. Use DP "
"across independent batches or requests for throughput. Then "
"benchmark several legal combinations using real kernel "
"shapes, collective measurements, workload distributions, and "
"end-to-end SLOs; choose based on exposed communication and "
"cost per token, not communication volume alone.**"
)
st.warning(
"**Honest caveat.** The specific configurations frontier labs "
"run are mostly unpublished, and they retune per architecture "
"+ hardware generation. What's stable is the *ordering "
"principle* and the *constraint list* — those follow from the "
"hardware, not from anyone's cleverness. Public counterexamples "
"(DeepSeek-V3, vLLM PP-on-non-NVLink) show that no single "
"recipe fits all workloads."
)
@@ -165,3 +165,14 @@ def test_ttft_tpot_section_wired_on_layout_tab():
pos = src.index('"Full-model latency (TTFT & TPOT)"')
tab_memory_pos = src.index("with tab_memory:")
assert pos < tab_memory_pos, "TTFT section must live inside tab_layout"
def test_parallelism_rules_tab_registered():
"""The 'Parallelism rules' tab is wired into st.tabs and has a
corresponding `with tab_prules:` block. Guards the new tab existing."""
from pathlib import Path
src = (
Path(__file__).parent / "app.py"
).resolve().read_text(encoding="utf-8")
assert src.count('"Parallelism rules"') == 1
assert "with tab_prules:" in src