Files
mukesh fc4747668e paper(§5): add roofline, capacity planning, parallelism selection subsections
Three new subsections at the top of §5, motivated by the deployment-sizing
questions the analytical visualization tool exposes:

- 5.1 Roofline Analysis — B*, L*, short vs long context batch effect
  (measured: 19× per-token cost reduction at short context, 28% at long)
- 5.2 Capacity Planning — 3-axis sizing (weights / KV / SLO), SLO targets,
  regime rules, deployment templates, per-axis playbook
- 5.3 Parallelism Selection — TP × CP × PP × DP × EP comparison,
  symptom-driven axis selection, add/stop criteria, misconceptions

Restructure:
- 05-gqa.tex trimmed to section header + intro
- Existing 5.4-5.7 content (placement, short/long ctx, composite) moved
  to 05x-fused-kernel.tex
- Summary extracted to 05z-summary.tex, cross-refs updated
- main.tex \input order sets the requested subsection sequence
- lmodern loaded to satisfy microtype font-expansion

Two new figures generated from tests/analytical_visualization/chip_roofline:
- roofline_short_context.png (S_kv=8K, batch wins)
- roofline_long_context.png (S_kv=1M, batch stalls at KV floor)
- Generator: tests/analytical_visualization/_gen_roofline_paper_figs.py

4 tables (regime rules, deployment templates, playbook, parallelism
criteria) use table* so they span both columns without overflow.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-30 15:56:41 -07:00

86 lines
4.5 KiB
TeX

\subsection{Roofline Analysis: Batch and Context Regimes}
\label{sec:roofline}
Before deploying any sharding strategy, the workload's position on the
decode roofline sets what \emph{can} be improved and what \emph{cannot}.
Two quantities matter: the machine's \textbf{critical batch}
$B^\ast = C \cdot b / (2 W)$ (where $C$ is per-PE FLOPs, $W$ is per-PE
HBM bandwidth, $b$ is bytes per parameter), and its \textbf{balance
context} $L^\ast = 2 N / (\mathrm{AI} \cdot \mathrm{kv\_bpt})$ (where
$N$ is active parameters and $\mathrm{AI} = C/W$ is arithmetic
intensity). Above $B^\ast$ the deployment leaves the memory-bound
regime; above $L^\ast$ per-user KV streaming dominates and no amount
of batching hides it. For Llama-3.1-70B on the default AHBM machine
($C \!=\! \SI{8}{TFLOPS}$/PE, $W\!=\! \SI{256}{GB/s}$/PE), we get
$B^\ast \!\approx\! 31$ and $L^\ast \!\approx\! \SI{13.4}{K}$ tokens.
\paragraph{Short-context regime ($S_{kv} \!<\! L^\ast$).}
Figure~\ref{fig:roofline-short} shows one decode step at
$S_{kv}\!=\! \SI{8}{K}$. The step-latency panel (left) makes it
obvious that weight fetch dominates at low $B$: at $B\!=\!1$, one
step spends $\sim\!\SI{535}{ms}$ streaming weights and only
$\SI{28}{ms}$ on per-sequence compute and KV combined. The
cost-per-token panel (right) is where the batch story lives: dividing
step latency by $B$ shrinks the weight term as $1/B$, so per-token
cost drops from $\SI{562}{ms}$ at $B\!=\!1$ to $\SI{29.7}{ms}$ at
$B\!=\!256$ --- a $19\times$ reduction. This is the memory-bound-to-
compute-bound crossover: past $B^\ast$, weight cost is fully
amortized and per-token time asymptotes to the compute + KV floor
($\sim\!\SI{27}{ms}$). \textbf{At short context, batching directly
lowers cost per token.}
\begin{figure*}[h]
\centering
\includegraphics[width=\textwidth]{roofline_short_context.png}
\caption{Roofline decomposition for Llama-3.1-70B decode at
$S_{kv}\!=\! \SI{8}{K}$ (short, well below $L^\ast\!\approx\! \SI{13.4}{K}$).
\textbf{Left}: step latency vs.\ batch $B$; weight fetch is flat
(one HBM sweep per step), compute and KV grow linearly with $B$.
\textbf{Right}: per-token cost ($=$ step $\div B$); the $1/B$
weight-fetch term amortizes rapidly, dropping total per-token time
from \SI{562}{ms} at $B\!=\!1$ to \SI{29.7}{ms} at $B\!=\!256$
--- a $19\times$ throughput gain from batching alone.}
\label{fig:roofline-short}
\end{figure*}
\paragraph{Long-context regime ($S_{kv} \!\gg\! L^\ast$).}
Figure~\ref{fig:roofline-long} shows the same decomposition at
$S_{kv}\!=\! \SI{1}{M}$ ($\sim\!78 \times L^\ast$). The step-latency
panel shows KV fetch has swelled by two orders of magnitude:
per-token KV cost is now $\sim\!\SI{1342}{ms}$, dwarfing both
weight ($\SI{535}{ms}$ at $B\!=\!1$) and compute ($\SI{17}{ms}$).
The cost-per-token panel is the punchline: increasing $B$ still
shrinks the weight term but leaves the giant KV floor untouched,
because \textbf{KV fetch is per-sequence} --- adding another user
adds a full extra copy of their KV read. Total per-token cost falls
only from \SI{1894}{ms} at $B\!=\!1$ to \SI{1361}{ms} at $B\!=\!256$
--- a mere $28\%$ reduction despite $256\times$ the batch.
\textbf{At long context, batching stops paying because per-user KV
streaming dominates.}
\begin{figure*}[h]
\centering
\includegraphics[width=\textwidth]{roofline_long_context.png}
\caption{Same decomposition at $S_{kv}\!=\! \SI{1}{M}$ (long,
$\sim\!78\times L^\ast$). KV fetch has grown to \SI{1342}{ms}/token
and is now the dominant term at every $B$. Batching only shrinks
the weight component; the KV floor is unmovable because each new
user brings a full per-sequence KV read. Total per-token cost falls
just $28\%$ from $B\!=\!1$ to $B\!=\!256$, versus $19\times$ at
short context.}
\label{fig:roofline-long}
\end{figure*}
\paragraph{Implication for deployment.} The two regimes call for
opposite strategies. In the short-context regime, the operator packs
$B$ as high as HBM allows to sit on the compute floor --- this is
where cost-per-token is minimized and hardware utilization is
highest. In the long-context regime, per-user KV is the binding
resource; batching offers little benefit, so the operator instead
shrinks $\mathrm{kv\_bpt}$ (GQA / MQA / MLA, INT4 KV, sparse
attention) and shards the sequence dimension itself (CP), routing
long-context requests to a dedicated pool with more chips per user.
The next two subsections (\S\ref{sec:capacity-planning},
\S\ref{sec:parallelism-selection}) turn these regime observations
into concrete sizing and sharding rules.