Files
ywkang edb30326ce report(1H): add Agentic Workloads (§6) and HW-Spec Search (§7) sections
- §6 Supporting Agentic Workloads: how the fused GQA design extends to
  agentic fan-out/fan-in; three-layer split (framework/runtime/kernel);
  Stationary-KV vs Distributed-Q execution policies.
- §7 Hardware Performance-Spec Search for GQA: WIP stub (sweep intent
  over GEMM TFLOPS, MATH-engine ALUs, CUBE↔CUBE and SIP↔SIP BW).
- Renumber Discussion/Conclusion/Future-Work to 08/09/10; update
  main.tex input order and toc.md.
- Add Agentic_Runtime_Architecture.md design note; rebuild main.pdf.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-22 10:20:27 -07:00

68 lines
3.1 KiB
TeX

\section{Hardware Performance-Spec Search for GQA}
\label{sec:hwspec}
\emph{Work in progress --- this section states the study's intent and
method; experimental data, figures, and conclusions are not yet
available and will be added in a later revision.}
The studies so far fixed the modeled hardware (\S\ref{sec:hw}) and
varied the \emph{software}: the composite command, the reduction path, and
the KV placement. This section inverts the question. Given the fused GQA
kernel of \S\ref{sec:gqa} as the target workload, \emph{which hardware
specification best serves it}, and where does spending more silicon stop
paying off? The discussion of \S\ref{sec:discussion} already gives a
strong prior---attention is data-movement bound, so raw MAC throughput is
not the binding constraint---but that conclusion was drawn at one operating
point. A systematic sweep is needed to turn it into a defensible
performance-spec recommendation across context lengths, agent counts, and
sharding regimes.
\subsection{Sweep axes}
\label{sec:hwspec-axes}
Four hardware knobs are varied, spanning the compute side and both levels
of the interconnect so that the compute/communication balance can be
located rather than assumed. The KernBench cost model already exposes each
as a parameter, so the sweep reuses the same deterministic engine as the
rest of the report; no production change is required.
\begin{table}[h]
\centering
\caption{Hardware knobs varied in the performance-spec search. Ranges and
step counts are to be finalised with the experiment.}
\label{tab:hwspec-axes}
\small
\begin{tabular}{@{}p{0.30\textwidth}p{0.14\textwidth}@{}}
\toprule
\textbf{Knob} & \textbf{Axis} \\
\midrule
GEMM throughput (TFLOPS) & compute \\
MATH-engine ALU count & compute \\
CUBE-to-CUBE (die-to-die) bandwidth & interconnect \\
SIP-to-SIP (card-to-card) bandwidth & interconnect \\
\bottomrule
\end{tabular}
\end{table}
The first two axes scale the on-PE compute engines; the latter two scale
the two inter-device links that carry the KV reduction and cross-device
softmax merge. Sweeping them jointly---rather than one at a time---is what
exposes the interactions: for example, whether added die-to-die bandwidth
only helps once card-to-card bandwidth is also raised, or whether GEMM
throughput is genuinely inert for attention once issue overhead is removed.
\subsection{Method and target output}
\label{sec:hwspec-method}
The intended procedure is a joint sweep over the four axes, running the
fused GQA kernel at representative decode and prefill configurations
(including the agentic fan-out shapes of \S\ref{sec:agentic}), and reading
latency and per-engine busy time from the engine's completion timestamps
and op log---the same measurement path used in \S\ref{sec:gqa}. The target
deliverables are: (i) the sensitivity of GQA latency to each knob in
isolation, (ii) the Pareto frontier of latency against a simple
area/cost proxy, and (iii) a recommended balanced specification---the knob
combination past which further investment does not move GQA latency. These
results are the natural quantitative counterpart to the qualitative
ranking in \S\ref{sec:discussion}, and completing them is a 2H objective.