\section{Hardware Performance-Spec Search for GQA} \label{sec:hwspec} \emph{Work in progress --- this section states the study's intent and method; experimental data, figures, and conclusions are not yet available and will be added in a later revision.} The studies so far fixed the modeled hardware (\S\ref{sec:hw}) and varied the \emph{software}: the composite command, the reduction path, and the KV placement. This section inverts the question. Given the fused GQA kernel of \S\ref{sec:gqa} as the target workload, \emph{which hardware specification best serves it}, and where does spending more silicon stop paying off? The discussion of \S\ref{sec:discussion} already gives a strong prior---attention is data-movement bound, so raw MAC throughput is not the binding constraint---but that conclusion was drawn at one operating point. A systematic sweep is needed to turn it into a defensible performance-spec recommendation across context lengths, agent counts, and sharding regimes. \subsection{Sweep axes} \label{sec:hwspec-axes} Four hardware knobs are varied, spanning the compute side and both levels of the interconnect so that the compute/communication balance can be located rather than assumed. The KernBench cost model already exposes each as a parameter, so the sweep reuses the same deterministic engine as the rest of the report; no production change is required. \begin{table}[h] \centering \caption{Hardware knobs varied in the performance-spec search. Ranges and step counts are to be finalised with the experiment.} \label{tab:hwspec-axes} \small \begin{tabular}{@{}p{0.30\textwidth}p{0.14\textwidth}@{}} \toprule \textbf{Knob} & \textbf{Axis} \\ \midrule GEMM throughput (TFLOPS) & compute \\ MATH-engine ALU count & compute \\ CUBE-to-CUBE (die-to-die) bandwidth & interconnect \\ SIP-to-SIP (card-to-card) bandwidth & interconnect \\ \bottomrule \end{tabular} \end{table} The first two axes scale the on-PE compute engines; the latter two scale the two inter-device links that carry the KV reduction and cross-device softmax merge. Sweeping them jointly---rather than one at a time---is what exposes the interactions: for example, whether added die-to-die bandwidth only helps once card-to-card bandwidth is also raised, or whether GEMM throughput is genuinely inert for attention once issue overhead is removed. \subsection{Method and target output} \label{sec:hwspec-method} The intended procedure is a joint sweep over the four axes, running the fused GQA kernel at representative decode and prefill configurations (including the agentic fan-out shapes of \S\ref{sec:agentic}), and reading latency and per-engine busy time from the engine's completion timestamps and op log---the same measurement path used in \S\ref{sec:gqa}. The target deliverables are: (i) the sensitivity of GQA latency to each knob in isolation, (ii) the Pareto frontier of latency against a simple area/cost proxy, and (iii) a recommended balanced specification---the knob combination past which further investment does not move GQA latency. These results are the natural quantitative counterpart to the qualitative ranking in \S\ref{sec:discussion}, and completing them is a 2H objective.