\section{Supporting Agentic Workloads} \label{sec:agentic} The fused GQA kernel of \S\ref{sec:gqa} solved the efficient execution of a \emph{single} logical attention stream: one query stream against one KV cache, tiled and merged with an online softmax. Agentic inference introduces a different execution model, in which one request dynamically \emph{forks} into several cooperating reasoning branches and later \emph{joins} them. The purpose of this section is to show that this multi-stream model needs no new attention algorithm---the expensive hardware machinery built for \S\ref{sec:gqa} is reused unchanged---and to work out the resulting design across three implementation layers. The central claim, stated once here and defended throughout, is: \begin{quote} \emph{The proposed agentic execution does not change the semantics of transformer attention; it changes only how independent query rows are scheduled over a shared KV cache.} \end{quote} \noindent Attention stays identical; only the scheduling differs. What follows is a \emph{design} built on the measured \S\ref{sec:gqa} kernel; it is not yet a KernBench measurement, and points that go beyond the implemented path are flagged as such. \subsection{Motivation: from one attention stream to many} \label{sec:agentic-why} Agentic execution has two recurring shapes: a \emph{loop} (an agent that repeatedly generates, calls a tool, and continues) and a \emph{fan-out / fan-in} (a parent agent forks several specialised sub-agents and later synthesises their results). The loop is ordinary autoregressive decoding and needs nothing new. The fan-out is the interesting case, because the branches are \emph{structurally related} rather than independent: \[ \begin{aligned} \text{context}_i \;=\;& \underbrace{\text{shared prefix}}_{\text{common}} \\ &+\; \underbrace{\text{private role}_i + \text{private suffix}_i}_{\text{branch-specific}} . \end{aligned} \] Ordinary batching groups unrelated requests that merely arrive together; agentic fan-out groups requests that share an identical prefix KV cache and differ only in a short private suffix. That structure creates two levers that unrelated-request batching cannot pull: \begin{enumerate} \item \textbf{Shared-prefix KV reuse.} All branches attend to the same prefix KV, so it is read (and stored) once, not once per branch. \item \textbf{Query-row batching.} The new query rows of many sub-agents read the \emph{same} shared KV, so they can be concatenated into one taller GEMM. \end{enumerate} \noindent Both levers land directly on the \S\ref{sec:gqa} design, which already multicasts a (small) query against a stationary KV cache and merges the result hierarchically. Fan-out simply makes the query taller, and fan-in adds a join step; neither touches the attention math. \subsection{Architecture: three execution layers} \label{sec:agentic-arch} Before the mechanics, we fix the layering, because the recurring reader question is ``whose job is this?''. Agentic execution divides cleanly along the request-flow layers already used in this report: the \textbf{agentic framework} decides \emph{what} to run and how to combine it, the \textbf{runtime} decides \emph{how} to map it onto the hardware without copying shared state, and the \textbf{kernel} does the actual math and reduction on the PEs. \[ \text{Framework} \;\longrightarrow\; \text{Runtime} \;\longrightarrow\; \text{Kernel} \] \noindent Keeping the split this way preserves the topology-agnostic runtime boundary: the framework never sees the SIP/CUBE/PE hierarchy, and the kernel never sees the agent tree. The remainder of the section walks the fan-out $\rightarrow$ runtime $\rightarrow$ kernel $\rightarrow$ fan-in path through exactly these three layers. Table~\ref{tab:agentic-levels} states each layer's responsibilities; the kernel row is unchanged from \S\ref{sec:gqa}. \begin{table*}[t] \centering \caption{Division of responsibilities for agentic attention. Only the framework and runtime rows are new work; the kernel is the \S\ref{sec:gqa} fused GQA kernel, unchanged in its math and reduction.} \label{tab:agentic-levels} \small \begin{tabular}{@{}p{0.16\textwidth}p{0.40\textwidth}p{0.36\textwidth}@{}} \toprule \textbf{Level} & \textbf{Responsibilities} & \textbf{Explicitly not its job} \\ \midrule \textbf{Agentic framework} & Manage the agent tree: fork sub-agents and assign roles; identify the shared prefix; mark which branches are co-schedulable (same shared prefix). Fan-in: enforce schema-constrained results, deduplicate/rank findings, hold evidence behind pointers, insert hierarchical reducers, and drive the main agent's final synthesis. & No topology, routing, KV placement, or scheduling. Sees agents and text, not CUBEs or PEs. \\ \addlinespace \textbf{Runtime} & Fork the logical context by page table (shared-prefix pages $+$ private suffix pages) with \emph{no copy} of shared KV. Concatenate co-scheduled sub-agent query rows into $Q_{\text{cmb}}$. Select the execution policy (Stationary-KV vs.\ Distributed-Q, \S\ref{sec:agentic-choice}) from the cost model. Launch the fused GQA kernel (composite command $+$ PE\_IPCQ) and hand the batched work and policy to the simulation engine. & No attention math and no per-hop routing (delegated to the engine and policy); no agent-level semantics. \\ \addlinespace \textbf{Kernel} & Multicast $Q_{\text{cmb}}$ to all PEs; per-PE local attention over the stationary KV shard via composite-command $Q\!\cdot\!K^{\top}$ and $P\!\cdot\!V$; produce per-row online-softmax state; merge that state hierarchically over PE\_IPCQ by matching row index (row-parallel, not serialised). & No knowledge of agents, page tables, or policy selection. Identical to \S\ref{sec:gqa}; a taller $Q$ is the only difference. \\ \bottomrule \end{tabular} \end{table*} \subsection{Stationary-KV Execution Policy} \label{sec:agentic-policyA} The core attention operation is $Q\!\cdot\!K^{\top}$, and under fan-out multiple sub-agents contribute different query rows while reading a common $K$. Rather than issuing $Q_A\!\cdot\!K_{\text{shared}}$, $Q_B\!\cdot\!K_{\text{shared}}$, \dots\ as separate small GEMMs, the runtime concatenates the rows, \[ Q_{\text{cmb}} = \begin{bmatrix} Q_A \\ Q_B \\ \vdots \end{bmatrix}, \qquad Q_{\text{cmb}}\!\cdot\!K_{\text{shared}}, \] and issues one taller GEMM. The arithmetic is unchanged---each row is still independent---but the $M$ dimension grows. For a representative fan-out of $8$ sub-agents at $20$ new tokens each, $M = 8\times20 = 160$; with a per-PE KV shard of $256$ sequence positions this is a $160\times d$ by $d\times256$ local product---an $M{=}160$, $N{=}256$ shape that sits well above the per-command issue overhead the composite command (\S\ref{sec:gemm}) is built to amortise. Fan-out is therefore especially valuable during \emph{prefill} of the private suffixes, where the combined $M$ is large; during decode the query stays short and the workload remains, as \S\ref{sec:gqa} found, KV-bandwidth bound. The Stationary-KV policy maps onto the \S\ref{sec:gqa} placement essentially unchanged: logically replicated $Q$ over a stationary, sequence-sharded KV cache. One KV head spans four CUBEs and 32 PEs, with each PE owning a different KV \emph{sequence} shard $K_p[S_p,d]$, $V_p[S_p,d_v]$. The combined $Q$ is multicast to all of them; each PE forms a complete local score $Q_{\text{cmb}}\!\cdot\!K_p^{\top}$ over its own shard, and the per-row online-softmax state is reduced hierarchically---8-PE merge inside each CUBE, then a 4-CUBE merge---using the same PE\_IPCQ merge primitive from \S\ref{sec:allreduce}. \paragraph{The reduction algorithm does not change.} This is the point to emphasise. Agentic batching does \emph{not} require a new reduction algorithm: the hierarchical online-softmax reduction of \S\ref{sec:gqa} remains exactly as-is, and only the query batch becomes larger. The reduction is always \emph{same row index across sequence shards}, never \emph{different query rows against each other}, so $M$ batched rows do \emph{not} become $M$ serialised communication rounds; the $(m,\ell,O)$ row states are exchanged as vectors/tiles and merged in parallel. Because the merge is row-indexed, a taller $Q$ widens each payload but adds no rounds. This is what makes agentic support a reuse of \S\ref{sec:gqa} rather than a redesign. \paragraph{Logical vs.\ physical $Q$ replication.} ``Replicated $Q$'' is a \emph{logical} statement. Because the shard axis is the KV \emph{sequence} dimension $S$, every PE must form the full score $Q\!\cdot\!K_p^{\top}$ against its local $K_p$ and therefore needs $Q$'s \emph{entire} hidden dimension $d$; what is partitioned across PEs is $K$/$V$ along $S$, never $Q$ along its columns. Splitting $Q$ (and $K$) on the hidden dimension would instead make each PE's product \emph{partial} and force a pre-softmax hidden-dimension reduction ($QK^{\top}=\sum_i Q_iK_i^{\top}$)---that is tensor-/head-parallel attention, a different structure from the sequence-parallel one assumed here, and one that cannot coexist with using the PE axis for sequence shards. Logical replication also does not mean 32 physical copies: $Q$ can be multicast once into a CUBE-local shared buffer (shared SRAM) that all PEs in the CUBE read, and a large $Q$ can further be \emph{row}-tiled in time ($Q[0{:}16,:],\,Q[16{:}32,:],\dots$)---row tiling splits the $M$ dimension, not the hidden-dimension columns. In short: the Stationary-KV policy uses logically replicated $Q$ across sequence-parallel PEs while $K$ and $V$ are partitioned along the sequence dimension; $Q$ may be temporally row-tiled or physically shared through multicast buffers, but it is not partitioned along the hidden-dimension columns. \paragraph{Replication is not $32\times$ the compute work (attention FLOPs).} Multicasting $Q$ to 32 PEs does not multiply attention FLOPs, because each PE computes against a different KV sequence shard rather than the same one. Let the KV sequence length be $S$; with sequence parallelism over 32 PEs, each PE owns $S/32$ positions. The score GEMM $Q[M,d]\!\cdot\!K^{\top}[d,S]$ costs $\propto M\,d\,S$, so each PE performs $M\,d\,(S/32)$ and the 32 shards sum to \[ 32 \cdot M\,d\,\tfrac{S}{32} \;=\; M\,d\,S, \] identical to attention over one undivided sequence. Replication therefore changes $Q$ distribution, reduction traffic, buffering, and scheduling---not the total attention FLOPs. \subsection{Distributed-Q Execution Policy (alternative)} \label{sec:agentic-policyB} The natural alternative is to partition (distribute) the query rows across PE groups (\emph{the Distributed-Q policy}) rather than replicate them. It is not automatically better. Because the 32 PEs already shard the KV \emph{sequence}, every query row must still attend to \emph{all} shards; partitioning $Q$ across PE groups therefore forces each group to reach every KV shard, which requires one of: regrouping KV shards per $Q$ group, replicating KV across groups, or reading remote KV through symmetric memory. Each of these adds memory-system complexity that the Stationary-KV policy avoids entirely. Time-multiplexing the same PEs over $Q$ groups is the fourth option, but that is temporal tiling---already available under the Stationary-KV policy as row tiling---not true spatial $Q$ partitioning. The Distributed-Q policy is thus a proposed adaptive extension, not the baseline, and is only worth its complexity when $Q$ grows large enough that multicast and reduction traffic dominate remote/regrouped-KV cost. \subsection{Why the Stationary-KV policy is the baseline} \label{sec:agentic-choice} The choice reduces to a cost comparison the runtime can estimate. The Stationary-KV policy pays for $Q$ multicast and a hierarchical reduction; the Distributed-Q policy pays for moving or replicating KV plus extra scheduling: \[ T_{\text{SK}} = T_{Q\text{-mcast}} + T_{\text{local GEMM}} + T_{\text{hier.\ reduce}}, \] \[ \begin{aligned} T_{\text{DQ}} = {}& T_{Q\text{-part}} + T_{\text{remote/repl.\ KV}} + T_{\text{local GEMM}} \\ &+ T_{\text{grp.\ reduce}} + T_{\text{remap}}. \end{aligned} \] For the assumed mapping (one KV head $=$ 4 CUBEs $=$ 32 PEs), the KV cache is large and stationary while $Q$ is comparatively small, so $T_{\text{SK}}$ is the lower cost and the Stationary-KV policy is the recommended baseline. A future runtime can compute both estimates per launch---from the current agent count, $Q$ size, KV-head mapping, and interconnect state---and switch to the Distributed-Q policy only in the regime where a very large $Q$ batch makes multicast and reduction traffic outweigh the cost of remote or regrouped KV. Until that regime is measured, the Distributed-Q policy remains a designed, not-yet-implemented option. \subsection{Fan-in: joining sub-agent branches} \label{sec:agentic-fanin} Fan-out is only half of the pattern; after it, each branch produces a private continuation and the parent must synthesise them. We treat fan-in as a runtime/framework design problem with a clear optimization ladder. \paragraph{Problem.} The branch KV caches \emph{cannot} be concatenated. Each token's K/V depends on its full causal history, so stacking several branches' private KV does not form the cache of any single valid sequence. Join must therefore happen at the token/text level, and its cost is dominated by the number of join-input tokens, because the new main-agent KV grows in proportion to them. \paragraph{Baseline.} A naive full-text gather concatenates every branch's raw output: $8$ agents $\times\ 1000$ tokens $=\ 8000$ tokens pushed through every layer during continuation prefill---inflating prefill work, KV allocation, and later decode-time KV reads. This is the cost the ladder below drives down. \paragraph{Optimization 1 --- schema-constrained results.} Constrain each sub-agent to emit a compact structured result (claim, confidence, evidence handles) instead of free text, cutting join input by an order of magnitude ($8\times1000 \to 8\times50$). \paragraph{Optimization 2 --- deduplication and ranking.} Overlapping findings across branches are merged and ranked in the framework before they reach the main context. This is preprocessing, not reasoning, and shrinks the input further without involving the main model. \paragraph{Optimization 3 --- pointer-based evidence.} Detailed evidence stays outside the main context behind handles, materialised only when the main agent actually requests it, so the effective input is summaries plus only the evidence truly used. \paragraph{Optimization 4 --- hierarchical reducers.} For wide fan-out, intermediate reducer agents summarise groups of branches, shrinking the final join prompt and organising fan-in traffic hierarchically---mirroring how the hierarchical online-softmax merge organises attention reduction. \paragraph{Future --- latent-state join.} A more aggressive step replaces text outputs with a few learned latent tokens, further cutting join prefill and KV growth. This requires training the main model to consume latent tokens and aligning branch representations; it is a model--system co-design direction, not a drop-in runtime optimization, and is out of scope here. \medskip \noindent Throughout, the shared prefix KV is reused by page-table reference and only the join suffix is new, so shortening join input is the primary lever on continuation-prefill cost, KV growth, and later decode-time KV reads. Taken together, \S\ref{sec:agentic-policyA}--\ref{sec:agentic-fanin} show why this workload is a natural extension of the 1H design rather than a new one: the decisive hardware levers of this report---cheap composite issue and a fast on-device reduction path---are exactly what make agentic fan-out efficient, and no part of the attention math is altered to get there. Quantifying the fan-out speedup and the fan-in join savings on KernBench is left as measured 2H work.