\section{The KernBench Platform} \label{sec:platform} All results in this report are produced on \emph{KernBench}, a system-level, discrete-event simulator for LLM kernels running on SIP-based AI accelerators. This section explains why the platform exists, how it executes a kernel, how it computes latency and how accurate that number is, and the concrete hardware configuration used for every experiment that follows. \subsection{Why KernBench: source-level kernels without a software stack} \label{sec:why} In a production end-to-end (E2E) stack, kernel performance is entangled with every layer above the hardware: the compiler's tiling and scheduling choices, the framework's operator dispatch, the collective library, and the runtime. Good E2E numbers require \emph{all} of those layers to be co-optimized, which makes it hard to answer a narrower but more fundamental question: \emph{given the hardware, how fast can a well-written kernel be, and which hardware features actually make it faster?} KernBench is built to answer exactly that question. Kernels are written and executed at the \emph{source level}---as algorithmic descriptions in a small tile-oriented kernel API---with no dependency on a compiler or any other software-stack layer. The simulator takes the kernel and a hardware topology and reports the latency that the modeled hardware would deliver. This isolation is deliberate: it lets us study algorithm-level optimizations (how to tile a GEMM, how to schedule a collective, how to fuse an attention kernel) and the hardware features that support them, without the confound of compiler maturity or framework overhead. The cost is that KernBench numbers are \emph{not} E2E latencies; they are the achievable-kernel latencies that an ideal software stack would expose. \subsection{Execution model} \label{sec:exec} KernBench is layered along the flow of a request: \begin{itemize} \item The \textbf{runtime API} is host-facing and topology-agnostic---it deploys tensors and launches kernels but knows nothing about routing or interconnect. \item The \textbf{simulation engine} schedules discrete events, routes every request through the modeled graph, and tracks completion via correlation IDs. \item The \textbf{components} are device-side nodes that model hardware behavior: the per-PE blocks (scheduler, DMA, GEMM and vector-math engines, TCM, IPCQ), the NoC routers, the HBM controllers, and the inter-chiplet links. \end{itemize} The topology is compiled once at configuration time into an authoritative graph of components and links; it is never mutated during a run. Within a PE, work is expressed as \emph{composite commands}: a single command carries an ordered pipeline of operations (\textsf{DMA\_READ} $\rightarrow$ \textsf{FETCH} $\rightarrow$ \textsf{GEMM}/\textsf{MATH} $\rightarrow$ \textsf{STORE} $\rightarrow$ \textsf{DMA\_WRITE}) that the PE scheduler tiles and streams. This composite mechanism is the substrate for the GEMM optimization (\S\ref{sec:gemm}) and, combined with on-PE collectives, for fused attention (\S\ref{sec:gqa}). Data and timing are handled in two passes, so that a kernel's numeric results and its latency are computed consistently but independently. \subsection{Latency model and its accuracy} \label{sec:latency} KernBench obeys a set of golden invariants that keep its latency numbers physically meaningful. End-to-end latency is computed \emph{strictly by explicit traversal} over modeled components and links: every routed request incurs latency greater than zero, routing is deterministic, and every valid request flow has explicit connectivity. There are no hidden shortcuts, implicit waits, or magic delays---if a delay exists in the result, it came from a scheduled event on a modeled component or link. Latency accumulates from three kinds of contributions: per-node fixed overheads (each component carries an \texttt{overhead\_ns}), per-link transfer time, and per-service occupancy. The interconnect is modeled at fine granularity. Each directed link serializes traffic through a bandwidth-limited FIFO, so a busy link delays later flits. Payloads are decomposed into fixed-size flits (default \SI{256}{\byte} bursts) that arrive at $\text{prop} + \text{flit\_bytes}/\text{bw}$ intervals, so link bandwidth throttles arrival rate rather than being applied as a lump sum. HBM is modeled with per-pseudo-channel parallelism: a stateless array of channel-availability timestamps with address-based channel selection captures bank-level concurrency. The cost of \emph{issuing} a command is modeled structurally rather than with a per-operation calibration table. The PE control processor charges, per command, \[ d_{\text{cmd}} = \textsf{FIXED} + b_{\text{logical}} \cdot R, \] where $b_{\text{logical}}$ is the command's hardware-logical byte size, \textsf{FIXED} captures the fixed per-command cost (queue-tail update, completion registration) and $R$ captures the per-byte cost of serializing the command descriptor into the scheduler queue. The default anchoring (\textsf{FIXED} $=40$ cycles, $R = 0.0625$ cycles/byte, i.e.\ \SI{16}{\byte\per\cycle}, at \SI{1}{\giga\hertz}) places a typical composite at roughly \SI{43}{\nano\second}, and a hard cap on a composite's descriptor size prevents the model from rewarding arbitrarily large fused commands beyond what real descriptor queues accept. The dispatch cost is enabled and tuned per topology; in the configurations measured here, command issue is not the bottleneck---data movement is---so this term stays small relative to DMA and collective time. How accurate is all this? The model is precise about the things that dominate kernel latency on this class of hardware: link bandwidth occupancy and serialization, HBM channel parallelism, flit-level streaming, and per-component switching overhead. The GEMM study (\S\ref{sec:gemm}) provides a direct check: the measured MAC efficiency tracks the analytic (theoretical) efficiency within roughly \SIrange{10}{20}{\percent} across a wide range of tile counts, with the gap attributable to pipeline fill and DMA effects that the analytic model omits. The known simplifications---idealized arbitration, no thermal or refresh effects, fixed burst granularity---are the price of a deterministic, inspectable model; they bound the absolute accuracy but do not distort the \emph{relative} comparisons (tiling A vs.\ B, topology X vs.\ Y) that this report is built on. \subsection{Modeled hardware configuration} \label{sec:hw} Table~\ref{tab:hw} summarizes the hardware configuration used for every experiment in this report. It is read directly from the simulator's topology description; per-experiment workload parameters (matrix shapes, collective sizes, sequence lengths) are stated in their respective sections rather than here. \begin{table}[t] \centering \caption{Modeled hardware configuration (shared by all experiments).} \label{tab:hw} \small \begin{tabular}{@{}ll@{}} \toprule \textbf{Parameter} & \textbf{Value} \\ \midrule \multicolumn{2}{@{}l}{\emph{Hierarchy}} \\ SIPs & 2 (1D ring) \\ CUBEs per SIP & 16 ($4\times4$ mesh) \\ PEs per CUBE & 8 (4 corners $\times$ 2) \\ PEs total & 256 \\ \midrule \multicolumn{2}{@{}l}{\emph{Processing element (PE)}} \\ GEMM engine peak & \SI{8}{\tera\flop\per\second} (f16) \\ TCM (on-PE) & \SI{16}{\mega\byte}, \SI{512}{\giga\byte\per\second} R/W \\ \quad kernel scratch & \SI{1}{\mega\byte} \\ DMA engines & 1 read + 1 write \\ CPU / scheduler overhead & \SI{2}{\nano\second} / \SI{1}{\nano\second} \\ \midrule \multicolumn{2}{@{}l}{\emph{Memory (per CUBE)}} \\ HBM capacity & \SI{48}{\giga\byte} (8 slices) \\ HBM aggregate BW & \SI{1024}{\giga\byte\per\second} \\ HBM pseudo-channels & 64 (8 per PE), \SI{32}{\giga\byte\per\second} each \\ SRAM (shared) & \SI{32}{\mega\byte}, \SI{128}{\giga\byte\per\second} link \\ HBM burst & \SI{256}{\byte} \\ \midrule \multicolumn{2}{@{}l}{\emph{Interconnect}} \\ Intra-CUBE NoC link & \SI{256}{\giga\byte\per\second}, \SI{2}{\nano\second}/router \\ Inter-CUBE (UCIe PHY) & \SI{512}{\giga\byte\per\second}, \SI{8}{\nano\second}, XY routing \\ Inter-SIP (PCIe) & \SI{768}{\giga\byte\per\second} per endpoint \\ \midrule \multicolumn{2}{@{}l}{\emph{Command-issue cost model (defaults)}} \\ FIXED per command & 40 cycles \\ per-byte rate $R$ & 0.0625 cycles/byte (\SI{16}{\byte\per\cycle}) \\ composite size cap & \SI{1024}{\byte} \\ \bottomrule \end{tabular} \end{table}