Files
kernbench2/docs/report/1H-codesign-paper/sections/02-platform.tex
T
ywkang 7f437a20bd paper(platform): edit-pass through §2 KernBench Platform done
- Reordered §2 subsections to follow the SIP → CUBE → PE → graph
  flow: Why KernBench → Device and execution model → Latency model
  → Modeled hardware configuration. Readers now meet the device
  hierarchy before the graph abstraction that re-uses it.
- §2.2 Device and execution model: starts with the SIP/CUBE/PE
  hierarchy paragraphs (each anchoring fig:sip-arch, fig:cube-arch,
  fig:pe-arch); then the runtime-API/sim-engine/components bullet
  list; then the atomic-vs-composite command distinction (corrects
  the prior over-narrow framing that read every PE command as
  composite -- atomic single-engine commands exist too, and PE_CPU
  itself runs control-plane work directly).
- §2.3 Latency model: opens with the four-contribution decomposition
  (per-node overhead, per-edge transmission, drain, queuing delay)
  and the latency_model schematic; retains existing The hardware as
  a graph / From graph to DES / Latency contributions / Congestion
  / Control-plane cost model / Accuracy paragraphs. Accuracy
  paragraph now closes on KernBench's sufficiency for *relative*
  HW/SW design trade-offs given analytic + external-simulator
  agreement.
- New figures and assets:
  - figures/sip_architecture.pdf  (SIP-level graph view)
  - figures/cube_architecture.pdf (CUBE-level zoom-in)
  - figures/latency_model.png      (conceptual latency-model
                                   schematic with per-node /
                                   per-edge / drain / queuing-delay
                                   colour coding)
  - figures/pe_architecture.png    (carried over)
- Source-of-truth generator for the latency schematic:
  scripts/paper/paper_latency_model_diagram.py (a report-only
  harness under scripts/paper/ per the /paper isolation rule).
- main.tex preamble: \usepackage{tikz} added (kept from prior
  sequence-diagram draft -- harmless now that the latency model is
  a PNG; left in to keep paragraph numbering stable).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-06-15 17:17:19 -07:00

320 lines
16 KiB
TeX

\section{The KernBench Platform}
\label{sec:platform}
All results in this report are produced on \emph{KernBench}, a
system-level discrete-event simulator for LLM kernels running on AHBM.
This section explains why the platform exists, the device and
execution model it presents to a kernel writer, how its latency model
turns that execution into a number, and the concrete hardware
configuration used for every experiment that follows.
\subsection{Why KernBench: source-level kernels without a software stack}
\label{sec:why}
In a production end-to-end (E2E) stack, kernel performance is entangled
with every layer above the hardware: the compiler's tiling and
scheduling choices, the framework's operator dispatch, the collective
library, and the runtime. Good E2E numbers require \emph{all} of those
layers to be co-optimized, which makes it hard to answer a narrower but
more fundamental question: \emph{given the hardware, how fast can a
well-written kernel be, and which hardware features actually make it
faster?}
KernBench is built to answer exactly that question. Kernels are written
and executed at the \emph{source level}---as algorithmic descriptions
in a small tile-oriented kernel API---with no dependency on a compiler
or any other software-stack layer. The simulator takes the kernel and a
hardware topology and reports the latency the modeled hardware would
deliver. This isolation is deliberate: it lets us study algorithm-level
optimizations (how to tile a GEMM, how to schedule a collective, how to
fuse an attention kernel) and the hardware features that support them,
without the confound of compiler maturity or framework overhead. The
cost is that KernBench numbers are \emph{not} E2E latencies; they are
the achievable-kernel latencies an ideal software stack would expose.
\subsection{Device and execution model}
\label{sec:exec}
KernBench models AHBM as a hierarchy of SIPs, CUBEs, and processing
elements (PEs). At the system level, multiple SIPs are connected
through inter-package links, while each SIP contains a collection of
CUBEs joined by an on-package interconnect
(Fig.~\ref{fig:sip-arch}).
\begin{figure}[t]
\centering
\includegraphics[width=\linewidth]{sip_architecture.pdf}
\caption{Modeled hardware graph at the \emph{SIP} level (one
example configuration; specific parameters in \S\ref{sec:hw} /
Table~\ref{tab:hw}). The SIP holds a $4{\times}4$ mesh of CUBEs and
an IO chiplet, with each line a directed link labelled by its
physical distance and bandwidth.}
\label{fig:sip-arch}
\end{figure}
Each CUBE contains eight PEs, shared SRAM, HBM controllers, an
\textsf{M\_CPU} control processor, and an intra-CUBE router mesh
(Fig.~\ref{fig:cube-arch}). Together these components form the
execution substrate for all kernels evaluated in this report.
\begin{figure}[t]
\centering
\includegraphics[width=\linewidth]{cube_architecture.pdf}
\caption{Modeled hardware graph at the \emph{CUBE} level---a
zoom-in on one CUBE node from Fig.~\ref{fig:sip-arch}. The CUBE
holds 8 PEs (each with its own HBM channels), the \textsf{M\_CPU}
control processor, a shared SRAM, an intra-CUBE NoC router mesh, and
UCIe links to neighbouring CUBEs.}
\label{fig:cube-arch}
\end{figure}
Within a PE, commands are dispatched by \textsf{PE\_CPU} to
\textsf{PE\_SCHED}, which routes work to specialized execution
engines---DMA, FETCH/STORE, GEMM, vector-math, and IPCQ
(Fig.~\ref{fig:pe-arch}). Commands come in two flavours. \emph{Atomic}
commands target a single engine---a plain DMA read, a GEMM tile, a
vector-math op, an IPCQ send/receive---and are the natural unit for
short or special-purpose work; the \textsf{PE\_CPU} itself also runs
control-plane work directly. \emph{Composite} commands, by contrast,
carry an ordered pipeline of operations across multiple engines
(\textsf{DMA\_READ} $\rightarrow$ \textsf{FETCH} $\rightarrow$
\textsf{GEMM}/\textsf{MATH} $\rightarrow$ \textsf{STORE} $\rightarrow$
\textsf{DMA\_WRITE}) that the \textsf{PE\_SCHED} tiles and streams
without per-tile redispatch. The composite form is the substrate for
the GEMM optimization (\S\ref{sec:gemm}) and, combined with on-PE
collectives, for fused attention (\S\ref{sec:gqa}).
\begin{figure*}[t]
\centering
\includegraphics[width=\linewidth]{pe_architecture.png}
\caption{Modeled PE architecture used in this report---one example
configuration whose specific parameters are listed in
\S\ref{sec:hw} (Table~\ref{tab:hw}); KernBench is not tied to this
particular decomposition and supports arbitrary PE block layouts
provided each component is given a port and bandwidth model.
\textsf{PE\_CPU} dispatches commands to the \textsf{PE\_SCHED}, which
routes tile-token streams through the \textsf{PE\_DMA},
\textsf{PE\_FETCH\_STORE}, and \textsf{GEMM}/\textsf{MATH} engines
along on-chip links; the \textsf{PE\_IPCQ} provides the control plane
for on-device collective communication.}
\label{fig:pe-arch}
\end{figure*}
KernBench is layered along the flow of a request:
\begin{itemize}
\item The \textbf{runtime API} is host-facing and
topology-agnostic---it deploys tensors and launches kernels but knows
nothing about routing or interconnect.
\item The \textbf{simulation engine} schedules discrete events and
routes every request through the modeled graph.
\item The \textbf{components} are device-side nodes that model
hardware behaviour: the per-PE blocks (scheduler, DMA, GEMM and
vector-math engines, TCM, IPCQ), the NoC routers, the HBM
controllers, and the inter-chiplet links.
\end{itemize}
Data and timing are handled in two passes, so that a kernel's numeric
results and its latency are computed consistently but independently.
\subsection{Latency model: graph traversal and contention}
\label{sec:latency}
The modeled hardware hierarchy described above is represented
internally as a directed graph. Nodes correspond to hardware
components---PEs, routers, memory controllers, SRAM blocks, IO
chiplets---while edges represent communication links with associated
bandwidth (\si{\giga\byte\per\second}) and propagation
(\si{\nano\second}) attributes. Every operation in KernBench---DMA
transfers, remote memory accesses, collective communication, command
dispatch---is modelled as a traversal through this graph. End-to-end
latency is decomposed into four contributions accumulated along the
traversal path (Fig.~\ref{fig:latency-model}): \textbf{per-node
overhead} at each component, \textbf{per-edge transmission} on each
wire, \textbf{drain} (per-flit service occupancy) at the destination,
and \textbf{queuing delay} at the shared FIFOs that the wire and the
destination share between concurrent transactions.
\begin{figure*}[t]
\centering
\includegraphics[width=\textwidth]{latency_model.png}
\caption{Conceptual schematic of the latency model. Two source nodes
(Requester A/B) inject flits through a chain of routers into a
destination node; on the shared edge between routers, flits from the
two transactions are interleaved flit-by-flit by the wire's FIFO
arrival order. End-to-end latency is the sum of four contributions:
\textbf{per-node overhead} (the switch's fixed processing cost,
shown in light yellow---the same colour as the destination's
processing-logic block), \textbf{per-edge transmission}
($\textit{flit\_size}/\textit{BW}$ on each wire),
\textbf{drain} (per-flit service occupancy at the destination's
channel), and \textbf{queuing delay} (waiting in a FIFO when a
shared resource is busy). The places where queuing actually
accumulates are highlighted in green: the router output queue and
the destination's input queue. This is the model, not a
measurement; specific bandwidths and overheads are listed in
\S\ref{sec:hw} (Table~\ref{tab:hw}).}
\label{fig:latency-model}
\end{figure*}
\paragraph{The hardware as a graph.} The topology is compiled once at
configuration time into this graph and is never mutated during a run.
There are no hidden shortcuts, implicit bypasses, or magic paths: if a
request reaches its destination, the path it took is explicit in the
graph, and the latency it incurred is the sum of the per-node and
per-edge costs paid along that path. The same graph representation
applies recursively at every hierarchy level---system, SIP
(Fig.~\ref{fig:sip-arch}), CUBE (Fig.~\ref{fig:cube-arch}), and PE
(Fig.~\ref{fig:pe-arch}).
\paragraph{From graph to discrete-event simulation.} The graph is
driven by a discrete-event engine. Two kinds of events advance
simulation time: \emph{node events} (component switching overhead,
service completions such as an HBM channel commit or a GEMM tile
finish) and \emph{edge events} (the flit-by-flit serialization of a
payload across a bandwidth-limited link). The engine maintains a
priority queue of pending events ordered by their scheduled time and
fires them one at a time, with ties broken under a deterministic
policy so that the same kernel on the same topology always yields the
same trace. Per-request correlation IDs are stamped at injection and
carried through every hop, so the path from injection to completion
is fully traceable. Every nanosecond in a reported latency
corresponds to exactly one of these events on exactly one node or
edge---there is no slack in the budget.
\paragraph{Latency contributions.} Three kinds of latency accumulate
along a traversal: (i) \emph{per-node fixed overhead}---each component
carries a small switching cost (router decode, controller pickup,
scheduler handoff); (ii) \emph{per-edge transfer time}---each link's
payload is decomposed into fixed-size flits (default
\SI{256}{\byte}), and each flit arrives at
$\text{prop}+\text{flit\_bytes}/\text{bw}$ after the previous one,
giving wormhole semantics across multi-hop paths; and (iii)
\emph{per-service occupancy}---memory controllers, GEMM stages, and
collective engines hold the request for their service time before
releasing it downstream. Each of these is attached to a specific node
or edge in the graph; together they make up the entire latency budget.
\paragraph{Congestion: where bottlenecks emerge.} The simulator's
sharpness comes from how it models contention for those nodes and
edges. \emph{Every directed edge has a FIFO}: an arriving flit takes
its bandwidth-limited transfer time on top of whatever earlier flits
are still being served, so a busy link queues later traffic behind
earlier traffic rather than transferring everything at peak BW.
\emph{HBM is modelled with per-pseudo-channel parallelism}: a
stateless array of channel-availability timestamps with address-based
channel selection captures the bank-level concurrency that real HBM
exposes, so 64 channels per CUBE deliver real parallelism on uniform
addresses but a hot channel surfaces as the bottleneck on skewed ones.
\emph{Every component has a serial worker}: a router carrying two
heavy streams interleaves them at flit granularity in arrival order
rather than fanning out for free, so two concurrent collectives sharing
a link share its bandwidth, not double it. Without these mechanisms
the simulator would simply re-confirm the peak-BW roofline; with them,
it reveals where the real bottlenecks form and which hardware levers
actually relieve them---which is exactly the question the codesign
work in this report turns on.
\paragraph{Control-plane (issue) cost model.} The cost of \emph{issuing}
a command is modelled structurally rather than with a per-operation
calibration table. The PE control processor charges, per command,
\[
d_{\text{cmd}} = \textsf{FIXED} + b_{\text{logical}} \cdot R,
\]
where $b_{\text{logical}}$ is the command's hardware-logical byte size,
\textsf{FIXED} captures the fixed per-command cost (queue-tail update,
completion registration) and $R$ captures the per-byte cost of
serializing the command descriptor into the scheduler queue. The
default anchoring (\textsf{FIXED} $= 40$ cycles, $R = 0.0625$
cycles/byte, i.e.\ \SI{16}{\byte\per\cycle}, at \SI{1}{\giga\hertz})
places a typical composite at roughly \SI{43}{\nano\second}, and a
hard cap on a composite's descriptor size prevents the model from
rewarding arbitrarily large fused commands beyond what real descriptor
queues accept. In the configurations measured here, command issue is
not the bottleneck---data movement is---so this term stays small
relative to DMA and collective time.
\paragraph{Accuracy.} The model is precise about the effects that
dominate kernel latency on this class of hardware: per-edge bandwidth
occupancy and flit-level serialization, HBM pseudo-channel parallelism,
and per-component switching overhead. Two independent cross-checks
drawn from the experiments in this report confirm that this precision
translates into physically reasonable kernel latencies. First, in the
GEMM study (\S\ref{sec:gemm}), simulator-measured MAC efficiency
tracks an analytic ideal-pipeline model within roughly
\SIrange{10}{20}{\percent} across a wide range of tile counts; the
residual gap is attributable to pipeline-fill and DMA effects the
analytic model omits. Second, in the all-reduce study
(\S\ref{sec:allreduce}, Fig.~\ref{fig:allreduce-cmp}), simulator
latency for a 2D-torus over six devices follows the expected
startup-plus-per-packet shape across the entire payload sweep---tight
at small payloads where startup dominates, and within a single-digit
multiplicative factor at the largest payloads, where the residual gap
is explained by per-router switching the analytic shape elides. A
single-device point from an external full-system simulator (FSIM) at
the largest payload sits an order of magnitude above the KernBench
multi-device torus, illustrating the well-known gap between an
achievable-kernel number and a full end-to-end-stack number rather
than a model error. The known simplifications---FIFO router
arbitration (instead of round-robin), HBM scheduler without
write-buffer reordering, no bank conflict, no refresh or thermal
effects, and no upstream backpressure---are the price of a
deterministic, inspectable model. These simplifications bound the
absolute accuracy, but the agreement with both analytic models and the
external full-system simulator data above indicates that KernBench is
sufficiently accurate for evaluating the \emph{relative}
hardware--software design trade-offs (tiling A vs.\ B, topology X
vs.\ Y, with vs.\ without composite command, mesh vs.\ torus) that
are the primary objective of this work.
\subsection{Modeled hardware configuration}
\label{sec:hw}
Table~\ref{tab:hw} summarizes the hardware configuration used for
every experiment in this report. It is read directly from the
simulator's topology description; per-experiment workload parameters
(matrix shapes, collective sizes, sequence lengths) are stated in
their respective sections rather than here.
\begin{table}[t]
\centering
\caption{Modeled hardware configuration (shared by all experiments).}
\label{tab:hw}
\small
\begin{tabular}{@{}ll@{}}
\toprule
\textbf{Parameter} & \textbf{Value} \\
\midrule
\multicolumn{2}{@{}l}{\emph{Hierarchy}} \\
SIPs & 2 (1D ring) \\
CUBEs per SIP & 16 ($4\times4$ mesh) \\
PEs per CUBE & 8 (4 corners $\times$ 2) \\
PEs total & 256 \\
\midrule
\multicolumn{2}{@{}l}{\emph{Processing element (PE)}} \\
GEMM engine peak & \SI{8}{\tera\flop\per\second} (f16) \\
TCM (on-PE) & \SI{16}{\mega\byte}, \SI{512}{\giga\byte\per\second} R/W \\
\quad kernel scratch & \SI{1}{\mega\byte} \\
DMA engines & 1 read + 1 write \\
CPU / scheduler overhead & \SI{2}{\nano\second} / \SI{1}{\nano\second} \\
\midrule
\multicolumn{2}{@{}l}{\emph{Memory (per CUBE)}} \\
HBM capacity & \SI{48}{\giga\byte} (8 slices) \\
HBM aggregate BW & \SI{1024}{\giga\byte\per\second} \\
HBM pseudo-channels & 64 (8 per PE), \SI{32}{\giga\byte\per\second} each \\
SRAM (shared) & \SI{32}{\mega\byte}, \SI{128}{\giga\byte\per\second} link \\
HBM burst & \SI{256}{\byte} \\
\midrule
\multicolumn{2}{@{}l}{\emph{Interconnect}} \\
Intra-CUBE NoC link & \SI{256}{\giga\byte\per\second}, \SI{2}{\nano\second}/router \\
Inter-CUBE (UCIe PHY) & \SI{512}{\giga\byte\per\second}, \SI{8}{\nano\second}, XY routing \\
Inter-SIP (PCIe) & \SI{768}{\giga\byte\per\second} per endpoint \\
\midrule
\multicolumn{2}{@{}l}{\emph{Command-issue cost model (defaults)}} \\
FIXED per command & 40 cycles \\
per-byte rate $R$ & 0.0625 cycles/byte (\SI{16}{\byte\per\cycle}) \\
composite size cap & \SI{1024}{\byte} \\
\bottomrule
\end{tabular}
\end{table}