paper(platform): WIP — §2 KernBench Platform mid-edit

Mid-pass through §2. Captured progress so far:

- §2 intro previews the three new threads (graph view, DES engine,
  congestion) the rewritten latency section weaves together.
- §2.3 retitled to "Latency model: a graph traversed by events" and
  restructured into bold-led paragraphs:
  - The hardware as a graph (nodes = components, edges = links).
  - From graph to discrete-event simulation (node/edge events,
    deterministic ordering, correlation-ID trace).
  - Latency contributions (per-node fixed, per-edge size-aware,
    per-service occupancy).
  - Congestion: per-edge FIFO BW occupancy, HBM per-PC parallelism,
    component serial workers — the mechanisms that surface real
    bottlenecks instead of peak-BW roofline.
  - Control-plane cost model (FIXED + b·R) — unchanged.
  - Accuracy: extended with a second cross-check from the all-reduce
    study (torus vs. analytic startup-plus-per-packet model + FSIM
    external single-device reference). Existing GEMM analytic-vs-
    measured 10-20% check retained.
- §2.4 Hardware-configuration table unchanged.

Still TODO: cross-check accuracy numbers against a fresh
milestone-1h-gemm/ccl re-run; the current artifacts predate the
ADR-0064 Rev2 cost-model and IPCQ Phase-2 race fix and may need a
refresh before §2.3 final lock-in.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-06-11 16:31:42 -07:00
parent d5267ec718
commit 9270d3435a
2 changed files with 132 additions and 75 deletions
Binary file not shown.
@@ -2,11 +2,12 @@
\label{sec:platform}
All results in this report are produced on \emph{KernBench}, a
system-level, discrete-event simulator for LLM kernels running on
SIP-based AI accelerators. This section explains why the platform exists,
how it executes a kernel, how it computes latency and how accurate that
number is, and the concrete hardware configuration used for every
experiment that follows.
system-level discrete-event simulator for LLM kernels running on AHBM.
This section explains why the platform exists, how it executes a
kernel, how its latency model works---specifically how the hardware is
viewed as a graph, how that graph is driven by a discrete-event engine,
and how congestion is captured---and the concrete hardware configuration
used for every experiment that follows.
\subsection{Why KernBench: source-level kernels without a software stack}
\label{sec:why}
@@ -21,16 +22,16 @@ well-written kernel be, and which hardware features actually make it
faster?}
KernBench is built to answer exactly that question. Kernels are written
and executed at the \emph{source level}---as algorithmic descriptions in a
small tile-oriented kernel API---with no dependency on a compiler or any
other software-stack layer. The simulator takes the kernel and a hardware
topology and reports the latency that the modeled hardware would deliver.
This isolation is deliberate: it lets us study algorithm-level
and executed at the \emph{source level}---as algorithmic descriptions
in a small tile-oriented kernel API---with no dependency on a compiler
or any other software-stack layer. The simulator takes the kernel and a
hardware topology and reports the latency the modeled hardware would
deliver. This isolation is deliberate: it lets us study algorithm-level
optimizations (how to tile a GEMM, how to schedule a collective, how to
fuse an attention kernel) and the hardware features that support them,
without the confound of compiler maturity or framework overhead. The cost
is that KernBench numbers are \emph{not} E2E latencies; they are the
achievable-kernel latencies that an ideal software stack would expose.
without the confound of compiler maturity or framework overhead. The
cost is that KernBench numbers are \emph{not} E2E latencies; they are
the achievable-kernel latencies an ideal software stack would expose.
\subsection{Execution model}
\label{sec:exec}
@@ -43,90 +44,146 @@ KernBench is layered along the flow of a request:
nothing about routing or interconnect.
\item The \textbf{simulation engine} schedules discrete events, routes
every request through the modeled graph, and tracks completion via
correlation IDs.
\item The \textbf{components} are device-side nodes that model hardware
behavior: the per-PE blocks (scheduler, DMA, GEMM and vector-math
engines, TCM, IPCQ), the NoC routers, the HBM controllers, and the
inter-chiplet links.
per-request correlation IDs.
\item The \textbf{components} are device-side nodes that model
hardware behavior: the per-PE blocks (scheduler, DMA, GEMM and
vector-math engines, TCM, IPCQ), the NoC routers, the HBM
controllers, and the inter-chiplet links.
\end{itemize}
The topology is compiled once at configuration time into an authoritative
graph of components and links; it is never mutated during a run. Within a
PE, work is expressed as \emph{composite commands}: a single command
carries an ordered pipeline of operations
Within a PE, work is expressed as \emph{composite commands}: a single
command carries an ordered pipeline of operations
(\textsf{DMA\_READ} $\rightarrow$ \textsf{FETCH} $\rightarrow$
\textsf{GEMM}/\textsf{MATH} $\rightarrow$ \textsf{STORE} $\rightarrow$
\textsf{DMA\_WRITE}) that the PE scheduler tiles and streams. This
composite mechanism is the substrate for the GEMM optimization
(\S\ref{sec:gemm}) and, combined with on-PE collectives, for fused
attention (\S\ref{sec:gqa}). Data and timing are handled in two passes, so
that a kernel's numeric results and its latency are computed
attention (\S\ref{sec:gqa}). Data and timing are handled in two passes,
so that a kernel's numeric results and its latency are computed
consistently but independently.
\subsection{Latency model and its accuracy}
\subsection{Latency model: a graph traversed by events}
\label{sec:latency}
KernBench obeys a set of golden invariants that keep its latency numbers
physically meaningful. End-to-end latency is computed \emph{strictly by
explicit traversal} over modeled components and links: every routed
request incurs latency greater than zero, routing is deterministic, and
every valid request flow has explicit connectivity. There are no hidden
shortcuts, implicit waits, or magic delays---if a delay exists in the
result, it came from a scheduled event on a modeled component or link.
\paragraph{The hardware as a graph.} KernBench views the modeled
hardware as a directed graph. \emph{Nodes} are the components listed
above; \emph{edges} are the interconnect links between them, each
carrying bandwidth (\si{\giga\byte\per\second}) and propagation
(\si{\nano\second}) attributes. The topology is compiled once at
configuration time into this graph and is never mutated during a run.
Every routed request---a DMA, a remote read, an IPCQ message, a
kernel-launch command---is a \emph{traversal} of this graph from a
source node to a destination service, hopping through routers and
links along the way. There are no hidden shortcuts, implicit bypasses,
or magic paths: if a request reaches its destination, the path it took
is explicit in the graph, and the latency it incurred is the sum of
the per-node and per-edge costs paid along that path.
Latency accumulates from three kinds of contributions: per-node fixed
overheads (each component carries an \texttt{overhead\_ns}), per-link
transfer time, and per-service occupancy. The interconnect is modeled at
fine granularity. Each directed link serializes traffic through a
bandwidth-limited FIFO, so a busy link delays later flits. Payloads are
decomposed into fixed-size flits (default \SI{256}{\byte} bursts) that
arrive at $\text{prop} + \text{flit\_bytes}/\text{bw}$ intervals, so link
bandwidth throttles arrival rate rather than being applied as a lump sum.
HBM is modeled with per-pseudo-channel parallelism: a stateless array of
channel-availability timestamps with address-based channel selection
captures bank-level concurrency.
\paragraph{From graph to discrete-event simulation.} The graph is
driven by a discrete-event engine. Two kinds of events advance
simulation time: \emph{node events} (component switching overhead,
service completions such as an HBM channel commit or a GEMM tile
finish) and \emph{edge events} (the flit-by-flit serialization of a
payload across a bandwidth-limited link). The engine maintains a
priority queue of pending events ordered by their scheduled time,
fires them one at a time, and treats ties under a deterministic
ordering policy so that the same kernel on the same topology always
yields the same trace. Per-request correlation IDs are stamped at
injection and carried through every hop, so the trace is recoverable
from injection to completion. Every nanosecond in a reported latency
traces back to exactly one of these events on exactly one node or
edge.
The cost of \emph{issuing} a command is modeled structurally rather than
with a per-operation calibration table. The PE control processor charges, per command,
\paragraph{Latency contributions.} Three kinds of latency accumulate
along a traversal: (i) \emph{per-node fixed overhead}---each component
carries a small switching cost (router decode, controller pickup,
scheduler handoff); (ii) \emph{per-edge transfer time}---each link's
payload is decomposed into fixed-size flits (default
\SI{256}{\byte}), and each flit arrives at
$\text{prop}+\text{flit\_bytes}/\text{bw}$ after the previous one,
giving wormhole semantics across multi-hop paths; and (iii)
\emph{per-service occupancy}---memory controllers, GEMM stages, and
collective engines hold the request for their service time before
releasing it downstream. Each of these is attached to a specific node
or edge in the graph; together they make up the entire latency budget.
\paragraph{Congestion: where bottlenecks emerge.} The simulator's
sharpness comes from how it models contention for those nodes and
edges. \emph{Every directed edge has a FIFO}: an arriving flit takes
its bandwidth-limited transfer time on top of whatever earlier flits
are still being served, so a busy link queues later traffic behind
earlier traffic rather than transferring everything at peak BW.
\emph{HBM is modeled with per-pseudo-channel parallelism}: a stateless
array of channel-availability timestamps with address-based channel
selection captures the bank-level concurrency that real HBM exposes,
so 64 channels per CUBE deliver real parallelism on uniform addresses
but a hot channel surfaces as the bottleneck on skewed ones.
\emph{Every component has a serial worker}: a router carrying two
heavy streams interleaves them at flit granularity in arrival order
rather than fanning out for free, so two concurrent collectives sharing
a link share its bandwidth, not double it. Without these mechanisms
the simulator would simply re-confirm the peak-BW roofline; with them,
it reveals where the real bottlenecks form and which hardware levers
actually relieve them---which is exactly the question the codesign
work in this report turns on.
\paragraph{Control-plane (issue) cost model.} The cost of \emph{issuing}
a command is modeled structurally rather than with a per-operation
calibration table. The PE control processor charges, per command,
\[
d_{\text{cmd}} = \textsf{FIXED} + b_{\text{logical}} \cdot R,
\]
where $b_{\text{logical}}$ is the command's hardware-logical byte size,
\textsf{FIXED} captures the fixed per-command cost (queue-tail
update, completion registration) and $R$ captures the per-byte cost of
serializing the command descriptor into the scheduler queue. The default
anchoring (\textsf{FIXED} $=40$ cycles, $R = 0.0625$ cycles/byte, i.e.\
\SI{16}{\byte\per\cycle}, at \SI{1}{\giga\hertz}) places a typical
composite at roughly \SI{43}{\nano\second}, and a hard cap on a
composite's descriptor size prevents the model from rewarding
arbitrarily large fused commands beyond what real descriptor queues
accept. The dispatch cost is enabled and tuned per topology; in the
configurations measured here, command issue is not the bottleneck---data
movement is---so this term stays small relative to DMA and collective
time.
\textsf{FIXED} captures the fixed per-command cost (queue-tail update,
completion registration) and $R$ captures the per-byte cost of
serializing the command descriptor into the scheduler queue. The
default anchoring (\textsf{FIXED} $= 40$ cycles, $R = 0.0625$
cycles/byte, i.e.\ \SI{16}{\byte\per\cycle}, at \SI{1}{\giga\hertz})
places a typical composite at roughly \SI{43}{\nano\second}, and a
hard cap on a composite's descriptor size prevents the model from
rewarding arbitrarily large fused commands beyond what real descriptor
queues accept. In the configurations measured here, command issue is
not the bottleneck---data movement is---so this term stays small
relative to DMA and collective time.
How accurate is all this? The model is precise about the things that
dominate kernel latency on this class of hardware: link bandwidth
occupancy and serialization, HBM channel parallelism, flit-level
streaming, and per-component switching overhead. The GEMM study
(\S\ref{sec:gemm}) provides a direct check: the measured MAC efficiency
tracks the analytic (theoretical) efficiency within roughly
\SIrange{10}{20}{\percent} across a wide range of tile counts, with the
gap attributable to pipeline fill and DMA effects that the analytic model
omits. The known simplifications---idealized arbitration, no thermal or
refresh effects, fixed burst granularity---are the price of a
deterministic, inspectable model; they bound the absolute accuracy but do
not distort the \emph{relative} comparisons (tiling A vs.\ B, topology X
vs.\ Y) that this report is built on.
\paragraph{Accuracy.} The model is precise about the effects that
dominate kernel latency on this class of hardware: per-edge bandwidth
occupancy and flit-level serialization, HBM pseudo-channel parallelism,
and per-component switching overhead. Two independent cross-checks
drawn from the experiments in this report confirm that this precision
translates into physically reasonable kernel latencies. First, in the
GEMM study (\S\ref{sec:gemm}), simulator-measured MAC efficiency
tracks an analytic ideal-pipeline model within roughly
\SIrange{10}{20}{\percent} across a wide range of tile counts; the
residual gap is attributable to pipeline-fill and DMA effects the
analytic model omits. Second, in the all-reduce study
(\S\ref{sec:allreduce}, Fig.~\ref{fig:allreduce-cmp}), simulator
latency for a 2D-torus over six devices follows the expected
startup-plus-per-packet shape across the entire payload sweep---tight
at small payloads where startup dominates, and within a single-digit
multiplicative factor at the largest payloads, where the residual gap
is explained by per-router switching the analytic shape elides. A
single-device point from an external full-system simulator (FSIM) at
the largest payload sits an order of magnitude above the KernBench
multi-device torus, illustrating the well-known gap between an
achievable-kernel number and a full end-to-end-stack number rather
than a model error. The known simplifications---FIFO router
arbitration (instead of round-robin), HBM scheduler without
write-buffer reordering, no bank conflict, no refresh or thermal
effects, and no upstream backpressure---are the price of a
deterministic, inspectable model. They bound the absolute accuracy but
do not distort the \emph{relative} comparisons (tiling A vs.\ B,
topology X vs.\ Y, with vs.\ without composite command, mesh vs.\
torus) that this report is built on.
\subsection{Modeled hardware configuration}
\label{sec:hw}
Table~\ref{tab:hw} summarizes the hardware configuration used for every
experiment in this report. It is read directly from the simulator's
topology description; per-experiment workload parameters (matrix shapes,
collective sizes, sequence lengths) are stated in their respective
sections rather than here.
Table~\ref{tab:hw} summarizes the hardware configuration used for
every experiment in this report. It is read directly from the
simulator's topology description; per-experiment workload parameters
(matrix shapes, collective sizes, sequence lengths) are stated in
their respective sections rather than here.
\begin{table}[t]
\centering