23992548f7
Three coupled fixes that recover small-tile GEMM pipeline efficiency from 53% to 88% (32x3072x32 load_ref, composite_window basis). 1. PE_DMA channel-hold (ADR-0014 D4 clarified): both the _handle_with_hooks (PeInternalTxn) and _pipeline_process (TileToken) paths used to hold the cap=1 DMA channel through the full HBM round-trip, which double-serialized with the HBM_CTRL's own per-PC `available_at` model and prevented back-to-back tile DMAs from amortizing their per-request head latency. Channel is now released after the request is enqueued onto the next hop; HBM serialization is HBM_CTRL's responsibility alone. Tests: new test_pe_dma_back_to_back_pipelining as the oracle (asserts wall < 75% of strict-serialized N x single_op). Existing test_pe_dma_record_start_after_channel_acquire rewritten to assert t_start clustering (channel released fast) instead of the old round-trip-hold invariant. test_pe_dma_same_channel_serializes still passes — HBM_CTRL preserves ordering. Probe regression: PE→local-HBM 32 KiB stays at 141 ns (single-request, unaffected). 2. milestone_1h_gemm bench: matmul_composite was reading MATMUL_M/K/N env vars at module load, so every sweep row replayed the cached 256³ result; values now read inside run(). Drops the stale sys.modules deletion hack. 3. Analytic ideal-pipeline model: dropped the (n_mn-1)·dma_w_per_pair penalty (over-pessimistic for under-tile shapes — it pushed measured > theoretical) and replaced the D_STAGES-derived head with empirical T_PIPELINE_FILL=60 ns / T_PIPELINE_TAIL=30 ns. Max analytic-vs-measured gap across all 7 swept shapes now 2.2 ppt (was 9-44 ppt under the old constants). Paper updates: - §3 (GEMM): 78%→88% measured at 48 tiles, 23%→15% at 1 tile, stage breakdown numbers refreshed (DMA in / Fetch / GEMM all ~785 ns at K=3072), analytic-vs-measured agreement tightened to "within 2.2 ppt". - §2.4 (Accuracy): GEMM tracking claim refreshed accordingly. - §5 (GQA): restore long-ctx 4-cases figures into figures/ (they were dropped from bench output dir as derived artifacts in92b9221/e45626cbut §5 still cites them by name). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
461 lines
22 KiB
TeX
461 lines
22 KiB
TeX
\begin{figure*}[t]
|
|
\centering
|
|
\begin{subfigure}[b]{0.495\textwidth}
|
|
\centering
|
|
\includegraphics[width=\linewidth,height=0.7\linewidth,keepaspectratio]{sip_architecture.pdf}
|
|
\caption{SIP Architecture}
|
|
\label{fig:sip-arch}
|
|
\end{subfigure}\hfill
|
|
\begin{subfigure}[b]{0.42\textwidth}
|
|
\centering
|
|
\includegraphics[width=\linewidth,height=0.84\linewidth]{cube_architecture.pdf}
|
|
\caption{CUBE Architecture}
|
|
\label{fig:cube-arch}
|
|
\end{subfigure}
|
|
|
|
\vspace{0.6em}
|
|
|
|
\begin{subfigure}[t]{\textwidth}
|
|
\centering
|
|
\includegraphics[width=\linewidth]{pe_architecture.png}
|
|
\caption{PE level (zoom-in of one PE in (b)): \textsf{PE\_CPU}
|
|
dispatches commands to \textsf{PE\_SCHED}, which routes tile-token
|
|
streams through \textsf{PE\_DMA}, \textsf{PE\_FETCH\_STORE}, and
|
|
\textsf{GEMM}/\textsf{MATH} engines; \textsf{PE\_IPCQ} is the
|
|
on-device collective control plane.}
|
|
\label{fig:pe-arch}
|
|
\end{subfigure}
|
|
\caption{Modeled hardware graph at the SIP, CUBE, and PE levels.
|
|
This figure is an \emph{illustrative topology} chosen for
|
|
readability; the \emph{experimental configuration} used throughout
|
|
this report (port counts, slice counts, and link parameters) is
|
|
specified in \S\ref{sec:hw} / Table~\ref{tab:hw}, and numbers in the
|
|
two may differ. KernBench is not tied to this particular
|
|
arrangement: each box is a modeled component node, each line a
|
|
directed link with bandwidth and propagation attributes, and any
|
|
topology that respects those attributes is supported.}
|
|
\label{fig:hw-arch}
|
|
\end{figure*}
|
|
|
|
\begin{figure*}[t]
|
|
\centering
|
|
\includegraphics[width=\textwidth]{latency_model.png}
|
|
\caption{Conceptual schematic of the latency model. Two source nodes
|
|
(Requester A/B) inject flits through a chain of routers into a
|
|
destination node; on the shared edge between routers, flits from the
|
|
two transactions are interleaved flit-by-flit by the wire's FIFO
|
|
arrival order. End-to-end latency is the sum of four contributions:
|
|
\textbf{per-node overhead} (the switch's fixed processing cost,
|
|
shown in light yellow---the same colour as the destination's
|
|
processing-logic block), \textbf{per-edge transmission}
|
|
($\textit{flit\_size}/\textit{BW}$ on each wire),
|
|
\textbf{drain} (per-flit service occupancy at the destination's
|
|
channel), and \textbf{queuing delay} (waiting in a FIFO when a
|
|
shared resource is busy). The places where queuing actually
|
|
accumulates are highlighted in green: the router output queue and
|
|
the destination's input queue. This is the model, not a
|
|
measurement; specific bandwidths and overheads are listed in
|
|
\S\ref{sec:hw} (Table~\ref{tab:hw}).}
|
|
\label{fig:latency-model}
|
|
\end{figure*}
|
|
|
|
\section{The KernBench Platform}
|
|
\label{sec:platform}
|
|
|
|
All results in this report are produced on \emph{KernBench}, a
|
|
system-level discrete-event simulator for LLM kernels running on AHBM.
|
|
This section explains why the platform exists, the device and
|
|
execution model it presents to a kernel writer, how its latency model
|
|
turns that execution into a number, and the concrete hardware
|
|
configuration used for every experiment that follows.
|
|
|
|
\subsection{Why KernBench}
|
|
\label{sec:why}
|
|
|
|
In a production end-to-end (E2E) stack, kernel performance is entangled
|
|
with every layer above the hardware: the compiler's tiling and
|
|
scheduling choices, the framework's operator dispatch, the collective
|
|
library, and the runtime. Good E2E numbers require \emph{all} of those
|
|
layers to be co-optimized, which makes it hard to answer a narrower but
|
|
more fundamental question: \emph{given the hardware, how fast can a
|
|
well-written kernel be, and which hardware features actually make it
|
|
faster?}
|
|
|
|
KernBench is built to answer exactly that question. Kernels are written
|
|
and executed at the \emph{source level}---as algorithmic descriptions
|
|
in a small tile-oriented kernel API---with no dependency on a compiler
|
|
or any other software-stack layer. The simulator takes the kernel and a
|
|
hardware topology and reports the latency the modeled hardware would
|
|
deliver. This isolation is deliberate: it lets us study algorithm-level
|
|
optimizations (how to tile a GEMM, how to schedule a collective, how to
|
|
fuse an attention kernel) and the hardware features that support them,
|
|
without the confound of compiler maturity or framework overhead. The
|
|
cost is that KernBench numbers are \emph{not} E2E latencies; they are
|
|
the achievable-kernel latencies an ideal software stack would expose.
|
|
|
|
\subsection{Device and execution model}
|
|
\label{sec:exec}
|
|
|
|
KernBench models AHBM as a hierarchy of SIPs, CUBEs, and processing
|
|
elements (PEs). At the system level, multiple SIPs are connected
|
|
through inter-package links, while each SIP contains a collection of
|
|
CUBEs joined by an on-package interconnect
|
|
(Fig.~\ref{fig:hw-arch}\subref{fig:sip-arch}).
|
|
|
|
Each CUBE contains eight PEs, shared SRAM, HBM controllers, an
|
|
\textsf{M\_CPU} control processor, and an intra-CUBE router mesh
|
|
(Fig.~\ref{fig:hw-arch}\subref{fig:cube-arch}). Together these
|
|
components form the execution substrate for all kernels evaluated in
|
|
this report. This organization reflects the memory-centric nature of
|
|
AHBM: each PE is paired with a dedicated slice of HBM bandwidth and
|
|
local on-PE storage (TCM), so compute lives next to the data it
|
|
consumes rather than fetching it through a far-away memory
|
|
controller. The platform's performance question is therefore not
|
|
``how many FLOPs can the chip do'' but ``how well can a kernel keep
|
|
each compute engine fed from its locally-attached memory while
|
|
moving the unavoidable traffic between PEs efficiently.''
|
|
|
|
Within a PE, commands are dispatched by \textsf{PE\_CPU} to
|
|
\textsf{PE\_SCHED}, which routes work to specialized execution
|
|
engines---DMA, FETCH/STORE, GEMM, vector-math, and IPCQ
|
|
(Fig.~\ref{fig:hw-arch}\subref{fig:pe-arch}). Commands come in two
|
|
flavours. \emph{Atomic} commands target a single engine---a plain
|
|
DMA read, a GEMM tile, a vector-math op, an IPCQ send/receive---and
|
|
are the natural unit for short or special-purpose work; the
|
|
\textsf{PE\_CPU} itself also runs control-plane work directly.
|
|
\emph{Composite} commands, by contrast, carry an ordered pipeline of
|
|
operations across multiple engines (\textsf{DMA\_READ} $\rightarrow$
|
|
\textsf{FETCH} $\rightarrow$ \textsf{GEMM}/\textsf{MATH} $\rightarrow$
|
|
\textsf{STORE} $\rightarrow$ \textsf{DMA\_WRITE}) that the
|
|
\textsf{PE\_SCHED} tiles and streams without per-tile redispatch. The
|
|
composite form is the substrate for the GEMM optimization
|
|
(\S\ref{sec:gemm}) and, combined with on-PE collectives, for fused
|
|
attention (\S\ref{sec:gqa}).
|
|
|
|
KernBench is layered along the flow of a request:
|
|
|
|
\begin{itemize}
|
|
\item The \textbf{runtime API} is host-facing and
|
|
topology-agnostic---it deploys tensors and launches kernels but knows
|
|
nothing about routing or interconnect.
|
|
\item The \textbf{simulation engine} schedules discrete events and
|
|
routes every request through the modeled graph.
|
|
\item The \textbf{components} are device-side nodes that model
|
|
hardware behaviour: the per-PE blocks (scheduler, DMA, GEMM and
|
|
vector-math engines, TCM, IPCQ), the NoC routers, the HBM
|
|
controllers, and the inter-chiplet links.
|
|
\end{itemize}
|
|
|
|
Data and timing are handled in two passes, so that a kernel's
|
|
numeric results and its latency are computed consistently but
|
|
independently. \textbf{Pass~1 (timing)} runs the kernel under
|
|
the discrete-event engine: memory ops (\textsf{tl.load},
|
|
\textsf{tl.store}) execute against a host-side \textsf{MemoryStore}
|
|
and return real tensor data, while compute ops (GEMM, vector-math)
|
|
only emit records into an op-log carrying their operands, shapes,
|
|
dtypes, and scheduled time. \textbf{Pass~2 (data)} replays that
|
|
op-log offline in numpy, producing the actual numeric outputs and
|
|
comparing them against a reference computation under per-dtype
|
|
tolerances (e.g.\ \texttt{rtol}/\texttt{atol} $=10^{-3}$ for f16).
|
|
Pass~2 is optional---runs that need only latency skip it---but when
|
|
enabled it guarantees that every reported timing number corresponds
|
|
to a kernel whose numeric output has been verified end-to-end.
|
|
|
|
\subsection{Latency model: graph traversal and contention}
|
|
\label{sec:latency}
|
|
|
|
The modeled hardware hierarchy described above is represented
|
|
internally as a directed graph. Nodes correspond to hardware
|
|
components---PEs, routers, memory controllers, SRAM blocks, IO
|
|
chiplets---while edges represent communication links with associated
|
|
bandwidth (\si{\giga\byte\per\second}) and propagation
|
|
(\si{\nano\second}) attributes. Every operation in KernBench---DMA
|
|
transfers, remote memory accesses, collective communication, command
|
|
dispatch---is modelled as a traversal through this graph. End-to-end
|
|
latency is decomposed into four contributions accumulated along the
|
|
traversal path (Fig.~\ref{fig:latency-model}): \textbf{per-node
|
|
overhead} at each component, \textbf{per-edge transmission} on each
|
|
wire, \textbf{drain} (per-flit service occupancy) at the destination,
|
|
and \textbf{queuing delay} at the shared FIFOs that the wire and the
|
|
destination share between concurrent transactions.
|
|
|
|
\paragraph{The hardware as a graph.} The topology is compiled once at
|
|
configuration time into this graph and is never mutated during a run.
|
|
There are no hidden shortcuts, implicit bypasses, or magic paths: if a
|
|
request reaches its destination, the path it took is explicit in the
|
|
graph, and the latency it incurred is the sum of the per-node and
|
|
per-edge costs paid along that path. The same graph representation
|
|
applies recursively at every hierarchy level---system, SIP, CUBE, and
|
|
PE (Fig.~\ref{fig:hw-arch}).
|
|
|
|
\paragraph{From graph to discrete-event simulation.} The graph is
|
|
driven by a discrete-event engine. Two kinds of events advance
|
|
simulation time: \emph{node events} (component switching overhead,
|
|
service completions such as an HBM channel commit or a GEMM tile
|
|
finish) and \emph{edge events} (the flit-by-flit serialization of a
|
|
payload across a bandwidth-limited link). The engine maintains a
|
|
priority queue of pending events ordered by their scheduled time and
|
|
fires them one at a time, with ties broken under a deterministic
|
|
policy so that the same kernel on the same topology always yields the
|
|
same trace. Per-request correlation IDs are stamped at injection and
|
|
carried through every hop, so the path from injection to completion
|
|
is fully traceable. Every modeled latency contribution
|
|
corresponds to exactly one of these events on exactly one node or
|
|
edge---there is no slack in the budget.
|
|
|
|
\paragraph{Latency contributions.} Three kinds of latency accumulate
|
|
along a traversal: (i) \emph{per-node fixed overhead}---each component
|
|
carries a small switching cost (router decode, controller pickup,
|
|
scheduler handoff); (ii) \emph{per-edge transfer time}---each link's
|
|
payload is decomposed into fixed-size flits (default
|
|
\SI{256}{\byte}), and each flit arrives at
|
|
$\text{prop}+\text{flit\_bytes}/\text{bw}$ after the previous one,
|
|
giving wormhole semantics across multi-hop paths; and (iii)
|
|
\emph{per-service occupancy}---memory controllers, GEMM stages, and
|
|
collective engines hold the request for their service time before
|
|
releasing it downstream. Each of these is attached to a specific node
|
|
or edge in the graph; together they make up the entire latency budget.
|
|
|
|
Beyond data movement and execution latency, KernBench
|
|
also models the control-plane cost required to issue work
|
|
to accelerator engines. Since every DMA, GEMM, vector-
|
|
math, and IPCQ operation is initiated through the
|
|
\textsf{PE\_CPU} $\rightarrow$ \textsf{PE\_SCHED} path,
|
|
command dispatch latency contributes to the end-to-end
|
|
execution time of all kernels.
|
|
|
|
\paragraph{Command dispatch overhead model.} The cost incurred by
|
|
the \textsf{PE\_CPU} when it dispatches a command to one of the
|
|
accelerator engines is modelled structurally rather than with a
|
|
per-operation calibration table. The \textsf{PE\_CPU} charges, per
|
|
command,
|
|
\[
|
|
d_{\text{cmd}} = \textsf{FIXED} + b_{\text{logical}} \cdot R,
|
|
\]
|
|
This linear-in-size form reflects the serialization cost of writing a
|
|
command descriptor into the scheduler queue: a fixed per-command
|
|
bookkeeping cost plus a byte-wise descriptor-transfer cost.
|
|
Concretely, $b_{\text{logical}}$ is the command's hardware-logical byte
|
|
size, \textsf{FIXED} captures the fixed per-command cost (queue-tail
|
|
update, completion registration) and $R$ captures the per-byte cost of
|
|
serializing the command descriptor into the scheduler queue. This
|
|
\textsf{FIXED} is distinct from Table~\ref{tab:hw}'s
|
|
\SI{2}{\nano\second} / \SI{1}{\nano\second} \textsf{PE\_CPU} /
|
|
\textsf{PE\_SCHED} fixed costs: the latter are per-component
|
|
traversal overheads each command pays when it transits those nodes,
|
|
whereas the 40-cycle \textsf{FIXED} term is the command-descriptor
|
|
issue cost. The
|
|
default anchoring (\textsf{FIXED} $= 40$ cycles, $R = 0.0625$
|
|
cycles/byte, i.e.\ \SI{16}{\byte\per\cycle}, at \SI{1}{\giga\hertz})
|
|
places a typical composite at roughly \SI{43}{\nano\second}, and a
|
|
hard cap on a composite's descriptor size prevents the model from
|
|
rewarding arbitrarily large fused commands beyond what real descriptor
|
|
queues accept. In the configurations measured here, command issue is
|
|
not the bottleneck---data movement is---so this term stays small
|
|
relative to DMA and collective time.
|
|
|
|
\begin{table*}[t]
|
|
\centering
|
|
\caption{PE\,$\to$\,HBM DMA latency probe at varying hop distances
|
|
(\SI{32}{\kibi\byte} transfer; output captured from \texttt{kernbench
|
|
probe}). \emph{Util\%} is the achieved bandwidth as a fraction of the
|
|
path's bottleneck-edge bandwidth.}
|
|
\label{tab:probe-pe-dma}
|
|
\small
|
|
\begin{tabular}{@{}lrrr@{}}
|
|
\toprule
|
|
Case & Latency~(\si{\nano\second}) & Util\% (\,32\,KiB) & Util\% (\,1\,MiB) \\
|
|
\midrule
|
|
PE\,$\to$\,local HBM & 141.0 & 90.8 & 100.0 \\
|
|
PE\,$\to$\,same-half HBM & 147.9 & 86.6 & ~99.9 \\
|
|
PE\,$\to$\,cross-half HBM & 161.2 & 79.4 & ~99.8 \\
|
|
PE\,$\to$\,cross-CUBE (best) & 330.5 & 77.5 & ~99.6 \\
|
|
PE\,$\to$\,cross-CUBE (worst) & 677.1 & 37.8 & ~97.8 \\
|
|
\bottomrule
|
|
\end{tabular}
|
|
\end{table*}
|
|
|
|
\subsubsection{Congestion and contention modeling}
|
|
\label{sec:congestion}
|
|
|
|
The simulator's
|
|
sharpness comes from how it models contention for those nodes and
|
|
edges. \emph{Every directed edge has a FIFO}: an arriving flit takes
|
|
its bandwidth-limited transfer time on top of whatever earlier flits
|
|
are still being served, so a busy link queues later traffic behind
|
|
earlier traffic rather than transferring everything at peak BW.
|
|
\emph{HBM is modelled with per-pseudo-channel parallelism}: a
|
|
stateless array of channel-availability timestamps with address-based
|
|
channel selection captures the bank-level concurrency that real HBM
|
|
exposes, so 64 channels per CUBE deliver real parallelism on uniform
|
|
addresses but a hot channel surfaces as the bottleneck on skewed ones.
|
|
\emph{Every component has a serial worker}: a router carrying two
|
|
heavy streams interleaves them at flit granularity in arrival order
|
|
rather than fanning out for free, so two concurrent collectives sharing
|
|
a link share its bandwidth, not double it. Without these mechanisms
|
|
the simulator would simply re-confirm the peak-BW roofline; with them,
|
|
it reveals where the real bottlenecks form and which hardware levers
|
|
actually relieve them---which is exactly the question the codesign
|
|
work in this report turns on.
|
|
|
|
|
|
\subsection{Accuracy}
|
|
\label{sec:accuracy}
|
|
|
|
The model is precise about the effects that
|
|
dominate kernel latency on this class of hardware: per-edge bandwidth
|
|
occupancy and flit-level serialization, HBM pseudo-channel parallelism,
|
|
and per-component switching overhead. Two independent cross-checks
|
|
drawn from the experiments in this report confirm that this precision
|
|
translates into physically reasonable kernel latencies.
|
|
|
|
First, in the GEMM study (\S\ref{sec:gemm}), simulator-measured MAC
|
|
efficiency tracks an analytic ideal-pipeline model within
|
|
\SI{2.2}{ppt} across every swept shape---from a single-tile
|
|
$M{=}K{=}N{=}32$ at $\sim\SI{7.5}{\percent}$ up to the deep-$K$
|
|
$K{=}3072$ case at $\sim\SI{88}{\percent}$. The residual gap is
|
|
fill/tail overhead the closed-form pipeline model omits.
|
|
|
|
Second, in the all-reduce study (\S\ref{sec:allreduce},
|
|
Fig.~\ref{fig:allreduce-cmp}), simulator latency for a 2D-torus over
|
|
six devices follows the expected startup-plus-per-packet shape across
|
|
the entire payload sweep---tight at small payloads where startup
|
|
dominates, and within a single-digit multiplicative factor at the
|
|
largest payloads, where the residual gap is explained by per-router
|
|
switching the analytic shape elides.
|
|
|
|
Third, the simulator's per-traversal behaviour can be probed
|
|
directly. Running \texttt{kernbench probe} on the modelled topology
|
|
issues a sequence of PE-to-HBM DMA reads at progressively greater
|
|
hop distances and reports the per-component overhead, per-edge
|
|
serialization, and per-PC drain that the model charges
|
|
(Table~\ref{tab:probe-pe-dma}). Three properties stand out. First,
|
|
the reported latency increases \emph{monotonically} with hop
|
|
count---from \SI{141}{\nano\second} at the local HBM slice to
|
|
\SI{677}{\nano\second} at the worst-case remote-CUBE slice---with the
|
|
increment per added hop matching the per-router overhead and the
|
|
UCIe-link cost in Table~\ref{tab:hw}. Second, the bandwidth-saturation
|
|
curves track the wire's serialization model: at \SI{1}{\mebi\byte}
|
|
transfers the local PE DMA reaches \SI{100}{\percent} of the per-edge
|
|
bandwidth limit, while the longest cross-CUBE path saturates at
|
|
\SI{97.8}{\percent}---exactly the gap that the per-edge propagation
|
|
and per-router overhead model would predict. Third, the simulator
|
|
passes the internal-consistency invariants the probe builds in
|
|
(monotonic hop progression; D2H reads at least as long as the
|
|
equivalent H2D writes because of the reverse-path acknowledgement;
|
|
cross-CUBE best less than cross-CUBE worst). These are properties
|
|
that hold \emph{only} when the underlying graph traversal, FIFO
|
|
contention, and propagation models are internally
|
|
self-consistent---a model error in any of them would surface as a
|
|
non-monotonic or under-utilising curve here long before it polluted a
|
|
kernel-level measurement.
|
|
|
|
The known simplifications---FIFO router arbitration (instead of
|
|
round-robin), HBM scheduler without write-buffer reordering, no bank
|
|
conflict, no refresh or thermal effects, and no upstream
|
|
backpressure---are the price of a deterministic, inspectable model.
|
|
These simplifications bound the absolute accuracy, but the agreement
|
|
with both analytic models and the probe's internal-consistency
|
|
checks above indicates that KernBench is sufficiently accurate for
|
|
evaluating the \emph{relative} hardware--software design trade-offs
|
|
(tiling A vs.\ B, topology X vs.\ Y, with vs.\ without composite
|
|
command, mesh vs.\ torus) that are the primary objective of this
|
|
work.
|
|
|
|
|
|
|
|
\subsection{Modeled hardware configuration}
|
|
\label{sec:hw}
|
|
|
|
Table~\ref{tab:hw} summarizes the hardware configuration
|
|
used throughout this report. The intent of this
|
|
configuration is not to model a specific product, but to
|
|
represent a realistic memory-centric accelerator and to
|
|
provide a consistent baseline for evaluating hardware--
|
|
software co-design mechanisms.
|
|
|
|
The compute capability within each CUBE is provisioned
|
|
such that inference workloads can effectively saturate
|
|
the available HBM bandwidth. The aggregate compute
|
|
throughput is therefore balanced against memory-system
|
|
bandwidth rather than being intentionally over- or
|
|
under-provisioned. Concretely, 8 PEs $\times$
|
|
\SI{8}{\tera\flop\per\second} give roughly
|
|
\SI{64}{\tera\flop\per\second} of f16 compute per CUBE
|
|
against \SI{2048}{\giga\byte\per\second} of HBM bandwidth,
|
|
which places the balanced point at about
|
|
$\sim$31~FLOP/byte; inference decode kernels typically
|
|
operate well below that, so HBM bandwidth---not compute---is
|
|
the natural ceiling on this configuration. For reference, the
|
|
GQA decode kernels evaluated in \S\ref{sec:gqa} operate at
|
|
arithmetic intensity well below this balance point, so their
|
|
ceiling is set by HBM and inter-PE traffic rather than by GEMM
|
|
throughput. HBM bandwidth
|
|
is distributed evenly across the PEs within a CUBE, with
|
|
each PE responsible for servicing approximately one-eighth
|
|
of the CUBE's memory bandwidth through its dedicated HBM
|
|
channels.
|
|
|
|
The on-chip interconnect is configured using bandwidth
|
|
and latency parameters representative of commercially
|
|
available mesh-network IPs. The goal is not to study a
|
|
particular NoC implementation, but rather to evaluate
|
|
kernel behavior under a realistic baseline communication
|
|
fabric.
|
|
|
|
For die-to-die communication, UCIe-A bandwidth
|
|
characteristics are used as the reference point for
|
|
inter-CUBE links. Inter-SIP communication is modeled
|
|
using PCIe-class links. Together, these assumptions
|
|
provide a representative communication hierarchy for
|
|
evaluating collective communication and distributed
|
|
kernel execution.
|
|
|
|
Unless otherwise stated, all experiments use this
|
|
configuration. Workload-specific parameters such as
|
|
matrix dimensions, sequence lengths, and collective
|
|
payload sizes are introduced in their respective sections.
|
|
|
|
\begin{table}[t]
|
|
\centering
|
|
\caption{Modeled hardware configuration}
|
|
\label{tab:hw}
|
|
\small
|
|
\begin{tabular}{@{}ll@{}}
|
|
\toprule
|
|
\textbf{Parameter} & \textbf{Value} \\
|
|
\midrule
|
|
\multicolumn{2}{@{}l}{\emph{Hierarchy}} \\
|
|
SIPs & 2 (1D ring) \\
|
|
CUBEs per SIP & 16 ($4\times4$ mesh) \\
|
|
PEs per CUBE & 8 (4 corners $\times$ 2) \\
|
|
PEs total & 256 \\
|
|
\midrule
|
|
\multicolumn{2}{@{}l}{\emph{Processing element (PE)}} \\
|
|
GEMM engine peak & \SI{8}{\tera\flop\per\second} (f16) \\
|
|
TCM (on-PE) & \SI{16}{\mega\byte}, \SI{512}{\giga\byte\per\second} R/W \\
|
|
\quad kernel scratch & \SI{1}{\mega\byte} \\
|
|
DMA engines & 1 read + 1 write \\
|
|
\textsf{PE\_CPU} fixed cost & \SI{2}{\nano\second} \\
|
|
\textsf{PE\_SCHED} fixed cost & \SI{1}{\nano\second} \\
|
|
\midrule
|
|
\multicolumn{2}{@{}l}{\emph{Memory (per CUBE)}} \\
|
|
HBM capacity & \SI{48}{\giga\byte} (8 slices) \\
|
|
HBM aggregate BW & \SI{2048}{\giga\byte\per\second} \\
|
|
HBM pseudo-channels & 64 (8 per PE), \SI{32}{\giga\byte\per\second} each \\
|
|
SRAM (shared) & \SI{32}{\mega\byte}, \SI{128}{\giga\byte\per\second} link \\
|
|
HBM burst & \SI{256}{\byte} \\
|
|
\midrule
|
|
\multicolumn{2}{@{}l}{\emph{Interconnect}} \\
|
|
Intra-CUBE NoC link & \SI{256}{\giga\byte\per\second}, \SI{2}{\nano\second}/router \\
|
|
Inter-CUBE (UCIe PHY) & \SI{512}{\giga\byte\per\second}, \SI{8}{\nano\second}, XY routing \\
|
|
Inter-SIP (PCIe) & \SI{768}{\giga\byte\per\second} per endpoint \\
|
|
\midrule
|
|
\multicolumn{2}{@{}l}{\emph{Command-issue cost model (defaults)}} \\
|
|
FIXED per command & 40 cycles \\
|
|
per-byte rate $R$ & 0.0625 cycles/byte (\SI{16}{\byte\per\cycle}) \\
|
|
composite size cap & \SI{1024}{\byte} \\
|
|
\bottomrule
|
|
\end{tabular}
|
|
\end{table}
|