paper(gemm): inline §3.1 pipeline-stage chain; drop display equation
The five-stage chain was set as a centered display equation, which overflowed the column and clipped the last 'DMA_WRITE' label. The inline prose right below it also repeated the same chain in a shorter form. Consolidate into one inline parenthetical listing the five stage names, and reuse 'flows through the chain' for the self-routing description — no clipping, no repetition. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
Binary file not shown.
@@ -19,20 +19,17 @@ in command overhead?
|
|||||||
|
|
||||||
\subsection{Design}
|
\subsection{Design}
|
||||||
|
|
||||||
The answer is the \emph{composite command}. A single command carries the
|
The answer is the \emph{composite command}. A single command carries
|
||||||
ordered tile pipeline
|
the ordered five-stage tile pipeline (\textsf{DMA\_READ},
|
||||||
\[
|
\textsf{FETCH}, \textsf{GEMM}, \textsf{STORE}, \textsf{DMA\_WRITE}),
|
||||||
\textsf{DMA\_READ}\rightarrow\textsf{FETCH}\rightarrow\textsf{GEMM}
|
|
||||||
\rightarrow\textsf{STORE}\rightarrow\textsf{DMA\_WRITE},
|
|
||||||
\]
|
|
||||||
and the PE scheduler splits the payload into hardware tiles
|
and the PE scheduler splits the payload into hardware tiles
|
||||||
(here $32\times64\times32$), emitting one tile token per tile. Subsequent
|
(here $32\times64\times32$), emitting one tile token per tile.
|
||||||
stages are reached by \emph{token self-routing} between the on-PE engines,
|
Subsequent stages are reached by \emph{token self-routing} between
|
||||||
so a tile flows DMA\,$\rightarrow$\,fetch\,$\rightarrow$\,GEMM\,$%
|
the on-PE engines, so a tile flows through the chain without
|
||||||
\rightarrow$\,store without returning to the scheduler between stages.
|
returning to the scheduler between stages. Because the whole
|
||||||
Because the whole pipeline is described by one command, the issue cost is
|
pipeline is described by one command, the issue cost is paid once
|
||||||
paid once per GEMM rather than once per tile-stage, and the scheduler is
|
per GEMM rather than once per tile-stage, and the scheduler is free
|
||||||
free to keep every stage busy on different tiles simultaneously---tile
|
to keep every stage busy on different tiles simultaneously---tile
|
||||||
$i$'s GEMM overlaps tile $i{+}1$'s DMA read. A multi-operation composite
|
$i$'s GEMM overlaps tile $i{+}1$'s DMA read. A multi-operation composite
|
||||||
additionally lets an epilogue (for example a vector-math step) ride the
|
additionally lets an epilogue (for example a vector-math step) ride the
|
||||||
same tile loop, firing per $K$-tile, per output tile, or once per kernel
|
same tile loop, firing per $K$-tile, per output tile, or once per kernel
|
||||||
|
|||||||
Reference in New Issue
Block a user