paper: add /paper skill + 1H HW-SW codesign report (GEMM, All-Reduce, fused GQA)

New `/paper` slash-command skill that synthesizes ADR/SPEC content and live
KernBench benchmark results into a sectioned LaTeX technical paper compiled
to PDF with Tectonic (auto-installed). The skill negotiates a TOC, grounds
every number in committed artifacts or fresh bench runs, and keeps
report-only benches isolated.

This commit also includes the first generated report:
- docs/report/1H-codesign-paper/ — main.tex + per-section .tex, figures,
  toc.md contract, and the built 8-page main.pdf. Covers the platform
  (source-level kernels, latency model + accuracy, HW config from
  topology.yaml), GEMM via composite command, All-Reduce via PE_IPCQ, and
  fused GQA combining both, plus discussion/conclusion/2H future work.
- scripts/paper/ — isolated report harnesses (not registered benches):
  paper_gqa_latency.py harvests per-panel GQA end-to-end latency + engine
  occupancy (the milestone only emitted op-counts); paper_plot_gqa.py
  renders the GQA figures.

GEMM/All-Reduce reuse committed milestone figures/CSVs; GQA results are
generated fresh. Honest flags retained: PE_CPU dispatch cost is 0 in this
config, and the proposed two-composite softmax_merge decode is marked
designed-not-measured.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-06-10 22:15:14 -07:00
parent 4a55ae5c0b
commit dd525bfcb7
25 changed files with 1561 additions and 0 deletions
+417
View File
@@ -0,0 +1,417 @@
---
description: Generate the 1H HW-SW codesign technical paper (PDF) — propose/confirm a TOC, then synthesize ADRs + live KernBench benchmark results into a sectioned LaTeX paper compiled with Tectonic.
---
# `/paper` — HW-SW Codesign Technical Report Generator
Produces a **technical-paper-style PDF** at
`docs/report/1H-codesign-paper/build/main.pdf`, synthesizing the
ADR corpus, SPEC.md, and **live KernBench benchmark results** into a
publication-shaped report on the three 1H optimizations:
1. **GEMM Acceleration** using the composite command
2. **All-Reduce Acceleration** using PE_IPCQ
3. **Fused Grouped Query Attention** using both composite command and PE_IPCQ
plus an explanation of the simulation platform itself (why KernBench
exists, its execution model, and how it computes latency / how accurate
that is).
Audience is the **internal team** (engineers + architects), but the
report reads like a real technical paper: **do not cite `ADR-NNNN` or
`SPEC §X.Y` anywhere** — not in prose, not in a reference list. ADRs and
SPEC are *grounding sources only* (use their content; never name them).
The **References** section, if used, lists **external literature only**
(e.g. FlashAttention, Megatron-LM, GPT-3, Llama 3) — actual papers, not
internal docs.
---
## Operating principles (read first)
- **Confirm the TOC before writing.** The first thing `/paper` does is
present the outline (saved `toc.md` if present, else the default
below) and **stop for the user's confirmation / adjustments**. Do not
generate sections until the user approves the TOC.
- **Ground every number.** Every quantitative result in prose, tables,
or captions MUST trace to (a) a committed bench artifact
(`*.csv` / `*.json` / `*.png` under `1H_milestone_output/` or
`docs/diagrams/`), (b) a freshly-run bench output, or (c) a specific
ADR/SPEC quotation. **Never invent latencies, speedups, or
utilizations.** If a needed number does not exist, generate it via a
bench run (below) or state the gap — do not guess.
- **Reuse existing results by default.** The three milestone benches
already have committed outputs. Do **not** re-run a bench unless the
user asks to regenerate, or the artifact is missing/stale.
- **Design / proposal status honesty.** Some designs are `Proposed`,
not yet implemented (e.g. ADR-0061, ADR-0065). Present these as
proposed and clearly separate "modeled & measured" from
"designed, not yet measured." Do not report measurements for
unimplemented paths.
- **English prose.** Architectural terms (composite command, PE_IPCQ,
GQA, TCM, SIP, CUBE, PE) used as-is.
---
## Invocation
- `/paper`**default**. Resolve the period, present the TOC
(`toc.md` if it exists, else the default TOC), and **stop** for
confirmation. Once the user approves (possibly after editing the TOC
in conversation), persist `toc.md` and proceed to full build.
- `/paper toc` — only (re)negotiate and persist the TOC. No sections,
no build.
- `/paper build` — assume the saved `toc.md` is approved; run the full
pipeline (sections → figures → compile PDF) without re-confirming.
- `/paper section <id>` — regenerate a single section (`<id>` = a
filename stem under `sections/`, e.g. `04-allreduce`) and recompile.
- `/paper figures` — (re)generate only the benchmark figures/data
(runs/refreshes benches), no prose changes, then recompile.
- `/paper pdf` — only recompile the existing `.tex` to PDF (Tectonic).
If `toc.md` does not yet exist, any mode except `/paper`/`/paper toc`
first falls back to presenting the default TOC for confirmation.
---
## Output layout
```
docs/report/1H-codesign-paper/
toc.md # the agreed outline (persisted; human-editable)
main.tex # preamble + \input of each section
refs.bib # ADR/SPEC bibliography (optional; or inline refs)
sections/
00-exec-summary.tex
01-introduction.tex
02-platform.tex # includes 3.4 Modeled Hardware Configuration
03-gemm.tex
04-allreduce.tex
05-gqa.tex
06-discussion.tex
07-conclusion.tex
08-future-work.tex
figures/ # PNGs copied in from bench outputs (portable)
build/ # Tectonic output; main.pdf lives here
```
Report-specific benchmark/plot harnesses (when a figure must be
generated, see *Generating missing results*) live **isolated** under:
```
scripts/paper/ # report-only harnesses — NOT registered benches
```
This directory is intentionally separate from `src/kernbench/benches/`
so the report's harnesses never collide with other people's benches
(the bench registry only audits `src/kernbench/benches/`). Treat
`scripts/paper/*` and everything under `docs/report/1H-codesign-paper/`
as **derived artifacts** (like `docs/diagrams/`): creating/updating them
does not require Phase 2 approval, but they MUST stay consistent with
SPEC.md and ADRs.
---
## Default TOC
If `toc.md` is absent, propose this and ask the user to confirm or edit:
1. **Executive Summary** — the attention-optimization goal, the two
enabling optimizations (GEMM composite command + PE_IPCQ communication),
the fused GQA capstone, the headline results, and the bottom-line
recommendation on which HW changes are worth keeping. May run a
paragraph or two longer than a terse abstract; front-loads conclusions.
2. **Introduction** — the 1H focus is **attention-kernel optimization**:
FlashAttention-style tiling and Grouped Query Attention (GQA), building
on prior work on Multi-Head Attention (MHA, studied earlier — reference
it as the established baseline, not re-derived here). Optimizing a fused
attention kernel requires two enabling optimizations, each studied in
its own right and then **combined inside the fused kernel**:
(a) **GEMM optimization** (the composite command) and
(b) **communication optimization** (PE_IPCQ). Frame the three result
sections as *two enablers → the fused GQA capstone that uses both*.
Motivate why HW-SW codesign (not software alone) is required to realize
the gains.
3. **The KernBench Platform**
1. *Why KernBench* — source-level kernel execution with **no compiler /
software-stack dependency**; isolate algorithm-level optimization from
other layers so a kernel can fully exploit the hardware without E2E
cross-layer entanglement.
2. *Execution model* — discrete-event graph; `runtime API → sim_engine →
components`; PE pipeline + composite commands; 2-pass data/timing model.
3. *Latency model & accuracy* — traversal-based golden invariants
(every hop pays > 0, deterministic, explicit connectivity); the
structural CPU-dispatch cost model; per-edge BW occupancy, HBM
pseudo-channel parallelism, flit streaming; known simplifications &
what calibration anchors the constants.
4. *Modeled Hardware Configuration* — the shared, platform-wide
quantitative config used across **all** experiments, read from
`topology.yaml`: SIP/CUBE/PE hierarchy (counts & nesting), PE clock &
compute throughput, TCM/SRAM/HBM capacities & bandwidths, NOC /
inter-CUBE / inter-SIP link bandwidths & latencies, and the cost-model
constants. Per-experiment workload parameters (GEMM shapes, All-Reduce
topologies, GQA seq/head/user counts) stay in each evaluation section,
not here.
4. **GEMM Acceleration via Composite Command** — necessity · design · results · analysis.
5. **All-Reduce Acceleration via PE_IPCQ** — necessity · design · results · analysis.
6. **Fused Grouped Query Attention** (composite + PE_IPCQ) — necessity · design · results · analysis.
7. **Discussion** — which HW changes are meaningful, and why (cross-cutting).
8. **Conclusion**.
9. **Future Work — 2H** — add the FFN/MoE layer toward full LLM decoding;
analyze how compute & data should be distributed for agentic / MoE
workloads.
10. **References** *(optional)* — **external literature only** (e.g.
FlashAttention, Megatron-LM, GPT-3, Llama 3). **No ADR/SPEC entries.**
Each of §4/§5/§6 follows the same four-beat structure:
**(a) why the optimization is needed**, **(b) the design (how it works)**,
**(c) experimental results** (run/plot/measure), **(d) analysis & meaning**.
---
## Source map (ADR / SPEC → section)
Use these as **grounding sources only** — read them for content, but
**never name an ADR or SPEC in the paper**. Verify each quantitative
claim against the actual file before asserting it (CLAUDE.md: no
asserting architectural numbers from memory).
- **Modeled Hardware Configuration (§3.4)**: `topology.yaml`
(authoritative for counts, clocks, BW, link params, cost-model
constants) is the **primary** source — read it directly and build the
HW table from it. Cross-check against ADR-0003 (hierarchy/scope),
ADR-0004 (memory BW guarantee), ADR-0033 (latency model), ADR-0064
(cost-model defaults). If `topology.yaml` and an ADR disagree, the
loaded `topology.yaml` is what the sim actually runs — report that
value and flag the mismatch.
- **Platform — why/execution/latency**: `SPEC §0`, `SPEC §0.1`
(golden invariants), `SPEC §1 R1/R2/R3`, `SPEC §2.1` (graph execution);
ADR-0003 (hierarchy/scope), ADR-0007 (runtime↔sim_engine boundary),
ADR-0014 (PE pipeline / composite), ADR-0020 (2-pass data/timing),
ADR-0033 (latency model assumptions & simplifications),
ADR-0064 (structural CPU-dispatch cost model: `FIXED + logical_bytes×R`).
- **GEMM / composite**: ADR-0014 (composite command D3.2/D3.3),
ADR-0042 (tile-plan generators), ADR-0044 (GEMM eval harness),
ADR-0064 (cost model — composite's per-command BW reduction),
ADR-0065 (flat-ops `CompositeCmd`, *Proposed*).
- **All-Reduce / PE_IPCQ**: ADR-0023 (PE_IPCQ — control plane, ring
buffers, VC split), ADR-0025 (direction addressing), ADR-0032
(intercube all-reduce), ADR-0043 (allreduce eval harness),
ADR-0047 (AHBM CCL backend), ADR-0050 (CCL algorithm contract).
- **GQA / fused attention**: ADR-0060 (fused GQA kernel),
ADR-0061 (`tl.broadcast`/`tl.repeat`, *Proposed*),
ADR-0062 (lazy `tl.load`), ADR-0063 (`tl.scratch_scope`),
ADR-0064 (cost model), ADR-0065 (flat-ops composite + `softmax_merge`,
*Proposed*). KO mirrors in `docs/adr-ko/`; DDD-0060/0065 design docs if present.
- **Eval/milestone pattern**: ADR-0054 (self-contained milestone benches);
ADR-0045 (bench contract).
---
## Result map (bench → artifact → section) + run commands
All runs use the repo's `topology.yaml` unless a section needs a
specific topology. Reuse committed artifacts unless told to regenerate.
### §4 GEMM — `milestone-1h-gemm`
- Artifacts (committed): `src/kernbench/benches/1H_milestone_output/gemm/`
→ `gemm_sweep.json`, `gemm_stage_breakdown.png`,
`gemm_mac_utilization_measured.png`,
`gemm_mac_utilization_theoretical_vs_measured.png`.
- Regenerate: `kernbench run --topology topology.yaml --bench milestone-1h-gemm`
(full 24-shape sweep ≈ minutes; `MILESTONE_FAST=1` reuses committed JSON).
### §5 All-Reduce — `milestone-1h-ccl`
- Artifacts (committed): `src/kernbench/benches/1H_milestone_output/ccl/`
→ `summary.csv`, `topology.png`, per-topology
`AllReduce_LRAB_*` PNGs, buffer-kind sweep
`AllReduce_LRAB_2Dtorus_..._with_TCM_SRAM_HBM.{png,csv}`,
`comparison_mesh_vs_ring_vs_2DTorus_vs_theoretical_vs_fsim.png`.
- Regenerate: `kernbench run --topology topology.yaml --bench milestone-1h-ccl`.
### §6 GQA — `milestone-gqa-headline`
- Artifact (committed): `src/kernbench/benches/1H_milestone_output/gqa_headline/sweep.json`.
**Caveat:** this currently holds only `op_log_summary` **op-counts**
(`gemm_count`, `ipcq_copy_count`, `dma_read_count`, `dma_write_count`)
for 4 panels (single/multi-user × prefill/decode) — **no latency,
no figures**.
- Regenerate counts:
`GQA_HEADLINE_RUN=1 kernbench run --topology topology.yaml --bench milestone-gqa-headline`.
- For a real GQA **results** section, generate latency + figures via an
isolated report harness — see next.
---
## Generating missing results (isolation rule)
When a section needs an artifact that does not yet exist (today: GQA
latency and any GQA figure), create a **report-dedicated harness** under
`scripts/paper/` — never under `src/kernbench/benches/`. Rules:
- One file per concern, prefixed `paper_`, e.g.
`scripts/paper/paper_gqa_latency.py` (drives the engine, emits
latency + a baseline-vs-optimized comparison),
`scripts/paper/paper_plot_gqa.py` (renders figures from that data).
- Follow the **ADR-0054 self-contained pattern**: the harness builds its
own `GraphEngine`(s), imports the existing GQA kernel modules
(`_gqa_attention_{prefill,decode}_{short,long}.py`,
`milestone_gqa_headline.py`), runs them, and reads latency from the
engine's completion timestamps / op log. It does **not** register a
bench and does **not** import or mutate other people's benches.
- Output figures land in `docs/report/1H-codesign-paper/figures/` and,
if useful as a shared diagram, mirror into `docs/diagrams/gqa_plots/`.
- **Only measure what the current code actually runs.** ADR-0065's GQA
opt2 (`softmax_merge`, two-composite decode) is *Proposed*; if it is
not wired in production, present it as designed-not-measured and base
the measured results on the implemented path. Verify implementation
status by reading the code before claiming a measurement.
- These harnesses are derived-artifact tooling (no Phase 2 needed). **But**
if the report ever requires changing a *registered production bench*
under `src/kernbench/benches/` or any production module, that is a
non-trivial production change — STOP and follow CLAUDE.md Phase 1 →
approval → Phase 2.
---
## Procedure
### Step 0 — Toolchain (Tectonic)
Ensure a working Tectonic before any compile step. In order:
1. If `tectonic` is on PATH (`tectonic --version` succeeds), use it.
2. Else try, in order, whichever succeeds (capture output, don't hang):
- `winget install --id TectonicProject.Tectonic --silent --accept-package-agreements --accept-source-agreements`
- `scoop install tectonic`
- Direct download (no admin): fetch the latest Windows `x86_64-pc-windows-msvc`
release zip from `https://github.com/tectonic-typesetting/tectonic/releases`
via PowerShell `Invoke-WebRequest`, expand it, and place `tectonic.exe`
under `docs/report/1H-codesign-paper/build/.tools/`. Use that absolute
path for compilation.
3. Re-verify with `tectonic --version`. If all methods fail, write the
`.tex` files anyway, **report the failure**, and tell the user the
exact command to compile manually (or to use Overleaf). Do not block
the rest of the pipeline.
Tectonic fetches LaTeX packages on first compile (needs network once).
### Step 1 — Period + TOC
- Period from system date: month 16 → `{YYYY}-1H`, else `{YYYY}-2H`.
(Report dir is fixed at `1H-codesign-paper` for this half; adjust the
title page text with the period.)
- If `toc.md` exists, present it; else present the **Default TOC**.
- **Stop and ask the user to confirm or edit.** Incorporate edits.
Persist the agreed outline to `toc.md`. (In `/paper build` mode, skip
the stop and trust the saved `toc.md`.)
### Step 2 — Platform section (§3, incl. §3.4 HW config)
Write `sections/02-platform.tex` from the platform source map, covering
all four sub-parts:
- §3.13.3 (why / execution model / latency model + accuracy). Include
the cost-model formula `dispatch_cycles = FIXED + logical_bytes×R` with
the committed defaults (verify against ADR-0064 before quoting), and
explicitly address **how accurate** the latency is: what is modeled
precisely (per-edge BW occupancy, HBM pseudo-channels, flit streaming,
per-component overhead) and the known simplifications.
- §3.4 *Modeled Hardware Configuration*: read `topology.yaml` directly
and build a compact, table-driven summary — SIP/CUBE/PE hierarchy
(counts & nesting), PE clock & compute throughput, TCM/SRAM/HBM
capacity & bandwidth, NOC / inter-CUBE / inter-SIP link bandwidth &
latency, and the cost-model constants (`FIXED_PER_CMD`, `R`, composite
cap). Every number from `topology.yaml`. This is the shared config;
later sections say "the modeled hardware" and only add their own
workload params.
Do not name the source ADRs/SPEC in the prose.
### Step 3 — The three result sections (§4 GEMM, §5 All-Reduce, §6 GQA)
For each, in order, produce the four beats:
- **(a) Necessity** — from the ADR Context: the bottleneck the
optimization targets (e.g. per-command dispatch overhead for GEMM;
collective comm cost / HoL blocking for All-Reduce; KV-load-bound
decode + softmax merge for GQA).
- **(b) Design** — from the ADR Decisions: composite command tile
pipeline (ADR-0014); PE_IPCQ control/data split + ring buffers + VC
(ADR-0023); fused attention with online-softmax, lazy load, scratch
recycling (ADR-0060/0062/0063). Note Proposed-vs-Accepted status.
- **(c) Results** — reuse committed artifacts (GEMM, All-Reduce). For
GQA, generate latency + figures via the isolated `scripts/paper/`
harness (above). Copy chosen PNGs into `figures/`, `\includegraphics`
them, and build result tables from the CSV/JSON — quoting actual
numbers, with a one-line note of the run command and source artifact.
- **(d) Analysis** — interpret: what the curve/table shows, why the HW
feature produces it, where it saturates or fails to help, and the
cost (area/complexity) implied. Tie back to the cost model.
### Step 4 — Discussion, Conclusion, Future Work (§6/§7/§8)
- **Discussion**: synthesize across the three — which HW changes
(composite command, PE_IPCQ, the cost-model-exposed BW reduction)
carry their weight, and under what regimes.
- **Conclusion**: the codesign thesis, supported by the measured
results; state plainly which HW changes are meaningful.
- **Future Work (2H)**: FFN / MoE layer toward full LLM decoding; how
compute & data should be distributed for agentic / MoE workloads
(expert routing, token dispersion, all-to-all vs all-reduce, capacity).
Frame as the next codesign questions, not claims.
### Step 5 — Assemble & compile
- `main.tex`: article/IEEEtran-style preamble, title with the period,
author, executive-summary input, then `\input{sections/...}` in TOC
order, then references. Keep the preamble minimal (graphicx, booktabs,
hyperref, amsmath) so Tectonic compiles cleanly.
- Compile: `tectonic main.tex --outdir build` (or the resolved Tectonic
path) from `docs/report/1H-codesign-paper/`. Re-run if references need
a second pass. Confirm `build/main.pdf` exists and report its size.
### Step 6 — Chat report
Emit to chat (not to any file):
```
## /paper — Build Summary
**PDF:** docs/report/1H-codesign-paper/build/main.pdf (<N> pages, <size>)
**Toolchain:** tectonic <version | how installed>
**Sections written:** <list>
**Benches: reused vs regenerated**
- GEMM: <reused committed | regenerated via milestone-1h-gemm>
- All-Reduce: <reused | regenerated>
- GQA: <generated via scripts/paper/paper_gqa_latency.py | counts-only>
**Figures included:** <list of figures/*.png>
**Grounding check:** every result number traced to <artifact list>
**Proposed-not-measured items flagged:** <e.g. ADR-0065 opt2, ADR-0061>
**Open gaps / suggested next runs:** <…>
```
---
## Constraints (do not violate)
1. **TOC first.** Never write sections before the TOC is confirmed
(default mode). `toc.md` is the contract.
2. **No fabricated numbers.** Every quantitative claim traces to a
committed artifact, a fresh bench run, or an ADR/SPEC quote. Verify
architectural constants against the source file, not memory.
3. **Reuse by default.** Do not re-run a bench unless asked or the
artifact is missing/stale.
4. **Isolation.** Report harnesses live only in `scripts/paper/`; never
register a bench, never touch other people's benches.
5. **Production-change gate.** Any change to a registered bench or
production module triggers CLAUDE.md Phase 1 → approval → Phase 2.
Report tooling and `docs/report/...` are derived artifacts (exempt).
6. **Status honesty.** Mark Proposed designs as proposed; never present
an unimplemented path as measured.
7. **No internal references.** Never name `ADR-NNNN` or `SPEC §X.Y` in
prose, captions, or the reference list — they are grounding sources
only. A References section, if present, holds **external literature
only**. Hardware numbers come from `topology.yaml`.
8. **Determinism.** Sections in TOC order; same inputs → same paper.
9. **English prose**, internal-audience tone.
```
+2
View File
@@ -0,0 +1,2 @@
# Downloaded Tectonic toolchain binary (large, machine-local; re-fetched by /paper)
build/.tools/
Binary file not shown.
Binary file not shown.

After

Width:  |  Height:  |  Size: 75 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 86 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 194 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 40 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 46 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 42 KiB

@@ -0,0 +1,97 @@
{
"version": 1,
"panels": [
"single_user_prefill_gqa",
"multi_user_prefill_gqa",
"single_user_decode_gqa",
"multi_user_decode_gqa"
],
"rows": [
{
"panel": "single_user_prefill_gqa",
"kind": "prefill",
"C": 1,
"S_kv": 16,
"latency_ns": 445.1180000000004,
"op_log_summary": {
"gemm_count": 2,
"ipcq_copy_count": 0,
"dma_read_count": 3,
"dma_write_count": 1
},
"engine_occupancy_ns": {
"pe_gemm": 2.048000000000002,
"pe_math": 5.0,
"pe_dma": 72.0,
"pe_fetch_store": 0,
"pe_ipcq": 0,
"pe_cpu": 0
}
},
{
"panel": "multi_user_prefill_gqa",
"kind": "prefill",
"C": 4,
"S_kv": 16,
"latency_ns": 4630.408000000019,
"op_log_summary": {
"gemm_count": 32,
"ipcq_copy_count": 24,
"dma_read_count": 12,
"dma_write_count": 4
},
"engine_occupancy_ns": {
"pe_gemm": 8.192000000000917,
"pe_math": 236.0,
"pe_dma": 1446.945000000003,
"pe_fetch_store": 0,
"pe_ipcq": 0,
"pe_cpu": 0
}
},
{
"panel": "single_user_decode_gqa",
"kind": "decode",
"C": 1,
"P": 8,
"S_kv": 64,
"latency_ns": 3631.730500000015,
"op_log_summary": {
"gemm_count": 16,
"ipcq_copy_count": 21,
"dma_read_count": 24,
"dma_write_count": 1
},
"engine_occupancy_ns": {
"pe_gemm": 16.383999999998196,
"pe_math": 160.0,
"pe_dma": 1450.1150000000016,
"pe_fetch_store": 0,
"pe_ipcq": 0,
"pe_cpu": 0
}
},
{
"panel": "multi_user_decode_gqa",
"kind": "decode",
"C": 4,
"P": 8,
"S_kv": 128,
"latency_ns": 6692.578999999866,
"op_log_summary": {
"gemm_count": 64,
"ipcq_copy_count": 93,
"dma_read_count": 96,
"dma_write_count": 1
},
"engine_occupancy_ns": {
"pe_gemm": 32.76799999998184,
"pe_math": 688.0000000000009,
"pe_dma": 15919.751500000013,
"pe_fetch_store": 0,
"pe_ipcq": 0,
"pe_cpu": 0
}
}
]
}
Binary file not shown.

After

Width:  |  Height:  |  Size: 38 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 62 KiB

+40
View File
@@ -0,0 +1,40 @@
\documentclass[10pt,twocolumn]{article}
\usepackage[margin=0.75in]{geometry}
\usepackage{graphicx}
\usepackage{booktabs}
\usepackage{amsmath}
\usepackage{amssymb}
\usepackage{siunitx}
\DeclareSIUnit\flop{FLOP}
\DeclareSIUnit\cycle{cycle}
\DeclareSIUnit\byte{B}
\usepackage{xcolor}
\usepackage{hyperref}
\hypersetup{colorlinks=true,linkcolor=blue!50!black,citecolor=blue!50!black,urlcolor=blue!50!black}
\usepackage{caption}
\captionsetup{font=small,labelfont=bf}
\usepackage{microtype}
\graphicspath{{figures/}}
\title{\textbf{Hardware--Software Codesign for Attention-Centric\\
LLM Kernels: GEMM, Collective Communication, and\\
Fused Grouped-Query Attention}}
\author{KernBench Team}
\date{2026 First Half (1H)}
\begin{document}
\maketitle
\input{sections/00-exec-summary}
\input{sections/01-introduction}
\input{sections/02-platform}
\input{sections/03-gemm}
\input{sections/04-allreduce}
\input{sections/05-gqa}
\input{sections/06-discussion}
\input{sections/07-conclusion}
\input{sections/08-future-work}
\end{document}
@@ -0,0 +1,39 @@
\section*{Executive Summary}
\addcontentsline{toc}{section}{Executive Summary}
This report studies how to make attention-centric LLM kernels fast through
hardware--software codesign, using KernBench, a discrete-event simulator
that runs kernels at the source level---without a compiler or other
software-stack dependency---so that algorithm-level optimizations and the
hardware features that support them can be isolated and measured.
The 1H work centers on attention. Building on prior multi-head-attention
work, we target FlashAttention-style Grouped-Query Attention (GQA) and find
that a fast fused attention kernel rests on two enabling optimizations,
each studied in its own right and then combined inside the fused kernel:
a \textbf{composite command} that issues a tiled GEMM as a single,
self-routing pipeline, and \textbf{PE\_IPCQ}, a per-PE on-device collective
engine for all-reduce.
The measured results form a single, consistent story. For GEMM, the
composite command lets compute-rich shapes reach
\textasciitilde\SI{78}{\percent} of peak MAC efficiency, with the simulator's
measured efficiency tracking the analytic prediction within
\SIrange{10}{20}{\percent}---both a performance result and a validation of
the latency model. For all-reduce, PE\_IPCQ delivers on-device collectives
whose latency tracks the interconnect's physical limits; a 2D-torus
inter-device fabric is \SI{20}{}--\SI{25}{\percent} faster than a mesh or
ring at scale, and staging the collective in on-PE TCM rather than HBM or
SRAM saves a further \SI{14}{}--\SI{38}{\percent}. For fused GQA, the two
enablers combine: because the composite command keeps the MAC array nearly
idle, the kernel is exposed as overwhelmingly data-movement bound---the DMA
engine outweighs the compute engines by two to three orders of magnitude,
and on-device IPCQ reduction traffic scales with users and devices.
\textbf{Bottom line.} For an attention-dominated decoder, the hardware
changes worth keeping are the ones that move and reduce data faster: the
single-command self-routing GEMM pipeline (which makes the MAC array
usable and exposes the real bottleneck), the on-device PE\_IPCQ collective
with its compute/communication virtual-channel split, fast on-PE staging
memory, and wrap-around (torus) inter-device links. Additional raw MAC
throughput is not the limiting factor for this workload.
@@ -0,0 +1,49 @@
\section{Introduction}
\label{sec:intro}
The performance of a large language model on an accelerator is decided less
by peak FLOPS than by how well the kernel exploits the memory system and
the interconnect. This is especially true of attention, which in the
decode phase reads a large KV cache to do a small amount of arithmetic.
Optimizing such kernels is fundamentally a \emph{codesign} problem: the
algorithm (how to tile, fuse, and reduce) and the hardware (what issue
mechanism, what collective engine, what memory hierarchy) have to be
designed against each other. Optimizing only the software leaves the
hardware idle; adding hardware that the software cannot reach is wasted
area.
\paragraph{The 1H focus: attention.} This half's work targets attention.
Multi-head attention (MHA) was studied previously and is taken here as the
established baseline. The new work is FlashAttention-style tiling and
Grouped-Query Attention (GQA)---the form used by modern long-context
decoders, in which several query heads share a single KV head to shrink the
KV cache. Realizing a fast \emph{fused} GQA kernel, however, is not a
single optimization. It decomposes into two enabling optimizations that we
studied separately and then brought together inside the fused kernel:
\begin{enumerate}
\item \textbf{GEMM optimization} via a composite command
(\S\ref{sec:gemm}): a way to issue a tiled matrix multiply as one
self-routing pipeline so the MAC array stays fed without per-tile command
overhead. Attention's $Q\!\cdot\!K^{\top}$ and $P\!\cdot\!V$ products are
exactly such GEMMs.
\item \textbf{Communication optimization} via PE\_IPCQ
(\S\ref{sec:allreduce}): a per-PE on-device collective engine that
performs all-reduce overlapped with compute. Attention's multi-user and
sequence-parallel KV reductions are exactly such collectives.
\end{enumerate}
The report therefore reads as \emph{two enablers building to one capstone}:
GEMM (\S\ref{sec:gemm}) and all-reduce (\S\ref{sec:allreduce}) are
developed and measured on their own, and then fused GQA
(\S\ref{sec:gqa}) shows them operating together. Before the results,
\S\ref{sec:platform} describes the KernBench platform---why a
software-stack-independent, source-level simulator is the right tool for
this question, how it executes a kernel, how it computes latency and how
accurate that is, and the exact hardware configuration used throughout.
We close with a cross-cutting discussion of which hardware changes are
worth their cost (\S\ref{sec:discussion}), a conclusion
(\S\ref{sec:conclusion}), and the 2H agenda
(\S\ref{sec:future}): extending from attention to the feed-forward and
mixture-of-experts layers, and to the compute/data distribution questions
that full LLM decoding and agentic workloads raise.
@@ -0,0 +1,171 @@
\section{The KernBench Platform}
\label{sec:platform}
All results in this report are produced on \emph{KernBench}, a
system-level, discrete-event simulator for LLM kernels running on
SIP-based AI accelerators. This section explains why the platform exists,
how it executes a kernel, how it computes latency and how accurate that
number is, and the concrete hardware configuration used for every
experiment that follows.
\subsection{Why KernBench: source-level kernels without a software stack}
\label{sec:why}
In a production end-to-end (E2E) stack, kernel performance is entangled
with every layer above the hardware: the compiler's tiling and
scheduling choices, the framework's operator dispatch, the collective
library, and the runtime. Good E2E numbers require \emph{all} of those
layers to be co-optimized, which makes it hard to answer a narrower but
more fundamental question: \emph{given the hardware, how fast can a
well-written kernel be, and which hardware features actually make it
faster?}
KernBench is built to answer exactly that question. Kernels are written
and executed at the \emph{source level}---as algorithmic descriptions in a
small tile-oriented kernel API---with no dependency on a compiler or any
other software-stack layer. The simulator takes the kernel and a hardware
topology and reports the latency that the modeled hardware would deliver.
This isolation is deliberate: it lets us study algorithm-level
optimizations (how to tile a GEMM, how to schedule a collective, how to
fuse an attention kernel) and the hardware features that support them,
without the confound of compiler maturity or framework overhead. The cost
is that KernBench numbers are \emph{not} E2E latencies; they are the
achievable-kernel latencies that an ideal software stack would expose.
\subsection{Execution model}
\label{sec:exec}
KernBench is layered along the flow of a request:
\begin{itemize}
\item The \textbf{runtime API} is host-facing and
topology-agnostic---it deploys tensors and launches kernels but knows
nothing about routing or interconnect.
\item The \textbf{simulation engine} schedules discrete events, routes
every request through the modeled graph, and tracks completion via
correlation IDs.
\item The \textbf{components} are device-side nodes that model hardware
behavior: the per-PE blocks (scheduler, DMA, GEMM and vector-math
engines, TCM, IPCQ), the NoC routers, the HBM controllers, and the
inter-chiplet links.
\end{itemize}
The topology is compiled once at configuration time into an authoritative
graph of components and links; it is never mutated during a run. Within a
PE, work is expressed as \emph{composite commands}: a single command
carries an ordered pipeline of operations
(\textsf{DMA\_READ} $\rightarrow$ \textsf{FETCH} $\rightarrow$
\textsf{GEMM}/\textsf{MATH} $\rightarrow$ \textsf{STORE} $\rightarrow$
\textsf{DMA\_WRITE}) that the PE scheduler tiles and streams. This
composite mechanism is the substrate for the GEMM optimization
(\S\ref{sec:gemm}) and, combined with on-PE collectives, for fused
attention (\S\ref{sec:gqa}). Data and timing are handled in two passes, so
that a kernel's numeric results and its latency are computed
consistently but independently.
\subsection{Latency model and its accuracy}
\label{sec:latency}
KernBench obeys a set of golden invariants that keep its latency numbers
physically meaningful. End-to-end latency is computed \emph{strictly by
explicit traversal} over modeled components and links: every routed
request incurs latency greater than zero, routing is deterministic, and
every valid request flow has explicit connectivity. There are no hidden
shortcuts, implicit waits, or magic delays---if a delay exists in the
result, it came from a scheduled event on a modeled component or link.
Latency accumulates from three kinds of contributions: per-node fixed
overheads (each component carries an \texttt{overhead\_ns}), per-link
transfer time, and per-service occupancy. The interconnect is modeled at
fine granularity. Each directed link serializes traffic through a
bandwidth-limited FIFO, so a busy link delays later flits. Payloads are
decomposed into fixed-size flits (default \SI{256}{\byte} bursts) that
arrive at $\text{prop} + \text{flit\_bytes}/\text{bw}$ intervals, so link
bandwidth throttles arrival rate rather than being applied as a lump sum.
HBM is modeled with per-pseudo-channel parallelism: a stateless array of
channel-availability timestamps with address-based channel selection
captures bank-level concurrency.
The cost of \emph{issuing} a command is modeled structurally rather than
with a per-operation calibration table. The PE control processor charges, per command,
\[
d_{\text{cmd}} = \textsf{FIXED} + b_{\text{logical}} \cdot R,
\]
where $b_{\text{logical}}$ is the command's hardware-logical byte size,
\textsf{FIXED} captures the fixed per-command cost (queue-tail
update, completion registration) and $R$ captures the per-byte cost of
serializing the command descriptor into the scheduler queue. The default
anchoring (\textsf{FIXED} $=40$ cycles, $R = 0.0625$ cycles/byte, i.e.\
\SI{16}{\byte\per\cycle}, at \SI{1}{\giga\hertz}) places a typical
composite at roughly \SI{43}{\nano\second}, and a hard cap on a
composite's descriptor size prevents the model from rewarding
arbitrarily large fused commands beyond what real descriptor queues
accept. The dispatch cost is enabled and tuned per topology; in the
configurations measured here, command issue is not the bottleneck---data
movement is---so this term stays small relative to DMA and collective
time.
How accurate is all this? The model is precise about the things that
dominate kernel latency on this class of hardware: link bandwidth
occupancy and serialization, HBM channel parallelism, flit-level
streaming, and per-component switching overhead. The GEMM study
(\S\ref{sec:gemm}) provides a direct check: the measured MAC efficiency
tracks the analytic (theoretical) efficiency within roughly
\SIrange{10}{20}{\percent} across a wide range of tile counts, with the
gap attributable to pipeline fill and DMA effects that the analytic model
omits. The known simplifications---idealized arbitration, no thermal or
refresh effects, fixed burst granularity---are the price of a
deterministic, inspectable model; they bound the absolute accuracy but do
not distort the \emph{relative} comparisons (tiling A vs.\ B, topology X
vs.\ Y) that this report is built on.
\subsection{Modeled hardware configuration}
\label{sec:hw}
Table~\ref{tab:hw} summarizes the hardware configuration used for every
experiment in this report. It is read directly from the simulator's
topology description; per-experiment workload parameters (matrix shapes,
collective sizes, sequence lengths) are stated in their respective
sections rather than here.
\begin{table}[t]
\centering
\caption{Modeled hardware configuration (shared by all experiments).}
\label{tab:hw}
\small
\begin{tabular}{@{}ll@{}}
\toprule
\textbf{Parameter} & \textbf{Value} \\
\midrule
\multicolumn{2}{@{}l}{\emph{Hierarchy}} \\
SIPs & 2 (1D ring) \\
CUBEs per SIP & 16 ($4\times4$ mesh) \\
PEs per CUBE & 8 (4 corners $\times$ 2) \\
PEs total & 256 \\
\midrule
\multicolumn{2}{@{}l}{\emph{Processing element (PE)}} \\
GEMM engine peak & \SI{8}{\tera\flop\per\second} (f16) \\
TCM (on-PE) & \SI{16}{\mega\byte}, \SI{512}{\giga\byte\per\second} R/W \\
\quad kernel scratch & \SI{1}{\mega\byte} \\
DMA engines & 1 read + 1 write \\
CPU / scheduler overhead & \SI{2}{\nano\second} / \SI{1}{\nano\second} \\
\midrule
\multicolumn{2}{@{}l}{\emph{Memory (per CUBE)}} \\
HBM capacity & \SI{48}{\giga\byte} (8 slices) \\
HBM aggregate BW & \SI{1024}{\giga\byte\per\second} \\
HBM pseudo-channels & 64 (8 per PE), \SI{32}{\giga\byte\per\second} each \\
SRAM (shared) & \SI{32}{\mega\byte}, \SI{128}{\giga\byte\per\second} link \\
HBM burst & \SI{256}{\byte} \\
\midrule
\multicolumn{2}{@{}l}{\emph{Interconnect}} \\
Intra-CUBE NoC link & \SI{256}{\giga\byte\per\second}, \SI{2}{\nano\second}/router \\
Inter-CUBE (UCIe PHY) & \SI{512}{\giga\byte\per\second}, \SI{8}{\nano\second}, XY routing \\
Inter-SIP (PCIe) & \SI{768}{\giga\byte\per\second} per endpoint \\
\midrule
\multicolumn{2}{@{}l}{\emph{Command-issue cost model (defaults)}} \\
FIXED per command & 40 cycles \\
per-byte rate $R$ & 0.0625 cycles/byte (\SI{16}{\byte\per\cycle}) \\
composite size cap & \SI{1024}{\byte} \\
\bottomrule
\end{tabular}
\end{table}
@@ -0,0 +1,110 @@
\section{GEMM Acceleration via the Composite Command}
\label{sec:gemm}
\subsection{Why it is needed}
GEMM is the compute core of every transformer block---the QKV
projections, the attention score and context products, and the
feed-forward matrices are all matrix multiplications. On a tiled
accelerator a single logical GEMM expands into many hardware tiles, and
each tile must be read from HBM, fetched into the register file, computed,
stored, and written back. If every one of those tile-stage steps were an
independently issued command, two costs would dominate. First, the host
and the PE control processor would pay a per-command issue overhead
$O(\text{tiles}\times\text{stages})$ times, which for a deep-$K$ reduction
is hundreds to thousands of issues. Second, with stages dispatched
one-at-a-time the scheduler cannot overlap the DMA of the next tile with
the compute of the current one---the pipeline never fills, and the MAC
array sits idle waiting for data. The hardware question is therefore: what
issue mechanism lets a single GEMM saturate the MAC array without drowning
in command overhead?
\subsection{Design}
The answer is the \emph{composite command}. A single command carries the
ordered tile pipeline
\[
\textsf{DMA\_READ}\rightarrow\textsf{FETCH}\rightarrow\textsf{GEMM}
\rightarrow\textsf{STORE}\rightarrow\textsf{DMA\_WRITE},
\]
and the PE scheduler splits the payload into hardware tiles
(here $32\times64\times32$), emitting one tile token per tile. Subsequent
stages are reached by \emph{token self-routing} between the on-PE engines,
so a tile flows DMA\,$\rightarrow$\,fetch\,$\rightarrow$\,GEMM\,$%
\rightarrow$\,store without returning to the scheduler between stages.
Because the whole pipeline is described by one command, the issue cost is
paid once per GEMM rather than once per tile-stage, and the scheduler is
free to keep every stage busy on different tiles simultaneously---tile
$i$'s GEMM overlaps tile $i{+}1$'s DMA read. A multi-operation composite
additionally lets an epilogue (for example a vector-math step) ride the
same tile loop, firing per $K$-tile, per output tile, or once per kernel
according to its declared scope; this is what later allows an attention
kernel to fuse its softmax work into the GEMM pipeline
(\S\ref{sec:gqa}).
\subsection{Results}
We sweep eight GEMM shapes spanning square, tall, wide, and deep-$K$
geometries, under three operand-staging variants
(\textsf{ref\_ref}, both operands streamed from HBM; \textsf{load\_ref},
one operand resident in TCM; \textsf{load\_load}, both resident).
Figure~\ref{fig:gemm-util} reports MAC utilization and efficiency, and
Figure~\ref{fig:gemm-stages} breaks the kernel into per-stage engine
busy time.
\begin{figure}[t]
\centering
\includegraphics[width=\linewidth]{gemm_mac_utilization_theoretical_vs_measured.png}
\caption{GEMM MAC utilization and efficiency, theoretical vs.\ measured.
Tile-fill sets the ceiling: under-tile shapes (marked $\ast$) such as
$M{=}K{=}N{=}32$, $M{=}8$, and $K{=}8$ cannot fill the MAC tile and cap at
\SI{50}{\percent}, \SI{25}{\percent}, \SI{12.5}{\percent}. For
tile-filling shapes, efficiency climbs with tile count---from
\textasciitilde\SI{23}{\percent} at one tile to \textasciitilde%
\SI{78}{\percent} measured at 48 tiles ($K{=}3072$)---and the measured
bars track the analytic prediction within
\SIrange{10}{20}{\percent}.}
\label{fig:gemm-util}
\end{figure}
\begin{figure}[t]
\centering
\includegraphics[width=\linewidth]{gemm_stage_breakdown.png}
\caption{Per-stage engine wall-clock (DMA in, Fetch, GEMM, DMA out) under
\textsf{load\_ref} staging. For the deep-$K$ shape ($K{=}3072$, 48 tiles)
the Fetch and GEMM stages are large and comparable
(\textasciitilde\SI{770}{} and \SI{785}{\nano\second}) while DMA-out is
negligible---a compute-rich, well-pipelined regime. For the low-reuse
shape ($M{=}128,K{=}8,N{=}128$) DMA-out grows to
\textasciitilde\SI{350}{\nano\second} and compute is small---a
data-movement-bound regime.}
\label{fig:gemm-stages}
\end{figure}
Two regularities stand out. (i) Utilization is governed by how completely
the problem fills the MAC tile: the three under-tile shapes are hard-capped
well below \SI{100}{\percent}, independent of how the kernel is issued.
(ii) Among tile-filling shapes, efficiency is governed by tile count---more
tiles amortize the one-time pipeline fill and issue cost, so the deep-$K$
shape reaches \textasciitilde\SI{78}{\percent} of peak while a single-tile
shape reaches only \textasciitilde\SI{23}{\percent}. The stage breakdown
explains why: with 48 tiles the GEMM and Fetch stages overlap and stay
busy, whereas the low-reuse shape spends most of its time moving data in
and out.
\subsection{Analysis and meaning}
The composite command does not manufacture bandwidth or MAC throughput; it
removes the two software-shaped obstacles between a GEMM and its hardware
roofline. By paying issue cost once per GEMM and chaining tile-stages
through token self-routing, it lets the scheduler keep the pipeline full,
so compute-rich shapes actually reach the efficiency their arithmetic
intensity allows ($\sim$\SI{78}{\percent} measured at 48 tiles), and
data-bound shapes actually reach their DMA bound instead of stalling on
command overhead. The close agreement between measured and theoretical
efficiency (Figure~\ref{fig:gemm-util}) is also the report's primary
validation that KernBench's latency model is faithful in the regime that
matters. The hardware implication is concrete: a single-command,
self-routing tile pipeline is the issue mechanism that makes the MAC array
usable, and it is a prerequisite---not a luxury---for the fused attention
kernel of \S\ref{sec:gqa}.
@@ -0,0 +1,113 @@
\section{All-Reduce Acceleration via PE\_IPCQ}
\label{sec:allreduce}
\subsection{Why it is needed}
Distributing a transformer across devices turns every tensor-parallel
layer into a collective: partial results computed on different PEs, CUBEs,
and SIPs must be summed and redistributed with an all-reduce. If that
collective is handled by the host or by a generic DMA path, three problems
appear. The reduction traffic competes with the kernel's own compute DMA
on the same links, causing head-of-line blocking; there is no efficient
peer-to-peer ring primitive, so data takes extra hops; and the ordering is
hard to make deterministic. The hardware question is how to perform
collectives \emph{on the device}, overlapped with compute and reproducible
run-to-run.
\subsection{Design}
KernBench models a dedicated per-PE collective engine, \textbf{PE\_IPCQ}
(inter-PE communication queue). It is a control-plane block: it owns the
ring-buffer address arithmetic, head/tail pointers, peer-pointer caches,
backpressure, and the four-direction (N/S/E/W) neighbor map, with eight
ring buffers per PE (four directions $\times$ \{tx, rx\}). Crucially,
PE\_IPCQ does \emph{not} move data itself---it delegates the actual
transfer to PE\_DMA, keeping a clean control/data split. To stop
collective traffic from blocking compute, PE\_DMA is extended into a
two-channel virtual-channel model: \texttt{vc\_compute} carries tile
load/store for GEMM and vector math, \texttt{vc\_comm} carries IPCQ sends,
each with an independent state machine. The same physical link is shared
but progresses in chunks (\SI{256}{\byte}), so a large GEMM DMA does not
lock the link end-to-end against a pending reduction. On top of this
substrate the collective runs a hierarchical local-reduce / global
all-reduce-broadcast schedule across whatever inter-device topology the
configuration specifies.
\subsection{Results}
Following the milestone-evaluation convention, the collective sweep builds
its own six-device (six-SIP, $2\times3$) configurations---distinct from
the two-SIP default of Table~\ref{tab:hw}---and measures all-reduce
latency as a function of payload size for three inter-device topologies:
a 1D ring, a 2D mesh (no wrap), and a 2D torus. Table~\ref{tab:allreduce}
and Figure~\ref{fig:allreduce-cmp} report the result.
\begin{table}[t]
\centering
\caption{All-reduce latency (ns) across six devices, by topology and
per-PE payload. Lower is better; the torus wins at every size.}
\label{tab:allreduce}
\small
\begin{tabular}{@{}rrrr@{}}
\toprule
\textbf{Bytes/PE} & \textbf{2D mesh} & \textbf{Ring 1D} & \textbf{2D torus} \\
\midrule
256 & 2667 & 2365 & 1701 \\
4{,}096 & 4450 & 4082 & 3038 \\
16{,}384 & 8900 & 8217 & 6403 \\
65{,}536 & 26705 & 24766 & 19865 \\
98{,}304 & 38574 & 35798 & 28840 \\
\bottomrule
\end{tabular}
\end{table}
\begin{figure}[t]
\centering
\includegraphics[width=\linewidth]{allreduce_comparison.png}
\caption{All-reduce latency vs.\ per-PE payload for the three topologies,
against the analytic torus model and an external full-system simulator
(FSIM) reference. The measured torus tracks the analytic curve within a
small constant factor; the FSIM single-device point
(\SI{366}{\micro\second}) sits an order of magnitude above the
KernBench algorithmic latency, illustrating the difference between an
achievable-kernel number and a full end-to-end-stack number.}
\label{fig:allreduce-cmp}
\end{figure}
\begin{figure}[t]
\centering
\includegraphics[width=\linewidth]{allreduce_buffer_kind.png}
\caption{Effect of IPCQ staging-buffer placement (2D torus). At
\SI{64}{\kibi\byte}/PE, TCM staging (\SI{19865}{\nano\second}) beats HBM
(\SI{23081}{\nano\second}) by \textasciitilde\SI{14}{\percent} and SRAM
(\SI{32201}{\nano\second}) by \textasciitilde\SI{38}{\percent}; at small
payloads the three are indistinguishable.}
\label{fig:allreduce-buf}
\end{figure}
Three findings follow. First, topology matters and the effect grows with
size: at \SI{96}{\kilo\byte}/PE the 2D torus is \SI{25}{\percent} faster
than the mesh and \SI{19}{\percent} faster than the ring, because its
wrap-around links shorten the worst-case reduction path. Second, the
measured curves grow smoothly with payload and stay within a small
constant factor of the analytic torus model, with the gap reflecting real
link serialization that the analytic model idealizes away. Third, where
the IPCQ staging buffer lives is a first-order knob: keeping it in on-PE
TCM is materially faster than HBM or shared SRAM at large payloads
(Figure~\ref{fig:allreduce-buf}).
\subsection{Analysis and meaning}
PE\_IPCQ turns the collective from an off-device, contention-prone
operation into an on-device primitive whose latency tracks the
interconnect's physical limits. The control/data split keeps the engine
small while reusing PE\_DMA for movement, and the virtual-channel split is
what lets a reduction proceed without stalling the compute DMA that feeds
the very GEMMs producing the partials---the same property the fused
attention kernel relies on when it interleaves KV reduction with score
computation (\S\ref{sec:gqa}). For the hardware roadmap the results argue
two things: provisioning wrap-around (torus) inter-device links is worth
roughly a \SI{20}{}--\SI{25}{\percent} collective-latency reduction at
scale, and giving the IPCQ a fast on-PE staging buffer (TCM) rather than
forcing it through HBM or SRAM is worth another \SI{14}{}--%
\SI{38}{\percent} at large payloads.
@@ -0,0 +1,123 @@
\section{Fused Grouped-Query Attention}
\label{sec:gqa}
\subsection{Why it is needed}
Attention is the 1H focus, and it is where the two preceding optimizations
have to come together. Grouped-Query Attention (GQA) shrinks the KV cache
by sharing each KV head across a group of query heads (here $h_q=8$ query
heads to $h_{kv}=1$ KV head, a group factor $G=8$), which makes decoding
feasible at long context but also makes it acutely memory-bound: a decode
step processes a single query position ($T_q=1$) against the entire KV
history, so its arithmetic intensity is low and its time is dominated by
streaming the KV cache out of HBM. FlashAttention-style tiling with an
online-softmax merge avoids ever materializing the full score matrix, but
realizing it as a fast \emph{fused} kernel needs both building blocks from
this report: efficient GEMM issue (\S\ref{sec:gemm}) for the
$Q\!\cdot\!K^{\top}$ and $P\!\cdot\!V$ products, and an efficient on-device
reduction (\S\ref{sec:allreduce}) for the multi-user and
sequence-parallel KV reductions. This section is the capstone: the fused
kernel that uses the composite command and PE\_IPCQ at the same time.
Multi-head attention (MHA) was studied in prior work and serves here as
the established baseline rather than being re-derived.
\subsection{Design}
The fused GQA kernel issues its matrix products as scheduler-managed
composite commands and keeps the online-softmax merge and the cross-device
KV reduction inside the kernel, on PE\_IPCQ. Two kernel families cover the
two phases. The \emph{prefill} kernel is head-parallel and rotates the KV
shards around an inter-CUBE ring (``Ring KV''). The \emph{decode} kernel
is head-replicated with a statically sharded KV cache and reduces partial
attention outputs through an M-fold intra-CUBE chain and, for multiple
users, a two-level reduce-to-root. Two further primitives make long
context practical: a \emph{lazy load} that issues the KV \textsf{DMA\_READ}
and returns immediately, auto-waiting only at first use so that KV load
overlaps score computation; and per-tile \emph{scratch recycling} that
keeps the running softmax accumulators ($m,\ell,O$) in a persistent arena
while freeing per-tile temporaries, so the kernel fits the
\SI{1}{\mebi\byte} scratch budget across many tiles. A further refinement
that restructures the decode step into two stateful composites (a named
\textsf{softmax\_merge} recipe) is designed but not yet wired into the
measured path; results below reflect the implemented kernel only.
\subsection{Results}
We measure four headline panels that vary the user count $C$ and the phase:
single- and multi-user prefill ($T_q=4$, $S_{kv}=16$), and single- and
multi-user decode ($P=8$ PEs, $S_{kv}=64$ and $128$), all at $d_{\text{head}}=64$
and $G=8$. For each panel we harvest end-to-end latency
(max event end minus min event start, the same window convention as the
GEMM study) together with the per-engine busy time and the operation mix.
Figure~\ref{fig:gqa-lat} and Figure~\ref{fig:gqa-break} report the result;
the underlying numbers are in Table~\ref{tab:gqa}.
\begin{table}[t]
\centering
\caption{Fused GQA per-panel latency and operation mix. Compute (GEMM,
MATH) is a tiny fraction of DMA occupancy; IPCQ copies grow with users and
PEs.}
\label{tab:gqa}
\small
\begin{tabular}{@{}lrrrr@{}}
\toprule
\textbf{Panel} & \textbf{Lat.\ (ns)} & \textbf{GEMM} & \textbf{IPCQ} & \textbf{DMA rd} \\
\midrule
prefill C=1 & 445 & 2 & 0 & 3 \\
prefill C=4 (Ring) & 4630 & 32 & 24 & 12 \\
decode C=1, P=8 & 3632 & 16 & 21 & 24 \\
decode C=4, P=8 & 6693 & 64 & 93 & 96 \\
\bottomrule
\end{tabular}
\end{table}
\begin{figure}[t]
\centering
\includegraphics[width=\linewidth]{gqa_latency_by_panel.png}
\caption{Fused GQA end-to-end latency. Latency grows from
\SI{445}{\nano\second} (single-user prefill) to \SI{6693}{\nano\second}
(four-user decode) as the KV history and the number of participating
devices grow.}
\label{fig:gqa-lat}
\end{figure}
\begin{figure}[t]
\centering
\includegraphics[width=\linewidth]{gqa_op_engine_breakdown.png}
\caption{Where the work goes. Left: operation counts---GEMM and IPCQ-copy
volume both scale with users and PEs. Right: summed engine occupancy on a
log scale---the DMA engine dominates by two to three orders of magnitude
over the GEMM and MATH engines in every panel.}
\label{fig:gqa-break}
\end{figure}
The dominant observation is in Figure~\ref{fig:gqa-break}: the compute
engines are almost idle. The GEMM engine accumulates only
\SIrange{2}{33}{\nano\second} of busy time across the panels and the
vector-math engine \SIrange{5}{688}{\nano\second}, while the DMA engine
accumulates \SIrange{72}{15920}{\nano\second}. Fused GQA, as modeled here,
is overwhelmingly data-movement bound. The operation mix shows why the
collective machinery matters: IPCQ-copy count rises from zero (single-user
prefill) to 93 (four-user decode) as the kernel reduces partial outputs
across more PEs and CUBEs, and DMA-read count rises in step as more KV
shards are streamed. The PE control-processor dispatch cost registered as
zero in this configuration---command issue is simply not on the critical
path when data movement is this dominant.
\subsection{Analysis and meaning}
These panels are the clearest statement of the codesign thesis in the
report. Because the composite command keeps GEMM issue cheap and the MAC
array barely occupied, the fused attention kernel's latency is set almost
entirely by data movement: streaming the KV cache and reducing partials
across devices. That is precisely the cost that the communication-side
work targets---PE\_IPCQ for the on-device reduction, the lazy load for
load/compute overlap, fast TCM staging and torus links for the reduction
itself. In other words, the two enablers are not independent features that
happen to appear in the same kernel; the GEMM optimization is what
\emph{exposes} the data-movement bottleneck (by removing the compute and
issue overhead that would otherwise hide it), and the communication
optimization is what \emph{attacks} it. For an attention-dominated decoder
the meaningful hardware investments are therefore the ones that move data
faster and reduce it on-device---not additional MAC throughput, which this
workload cannot use.
@@ -0,0 +1,51 @@
\section{Discussion: which hardware changes are meaningful}
\label{sec:discussion}
Read together, the three studies point to a consistent ranking of where
hardware investment pays off for attention-centric decoding.
\paragraph{The composite command is foundational, but indirectly.} Its
direct effect---driving compute-rich GEMMs to
\textasciitilde\SI{78}{\percent} of peak---matters most for the
compute-bound parts of a model (the large feed-forward and projection
matrices). For attention itself, its more important effect is
\emph{diagnostic}: by making issue and compute nearly free, it removes the
overhead that would otherwise mask the true bottleneck, and the fused GQA
results then show unambiguously that the kernel is data-movement bound.
Without a cheap, self-routing issue mechanism we would not be able to tell
whether attention is slow because of compute or because of data movement;
with it, the answer is clear.
\paragraph{The communication path is where attention latency actually
lives.} Every all-reduce and fused-GQA measurement says the same thing:
the limiting resource is moving and reducing data, not multiplying it. That
makes the PE\_IPCQ design and the choices around it the highest-value
hardware levers for this workload:
\begin{itemize}
\item \textbf{On-device collectives with a compute/communication
virtual-channel split.} Performing the reduction on the device, with
\texttt{vc\_comm} separated from \texttt{vc\_compute} so the reduction
does not stall the compute DMA, is what lets attention overlap KV
movement with score computation at all.
\item \textbf{Fast on-PE staging memory.} Keeping the IPCQ buffer in TCM
rather than HBM or SRAM is worth \SI{14}{}--\SI{38}{\percent} of
collective latency at large payloads---a pure placement decision with a
first-order effect.
\item \textbf{Wrap-around (torus) inter-device links.} A torus fabric
buys \SI{20}{}--\SI{25}{\percent} over a mesh or ring at scale by
shortening the worst-case reduction path.
\end{itemize}
\paragraph{Raw MAC throughput is not the constraint for attention.} The
GQA panels leave the GEMM and vector-math engines two to three orders of
magnitude below the DMA engine in busy time. Adding MAC area would not move
decode latency; the workload cannot use it. This is the single most
actionable finding for an attention-dominated roadmap.
\paragraph{Caveats.} These conclusions are achievable-kernel results from a
deterministic model, not E2E measurements; absolute numbers carry the
model's idealizations (\S\ref{sec:latency}), though the relative rankings
that drive the recommendations are robust to them. The headline GQA panels
are also at modest scale (up to four users, single SIP), and one designed
refinement---the two-composite \textsf{softmax\_merge} decode---is not yet
in the measured path. Larger-scale and multi-SIP headline runs are 2H work.
@@ -0,0 +1,27 @@
\section{Conclusion}
\label{sec:conclusion}
This 1H work set out to make attention-centric LLM kernels fast through
hardware--software codesign, and to do so on a platform that isolates
algorithm-level behavior from the rest of the software stack. The result is
a coherent picture rather than three separate optimizations. A composite
command that issues a tiled GEMM as one self-routing pipeline makes the MAC
array usable---reaching \textasciitilde\SI{78}{\percent} of peak on
compute-rich shapes with measured efficiency tracking theory---and, just as
importantly, makes compute cheap enough that the real bottleneck becomes
visible. A per-PE on-device collective engine, PE\_IPCQ, turns all-reduce
into a primitive whose latency follows the interconnect's physical limits,
with topology and staging-memory choices each worth tens of percent. Fused
Grouped-Query Attention then combines the two and shows the payoff and the
lesson at once: the kernel is data-movement bound, so the optimizations
that move and reduce data---not those that add arithmetic---are what
determine its speed.
The practical conclusion for the hardware roadmap is therefore specific.
The changes worth keeping are the single-command self-routing GEMM
pipeline, the on-device PE\_IPCQ collective with its compute/communication
virtual-channel split, fast on-PE staging memory, and wrap-around
inter-device links. Additional MAC throughput is not, for this workload, a
meaningful investment. KernBench made these conclusions measurable by
holding everything except the algorithm and the hardware fixed; the next
half extends the same method beyond attention to the rest of the decoder.
@@ -0,0 +1,40 @@
\section{Future Work --- 2H}
\label{sec:future}
The 1H work covered attention end to end. The natural next step is to
complete the decoder and then to ask how compute and data should be
distributed for realistic, agentic workloads.
\paragraph{From attention to the full decoder: FFN and MoE.} A decoder
block is attention followed by a feed-forward network (FFN), and in modern
models that FFN is increasingly a mixture-of-experts (MoE) layer. The FFN
is the compute-rich counterpart to attention---it is where the composite
command's MAC-efficiency gains (\S\ref{sec:gemm}) should matter most---so
adding it gives a balanced view of a full block instead of its
memory-bound half alone. MoE adds a new dimension: a routing step selects a
few experts per token, turning the dense FFN GEMM into a sparse, data-
dependent dispatch. The open questions are how to issue expert GEMMs as
composites under data-dependent token counts, and how to move tokens to
experts efficiently---an all-to-all-shaped communication pattern distinct
from the all-reduce studied here, and a natural extension of the PE\_IPCQ
work.
\paragraph{Compute and data distribution for full LLM decoding.} With both
attention and FFN/MoE in hand, the question becomes where each layer's
compute and state should live. Attention is KV-bound and favors keeping the
KV cache close to the PEs that consume it; FFN/MoE is compute-bound and
favors spreading GEMM work across PEs; MoE routing makes the optimal
placement token- and time-dependent. A 2H goal is to use KernBench to
explore these placement and parallelization trade-offs---tensor vs.\ expert
vs.\ sequence parallelism---under a single, software-stack-independent
model, so the interconnect and memory implications of each choice are
measured rather than assumed.
\paragraph{Agentic workloads.} Agentic inference interleaves many
short, bursty decode requests with tool use and long shared contexts,
which stresses the system differently from a single long generation:
context reuse across requests, dynamic batching, and uneven expert load all
change how compute and data should be dispersed. Characterizing how total
compute and data movement distribute across the SIP/CUBE/PE hierarchy under
such workloads---and which of the 1H hardware levers still dominate when the
workload is this irregular---is the broader 2H agenda.
+73
View File
@@ -0,0 +1,73 @@
# KernBench — 2026 1H HW-SW Codesign Report — Table of Contents
> Agreed outline (the `/paper` contract). `/paper build` follows this.
> Period: 2026 1H. Language: English. References: external literature only
> (no ADR/SPEC named anywhere in the paper). PDF via Tectonic.
## Section order → files
1. **Executive Summary**`sections/00-exec-summary.tex`
The attention-optimization goal, the two enabling optimizations (GEMM
composite command + PE_IPCQ communication), the fused GQA capstone, the
headline results, and the bottom-line recommendation on which HW changes
are worth keeping. Front-loads conclusions; may run slightly longer than
a terse abstract.
2. **Introduction**`sections/01-introduction.tex`
1H focus is **attention-kernel optimization**: FlashAttention-style
tiling and Grouped Query Attention (GQA), building on prior work on
Multi-Head Attention (MHA, studied earlier — referenced as the
established baseline, not re-derived). Optimizing a fused attention
kernel requires two enabling optimizations, each studied in its own
right and then **combined inside the fused kernel**:
(a) **GEMM optimization** (the composite command), and
(b) **communication optimization** (PE_IPCQ).
Narrative arc: *two enablers → the fused GQA capstone that uses both.*
Motivate why HW-SW codesign (not software alone) is required.
3. **The KernBench Platform**`sections/02-platform.tex`
- 3.1 *Why KernBench* — source-level kernel execution, no compiler /
SW-stack dependency; isolate algorithm-level optimization.
- 3.2 *Execution model* — discrete-event graph; runtime API → sim_engine
→ components; PE pipeline + composite commands; 2-pass data/timing.
- 3.3 *Latency model & accuracy* — traversal-based golden invariants;
structural CPU-dispatch cost model; known simplifications + calibration.
- 3.4 *Modeled Hardware Configuration* — shared platform-wide config from
`topology.yaml` (hierarchy counts, clocks, capacities/BW, link params,
cost-model constants). Per-experiment workload params live in §4–§6.
4. **GEMM Acceleration via Composite Command**`sections/03-gemm.tex`
necessity · design · results · analysis. (+ GEMM workload params)
5. **All-Reduce Acceleration via PE_IPCQ**`sections/04-allreduce.tex`
necessity · design · results · analysis. (+ All-Reduce topologies/sizes)
6. **Fused Grouped Query Attention** (composite + PE_IPCQ) — `sections/05-gqa.tex`
necessity · design · results · analysis. (+ GQA seq/head/user configs)
7. **Discussion**`sections/06-discussion.tex`
Which HW changes are meaningful, and under what regimes (cross-cutting).
8. **Conclusion**`sections/07-conclusion.tex`
The codesign thesis, stated plainly, supported by the measured results.
9. **Future Work — 2H**`sections/08-future-work.tex`
Add the FFN/MoE layer toward full LLM decoding; how compute & data should
be distributed for agentic / MoE workloads.
10. **References** *(optional)* — external literature only (FlashAttention,
Megatron-LM, GPT-3, Llama 3). No ADR/SPEC entries.
## Per-section structure (§4/§5/§6)
Each of the three optimization sections follows the same four beats:
**(a) necessity → (b) design → (c) experimental results → (d) analysis & meaning.**
## Build notes
- GEMM (§4) and All-Reduce (§5): reuse committed figures/CSVs under
`src/kernbench/benches/1H_milestone_output/{gemm,ccl}/`.
- GQA (§6): committed `sweep.json` holds op-counts only — latency + figures
are generated by an isolated harness under `scripts/paper/` at build time.
Measure only the implemented path; mark proposed designs (e.g. the
flat-ops `softmax_merge` decode opt2) as designed-not-measured.
+101
View File
@@ -0,0 +1,101 @@
"""Report harness: GQA end-to-end latency + op/engine breakdown.
Isolated under ``scripts/paper/`` for the 1H codesign report only — it is
NOT a registered bench and does not touch other people's benches. It
reuses the existing GQA headline panels (the real GQA kernels wired in
``milestone_gqa_headline``) but, unlike that milestone (which records only
op-counts), it also harvests per-panel end-to-end latency and per-engine
occupancy from ``result.engine.op_log``.
Latency definition (same window convention as ``milestone_1h_gemm``):
end_to_end_ns = max(r.t_end) - min(r.t_start) over all op_log records.
Output: docs/report/1H-codesign-paper/figures/gqa_latency.json
Run:
python scripts/paper/paper_gqa_latency.py
"""
from __future__ import annotations
import json
from pathlib import Path
from kernbench.benches.milestone_gqa_headline import (
_PANEL_DISPATCH,
_PANELS,
_make_bench_fn,
_summarize_op_log,
)
_REPORT_DIR = Path(__file__).resolve().parents[2] / "docs" / "report" / "1H-codesign-paper"
_FIG_DIR = _REPORT_DIR / "figures"
_OUT_JSON = _FIG_DIR / "gqa_latency.json"
# PE engine component suffixes whose occupancy we break out.
_ENGINES = ("pe_gemm", "pe_math", "pe_dma", "pe_fetch_store", "pe_ipcq", "pe_cpu")
def _occupancy_ns(op_log, suffix: str) -> float:
return sum(
r.t_end - r.t_start
for r in op_log
if r.component_id.endswith("." + suffix)
)
def _end_to_end_ns(op_log) -> float:
if not op_log:
return 0.0
return max(r.t_end for r in op_log) - min(r.t_start for r in op_log)
def _run_panel(panel: str, topology: str) -> dict:
from kernbench.runtime_api.bench_runner import run_bench
from kernbench.runtime_api.types import resolve_device
from kernbench.sim_engine.engine import GraphEngine
from kernbench.topology.builder import resolve_topology
topo = resolve_topology(topology)
result = run_bench(
topology=topo,
bench_fn=_make_bench_fn(panel),
device=resolve_device(None),
engine_factory=lambda t, d: GraphEngine(
getattr(t, "topology_obj", t), enable_data=True,
),
)
if not result.completion.ok:
raise RuntimeError(f"panel {panel!r} failed: {result.completion}")
op_log = result.engine.op_log
kind, params = _PANEL_DISPATCH[panel]
return {
"panel": panel,
"kind": kind,
**params,
"latency_ns": _end_to_end_ns(op_log),
"op_log_summary": _summarize_op_log(op_log),
"engine_occupancy_ns": {
eng: _occupancy_ns(op_log, eng) for eng in _ENGINES
},
}
def main() -> None:
topology = "topology.yaml"
rows = [_run_panel(panel, topology) for panel in _PANELS]
_FIG_DIR.mkdir(parents=True, exist_ok=True)
out = {"version": 1, "panels": list(_PANELS), "rows": rows}
_OUT_JSON.write_text(json.dumps(out, indent=2))
print(f"wrote {_OUT_JSON}")
for r in rows:
s = r["op_log_summary"]
print(
f" {r['panel']:24s} latency={r['latency_ns']:10.1f} ns "
f"gemm={s['gemm_count']:3d} ipcq={s['ipcq_copy_count']:3d} "
f"dma_rd={s['dma_read_count']:3d} dma_wr={s['dma_write_count']:2d}"
)
if __name__ == "__main__":
main()
+108
View File
@@ -0,0 +1,108 @@
"""Report harness: render GQA figures from ``gqa_latency.json``.
Isolated under ``scripts/paper/`` for the 1H codesign report only.
Reads the JSON emitted by ``paper_gqa_latency.py`` and writes two PNGs into
``docs/report/1H-codesign-paper/figures/``:
gqa_latency_by_panel.png end-to-end latency per panel
gqa_op_engine_breakdown.png op-counts + engine occupancy per panel
Run (after paper_gqa_latency.py):
python scripts/paper/paper_plot_gqa.py
"""
from __future__ import annotations
import json
from pathlib import Path
import matplotlib
matplotlib.use("Agg")
import matplotlib.pyplot as plt # noqa: E402
_FIG_DIR = Path(__file__).resolve().parents[2] / "docs" / "report" / "1H-codesign-paper" / "figures"
_IN_JSON = _FIG_DIR / "gqa_latency.json"
_LABELS = {
"single_user_prefill_gqa": "prefill\nC=1",
"multi_user_prefill_gqa": "prefill\nC=4 (Ring KV)",
"single_user_decode_gqa": "decode\nC=1, P=8",
"multi_user_decode_gqa": "decode\nC=4, P=8",
}
def _load() -> list[dict]:
return json.loads(_IN_JSON.read_text())["rows"]
def _plot_latency(rows: list[dict]) -> Path:
labels = [_LABELS[r["panel"]] for r in rows]
lat = [r["latency_ns"] for r in rows]
fig, ax = plt.subplots(figsize=(7.0, 4.0))
bars = ax.bar(labels, lat, color="#3b6ea5", width=0.6)
ax.set_ylabel("end-to-end latency (ns)")
ax.set_title("Fused GQA — end-to-end latency per panel")
ax.bar_label(bars, fmt="%.0f", padding=3, fontsize=9)
ax.grid(axis="y", ls=":", alpha=0.5)
ax.set_ylim(0, max(lat) * 1.15)
fig.tight_layout()
out = _FIG_DIR / "gqa_latency_by_panel.png"
fig.savefig(out, dpi=150)
plt.close(fig)
return out
def _plot_breakdown(rows: list[dict]) -> Path:
labels = [_LABELS[r["panel"]] for r in rows]
x = range(len(rows))
fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(11.0, 4.2))
# Left: op-count breakdown (grouped bars).
keys = ["gemm_count", "ipcq_copy_count", "dma_read_count", "dma_write_count"]
disp = ["GEMM", "IPCQ copy", "DMA read", "DMA write"]
colors = ["#3b6ea5", "#c0504d", "#9bbb59", "#8064a2"]
w = 0.2
for i, (k, d, c) in enumerate(zip(keys, disp, colors)):
vals = [r["op_log_summary"][k] for r in rows]
ax1.bar([xi + (i - 1.5) * w for xi in x], vals, width=w, label=d, color=c)
ax1.set_xticks(list(x))
ax1.set_xticklabels(labels, fontsize=8)
ax1.set_ylabel("op count")
ax1.set_title("Per-panel op-count breakdown")
ax1.legend(fontsize=8)
ax1.grid(axis="y", ls=":", alpha=0.5)
# Right: engine occupancy (log scale — compute vs data movement).
engs = ["pe_gemm", "pe_math", "pe_dma"]
edisp = ["GEMM engine", "MATH engine", "DMA engine"]
ecolors = ["#3b6ea5", "#e8a33d", "#c0504d"]
w2 = 0.25
for i, (e, d, c) in enumerate(zip(engs, edisp, ecolors)):
vals = [max(r["engine_occupancy_ns"][e], 0.1) for r in rows]
ax2.bar([xi + (i - 1) * w2 for xi in x], vals, width=w2, label=d, color=c)
ax2.set_yscale("log")
ax2.set_xticks(list(x))
ax2.set_xticklabels(labels, fontsize=8)
ax2.set_ylabel("summed engine occupancy (ns, log)")
ax2.set_title("Compute vs data-movement occupancy")
ax2.legend(fontsize=8)
ax2.grid(axis="y", ls=":", alpha=0.5)
fig.suptitle("Fused GQA — where the work goes", fontsize=12)
fig.tight_layout()
out = _FIG_DIR / "gqa_op_engine_breakdown.png"
fig.savefig(out, dpi=150)
plt.close(fig)
return out
def main() -> None:
rows = _load()
p1 = _plot_latency(rows)
p2 = _plot_breakdown(rows)
print(f"wrote {p1}")
print(f"wrote {p2}")
if __name__ == "__main__":
main()