Files
kernbench2/.claude/commands/paper.md
ywkang dd525bfcb7 paper: add /paper skill + 1H HW-SW codesign report (GEMM, All-Reduce, fused GQA)
New `/paper` slash-command skill that synthesizes ADR/SPEC content and live
KernBench benchmark results into a sectioned LaTeX technical paper compiled
to PDF with Tectonic (auto-installed). The skill negotiates a TOC, grounds
every number in committed artifacts or fresh bench runs, and keeps
report-only benches isolated.

This commit also includes the first generated report:
- docs/report/1H-codesign-paper/ — main.tex + per-section .tex, figures,
  toc.md contract, and the built 8-page main.pdf. Covers the platform
  (source-level kernels, latency model + accuracy, HW config from
  topology.yaml), GEMM via composite command, All-Reduce via PE_IPCQ, and
  fused GQA combining both, plus discussion/conclusion/2H future work.
- scripts/paper/ — isolated report harnesses (not registered benches):
  paper_gqa_latency.py harvests per-panel GQA end-to-end latency + engine
  occupancy (the milestone only emitted op-counts); paper_plot_gqa.py
  renders the GQA figures.

GEMM/All-Reduce reuse committed milestone figures/CSVs; GQA results are
generated fresh. Honest flags retained: PE_CPU dispatch cost is 0 in this
config, and the proposed two-composite softmax_merge decode is marked
designed-not-measured.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-10 22:15:14 -07:00

20 KiB
Raw Permalink Blame History

description
description
Generate the 1H HW-SW codesign technical paper (PDF) — propose/confirm a TOC, then synthesize ADRs + live KernBench benchmark results into a sectioned LaTeX paper compiled with Tectonic.

/paper — HW-SW Codesign Technical Report Generator

Produces a technical-paper-style PDF at docs/report/1H-codesign-paper/build/main.pdf, synthesizing the ADR corpus, SPEC.md, and live KernBench benchmark results into a publication-shaped report on the three 1H optimizations:

  1. GEMM Acceleration using the composite command
  2. All-Reduce Acceleration using PE_IPCQ
  3. Fused Grouped Query Attention using both composite command and PE_IPCQ

plus an explanation of the simulation platform itself (why KernBench exists, its execution model, and how it computes latency / how accurate that is).

Audience is the internal team (engineers + architects), but the report reads like a real technical paper: do not cite ADR-NNNN or SPEC §X.Y anywhere — not in prose, not in a reference list. ADRs and SPEC are grounding sources only (use their content; never name them). The References section, if used, lists external literature only (e.g. FlashAttention, Megatron-LM, GPT-3, Llama 3) — actual papers, not internal docs.


Operating principles (read first)

  • Confirm the TOC before writing. The first thing /paper does is present the outline (saved toc.md if present, else the default below) and stop for the user's confirmation / adjustments. Do not generate sections until the user approves the TOC.
  • Ground every number. Every quantitative result in prose, tables, or captions MUST trace to (a) a committed bench artifact (*.csv / *.json / *.png under 1H_milestone_output/ or docs/diagrams/), (b) a freshly-run bench output, or (c) a specific ADR/SPEC quotation. Never invent latencies, speedups, or utilizations. If a needed number does not exist, generate it via a bench run (below) or state the gap — do not guess.
  • Reuse existing results by default. The three milestone benches already have committed outputs. Do not re-run a bench unless the user asks to regenerate, or the artifact is missing/stale.
  • Design / proposal status honesty. Some designs are Proposed, not yet implemented (e.g. ADR-0061, ADR-0065). Present these as proposed and clearly separate "modeled & measured" from "designed, not yet measured." Do not report measurements for unimplemented paths.
  • English prose. Architectural terms (composite command, PE_IPCQ, GQA, TCM, SIP, CUBE, PE) used as-is.

Invocation

  • /paperdefault. Resolve the period, present the TOC (toc.md if it exists, else the default TOC), and stop for confirmation. Once the user approves (possibly after editing the TOC in conversation), persist toc.md and proceed to full build.
  • /paper toc — only (re)negotiate and persist the TOC. No sections, no build.
  • /paper build — assume the saved toc.md is approved; run the full pipeline (sections → figures → compile PDF) without re-confirming.
  • /paper section <id> — regenerate a single section (<id> = a filename stem under sections/, e.g. 04-allreduce) and recompile.
  • /paper figures — (re)generate only the benchmark figures/data (runs/refreshes benches), no prose changes, then recompile.
  • /paper pdf — only recompile the existing .tex to PDF (Tectonic).

If toc.md does not yet exist, any mode except /paper//paper toc first falls back to presenting the default TOC for confirmation.


Output layout

docs/report/1H-codesign-paper/
  toc.md                      # the agreed outline (persisted; human-editable)
  main.tex                    # preamble + \input of each section
  refs.bib                    # ADR/SPEC bibliography (optional; or inline refs)
  sections/
    00-exec-summary.tex
    01-introduction.tex
    02-platform.tex           # includes 3.4 Modeled Hardware Configuration
    03-gemm.tex
    04-allreduce.tex
    05-gqa.tex
    06-discussion.tex
    07-conclusion.tex
    08-future-work.tex
  figures/                    # PNGs copied in from bench outputs (portable)
  build/                      # Tectonic output; main.pdf lives here

Report-specific benchmark/plot harnesses (when a figure must be generated, see Generating missing results) live isolated under:

scripts/paper/                # report-only harnesses — NOT registered benches

This directory is intentionally separate from src/kernbench/benches/ so the report's harnesses never collide with other people's benches (the bench registry only audits src/kernbench/benches/). Treat scripts/paper/* and everything under docs/report/1H-codesign-paper/ as derived artifacts (like docs/diagrams/): creating/updating them does not require Phase 2 approval, but they MUST stay consistent with SPEC.md and ADRs.


Default TOC

If toc.md is absent, propose this and ask the user to confirm or edit:

  1. Executive Summary — the attention-optimization goal, the two enabling optimizations (GEMM composite command + PE_IPCQ communication), the fused GQA capstone, the headline results, and the bottom-line recommendation on which HW changes are worth keeping. May run a paragraph or two longer than a terse abstract; front-loads conclusions.
  2. Introduction — the 1H focus is attention-kernel optimization: FlashAttention-style tiling and Grouped Query Attention (GQA), building on prior work on Multi-Head Attention (MHA, studied earlier — reference it as the established baseline, not re-derived here). Optimizing a fused attention kernel requires two enabling optimizations, each studied in its own right and then combined inside the fused kernel: (a) GEMM optimization (the composite command) and (b) communication optimization (PE_IPCQ). Frame the three result sections as two enablers → the fused GQA capstone that uses both. Motivate why HW-SW codesign (not software alone) is required to realize the gains.
  3. The KernBench Platform
    1. Why KernBench — source-level kernel execution with no compiler / software-stack dependency; isolate algorithm-level optimization from other layers so a kernel can fully exploit the hardware without E2E cross-layer entanglement.
    2. Execution model — discrete-event graph; runtime API → sim_engine → components; PE pipeline + composite commands; 2-pass data/timing model.
    3. Latency model & accuracy — traversal-based golden invariants (every hop pays > 0, deterministic, explicit connectivity); the structural CPU-dispatch cost model; per-edge BW occupancy, HBM pseudo-channel parallelism, flit streaming; known simplifications & what calibration anchors the constants.
    4. Modeled Hardware Configuration — the shared, platform-wide quantitative config used across all experiments, read from topology.yaml: SIP/CUBE/PE hierarchy (counts & nesting), PE clock & compute throughput, TCM/SRAM/HBM capacities & bandwidths, NOC / inter-CUBE / inter-SIP link bandwidths & latencies, and the cost-model constants. Per-experiment workload parameters (GEMM shapes, All-Reduce topologies, GQA seq/head/user counts) stay in each evaluation section, not here.
  4. GEMM Acceleration via Composite Command — necessity · design · results · analysis.
  5. All-Reduce Acceleration via PE_IPCQ — necessity · design · results · analysis.
  6. Fused Grouped Query Attention (composite + PE_IPCQ) — necessity · design · results · analysis.
  7. Discussion — which HW changes are meaningful, and why (cross-cutting).
  8. Conclusion.
  9. Future Work — 2H — add the FFN/MoE layer toward full LLM decoding; analyze how compute & data should be distributed for agentic / MoE workloads.
  10. References (optional)external literature only (e.g. FlashAttention, Megatron-LM, GPT-3, Llama 3). No ADR/SPEC entries.

Each of §4/§5/§6 follows the same four-beat structure: (a) why the optimization is needed, (b) the design (how it works), (c) experimental results (run/plot/measure), (d) analysis & meaning.


Source map (ADR / SPEC → section)

Use these as grounding sources only — read them for content, but never name an ADR or SPEC in the paper. Verify each quantitative claim against the actual file before asserting it (CLAUDE.md: no asserting architectural numbers from memory).

  • Modeled Hardware Configuration (§3.4): topology.yaml (authoritative for counts, clocks, BW, link params, cost-model constants) is the primary source — read it directly and build the HW table from it. Cross-check against ADR-0003 (hierarchy/scope), ADR-0004 (memory BW guarantee), ADR-0033 (latency model), ADR-0064 (cost-model defaults). If topology.yaml and an ADR disagree, the loaded topology.yaml is what the sim actually runs — report that value and flag the mismatch.

  • Platform — why/execution/latency: SPEC §0, SPEC §0.1 (golden invariants), SPEC §1 R1/R2/R3, SPEC §2.1 (graph execution); ADR-0003 (hierarchy/scope), ADR-0007 (runtime↔sim_engine boundary), ADR-0014 (PE pipeline / composite), ADR-0020 (2-pass data/timing), ADR-0033 (latency model assumptions & simplifications), ADR-0064 (structural CPU-dispatch cost model: FIXED + logical_bytes×R).

  • GEMM / composite: ADR-0014 (composite command D3.2/D3.3), ADR-0042 (tile-plan generators), ADR-0044 (GEMM eval harness), ADR-0064 (cost model — composite's per-command BW reduction), ADR-0065 (flat-ops CompositeCmd, Proposed).

  • All-Reduce / PE_IPCQ: ADR-0023 (PE_IPCQ — control plane, ring buffers, VC split), ADR-0025 (direction addressing), ADR-0032 (intercube all-reduce), ADR-0043 (allreduce eval harness), ADR-0047 (AHBM CCL backend), ADR-0050 (CCL algorithm contract).

  • GQA / fused attention: ADR-0060 (fused GQA kernel), ADR-0061 (tl.broadcast/tl.repeat, Proposed), ADR-0062 (lazy tl.load), ADR-0063 (tl.scratch_scope), ADR-0064 (cost model), ADR-0065 (flat-ops composite + softmax_merge, Proposed). KO mirrors in docs/adr-ko/; DDD-0060/0065 design docs if present.

  • Eval/milestone pattern: ADR-0054 (self-contained milestone benches); ADR-0045 (bench contract).


Result map (bench → artifact → section) + run commands

All runs use the repo's topology.yaml unless a section needs a specific topology. Reuse committed artifacts unless told to regenerate.

§4 GEMM — milestone-1h-gemm

  • Artifacts (committed): src/kernbench/benches/1H_milestone_output/gemm/gemm_sweep.json, gemm_stage_breakdown.png, gemm_mac_utilization_measured.png, gemm_mac_utilization_theoretical_vs_measured.png.
  • Regenerate: kernbench run --topology topology.yaml --bench milestone-1h-gemm (full 24-shape sweep ≈ minutes; MILESTONE_FAST=1 reuses committed JSON).

§5 All-Reduce — milestone-1h-ccl

  • Artifacts (committed): src/kernbench/benches/1H_milestone_output/ccl/summary.csv, topology.png, per-topology AllReduce_LRAB_* PNGs, buffer-kind sweep AllReduce_LRAB_2Dtorus_..._with_TCM_SRAM_HBM.{png,csv}, comparison_mesh_vs_ring_vs_2DTorus_vs_theoretical_vs_fsim.png.
  • Regenerate: kernbench run --topology topology.yaml --bench milestone-1h-ccl.

§6 GQA — milestone-gqa-headline

  • Artifact (committed): src/kernbench/benches/1H_milestone_output/gqa_headline/sweep.json. Caveat: this currently holds only op_log_summary op-counts (gemm_count, ipcq_copy_count, dma_read_count, dma_write_count) for 4 panels (single/multi-user × prefill/decode) — no latency, no figures.
  • Regenerate counts: GQA_HEADLINE_RUN=1 kernbench run --topology topology.yaml --bench milestone-gqa-headline.
  • For a real GQA results section, generate latency + figures via an isolated report harness — see next.

Generating missing results (isolation rule)

When a section needs an artifact that does not yet exist (today: GQA latency and any GQA figure), create a report-dedicated harness under scripts/paper/ — never under src/kernbench/benches/. Rules:

  • One file per concern, prefixed paper_, e.g. scripts/paper/paper_gqa_latency.py (drives the engine, emits latency + a baseline-vs-optimized comparison), scripts/paper/paper_plot_gqa.py (renders figures from that data).
  • Follow the ADR-0054 self-contained pattern: the harness builds its own GraphEngine(s), imports the existing GQA kernel modules (_gqa_attention_{prefill,decode}_{short,long}.py, milestone_gqa_headline.py), runs them, and reads latency from the engine's completion timestamps / op log. It does not register a bench and does not import or mutate other people's benches.
  • Output figures land in docs/report/1H-codesign-paper/figures/ and, if useful as a shared diagram, mirror into docs/diagrams/gqa_plots/.
  • Only measure what the current code actually runs. ADR-0065's GQA opt2 (softmax_merge, two-composite decode) is Proposed; if it is not wired in production, present it as designed-not-measured and base the measured results on the implemented path. Verify implementation status by reading the code before claiming a measurement.
  • These harnesses are derived-artifact tooling (no Phase 2 needed). But if the report ever requires changing a registered production bench under src/kernbench/benches/ or any production module, that is a non-trivial production change — STOP and follow CLAUDE.md Phase 1 → approval → Phase 2.

Procedure

Step 0 — Toolchain (Tectonic)

Ensure a working Tectonic before any compile step. In order:

  1. If tectonic is on PATH (tectonic --version succeeds), use it.
  2. Else try, in order, whichever succeeds (capture output, don't hang):
    • winget install --id TectonicProject.Tectonic --silent --accept-package-agreements --accept-source-agreements
    • scoop install tectonic
    • Direct download (no admin): fetch the latest Windows x86_64-pc-windows-msvc release zip from https://github.com/tectonic-typesetting/tectonic/releases via PowerShell Invoke-WebRequest, expand it, and place tectonic.exe under docs/report/1H-codesign-paper/build/.tools/. Use that absolute path for compilation.
  3. Re-verify with tectonic --version. If all methods fail, write the .tex files anyway, report the failure, and tell the user the exact command to compile manually (or to use Overleaf). Do not block the rest of the pipeline.

Tectonic fetches LaTeX packages on first compile (needs network once).

Step 1 — Period + TOC

  • Period from system date: month 16 → {YYYY}-1H, else {YYYY}-2H. (Report dir is fixed at 1H-codesign-paper for this half; adjust the title page text with the period.)
  • If toc.md exists, present it; else present the Default TOC.
  • Stop and ask the user to confirm or edit. Incorporate edits. Persist the agreed outline to toc.md. (In /paper build mode, skip the stop and trust the saved toc.md.)

Step 2 — Platform section (§3, incl. §3.4 HW config)

Write sections/02-platform.tex from the platform source map, covering all four sub-parts:

  • §3.13.3 (why / execution model / latency model + accuracy). Include the cost-model formula dispatch_cycles = FIXED + logical_bytes×R with the committed defaults (verify against ADR-0064 before quoting), and explicitly address how accurate the latency is: what is modeled precisely (per-edge BW occupancy, HBM pseudo-channels, flit streaming, per-component overhead) and the known simplifications.
  • §3.4 Modeled Hardware Configuration: read topology.yaml directly and build a compact, table-driven summary — SIP/CUBE/PE hierarchy (counts & nesting), PE clock & compute throughput, TCM/SRAM/HBM capacity & bandwidth, NOC / inter-CUBE / inter-SIP link bandwidth & latency, and the cost-model constants (FIXED_PER_CMD, R, composite cap). Every number from topology.yaml. This is the shared config; later sections say "the modeled hardware" and only add their own workload params.

Do not name the source ADRs/SPEC in the prose.

Step 3 — The three result sections (§4 GEMM, §5 All-Reduce, §6 GQA)

For each, in order, produce the four beats:

  • (a) Necessity — from the ADR Context: the bottleneck the optimization targets (e.g. per-command dispatch overhead for GEMM; collective comm cost / HoL blocking for All-Reduce; KV-load-bound decode + softmax merge for GQA).
  • (b) Design — from the ADR Decisions: composite command tile pipeline (ADR-0014); PE_IPCQ control/data split + ring buffers + VC (ADR-0023); fused attention with online-softmax, lazy load, scratch recycling (ADR-0060/0062/0063). Note Proposed-vs-Accepted status.
  • (c) Results — reuse committed artifacts (GEMM, All-Reduce). For GQA, generate latency + figures via the isolated scripts/paper/ harness (above). Copy chosen PNGs into figures/, \includegraphics them, and build result tables from the CSV/JSON — quoting actual numbers, with a one-line note of the run command and source artifact.
  • (d) Analysis — interpret: what the curve/table shows, why the HW feature produces it, where it saturates or fails to help, and the cost (area/complexity) implied. Tie back to the cost model.

Step 4 — Discussion, Conclusion, Future Work (§6/§7/§8)

  • Discussion: synthesize across the three — which HW changes (composite command, PE_IPCQ, the cost-model-exposed BW reduction) carry their weight, and under what regimes.
  • Conclusion: the codesign thesis, supported by the measured results; state plainly which HW changes are meaningful.
  • Future Work (2H): FFN / MoE layer toward full LLM decoding; how compute & data should be distributed for agentic / MoE workloads (expert routing, token dispersion, all-to-all vs all-reduce, capacity). Frame as the next codesign questions, not claims.

Step 5 — Assemble & compile

  • main.tex: article/IEEEtran-style preamble, title with the period, author, executive-summary input, then \input{sections/...} in TOC order, then references. Keep the preamble minimal (graphicx, booktabs, hyperref, amsmath) so Tectonic compiles cleanly.
  • Compile: tectonic main.tex --outdir build (or the resolved Tectonic path) from docs/report/1H-codesign-paper/. Re-run if references need a second pass. Confirm build/main.pdf exists and report its size.

Step 6 — Chat report

Emit to chat (not to any file):

## /paper — Build Summary
**PDF:** docs/report/1H-codesign-paper/build/main.pdf  (<N> pages, <size>)
**Toolchain:** tectonic <version | how installed>
**Sections written:** <list>
**Benches: reused vs regenerated**
  - GEMM: <reused committed | regenerated via milestone-1h-gemm>
  - All-Reduce: <reused | regenerated>
  - GQA: <generated via scripts/paper/paper_gqa_latency.py | counts-only>
**Figures included:** <list of figures/*.png>
**Grounding check:** every result number traced to <artifact list>
**Proposed-not-measured items flagged:** <e.g. ADR-0065 opt2, ADR-0061>
**Open gaps / suggested next runs:** <…>

Constraints (do not violate)

  1. TOC first. Never write sections before the TOC is confirmed (default mode). toc.md is the contract.
  2. No fabricated numbers. Every quantitative claim traces to a committed artifact, a fresh bench run, or an ADR/SPEC quote. Verify architectural constants against the source file, not memory.
  3. Reuse by default. Do not re-run a bench unless asked or the artifact is missing/stale.
  4. Isolation. Report harnesses live only in scripts/paper/; never register a bench, never touch other people's benches.
  5. Production-change gate. Any change to a registered bench or production module triggers CLAUDE.md Phase 1 → approval → Phase 2. Report tooling and docs/report/... are derived artifacts (exempt).
  6. Status honesty. Mark Proposed designs as proposed; never present an unimplemented path as measured.
  7. No internal references. Never name ADR-NNNN or SPEC §X.Y in prose, captions, or the reference list — they are grounding sources only. A References section, if present, holds external literature only. Hardware numbers come from topology.yaml.
  8. Determinism. Sections in TOC order; same inputs → same paper.
  9. English prose, internal-audience tone.