Files
kernbench2/docs/adr-proposed/ADR-0064-perf-cpu-issue-cost-model.md
T
ywkang 7a9d4ec47b gqa(adr): pivot ADR-0060 to composite hybrid + lazy tl.load; add ADR-0064 cost model
- ADR-0060: GEMMs (Q.Kt, P.V) via existing tl.composite (scheduler-managed
  tiling + K/V DMA streaming); softmax merge + IPCQ tree reduction stay in
  kernel. Front TL;DR pseudocode of the final composite kernel; new section
  B lists open design items (DDD sync, K pre-transpose, dma_read lever,
  kernel-vs-scheduler tiling, ring path).
- ADR-0062: redefined from a new load_async op to global lazy tl.load
  (non-blocking + auto-wait on first use; API unchanged; goldens regenerate).
- ADR-0064 (new): per-op-type CPU issue cost model (composite ~40ns >>
  primitive) so the hybrid's CPU-saturation win becomes measurable
  (currently dispatch_cycles=0 hides it). Cost-model impl deferred.
- KO mirrors for ADR-0060/0062/0064 (-ko suffix, adr-proposed).

Rationale: non-blocking CompositeCmd offloads tiling to PE_SCHEDULER,
decoupling CPU issue-rate from execution so the CPU can saturate the
engines; the prior 'composite = no latency benefit' claim was an artifact
of dispatch_cycles=0. Docs only; no production code changed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-04 10:19:37 -07:00

180 lines
8.5 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# ADR-0064: Per-op-type CPU issue cost model (command construct + dispatch)
## Status
Proposed
> Supporting ADR for **ADR-0060** (AHBM GQA Fused Attention). The hybrid
> decision there (GEMMs via `tl.composite`, softmax merge in the kernel)
> wins by **offloading tiling to PE_SCHEDULER so the CPU issues coarse
> descriptors and runs ahead, keeping the engines saturated**. That win is
> currently **invisible in the simulator** because per-op CPU issue cost is
> zero. This ADR makes the issue cost real and **op-type-differentiated**,
> so composite-vs-primitive trade-offs (and CPU saturation) are measurable.
## Context
### What exists today
- Every `tl.*` op calls `_emit_dispatch_overhead()` before emitting its
command (`tl_context.py:123-125, 190, 227, 235, 612, …`), which emits
`PeCpuOverheadCmd(cycles=dispatch_cycles)` **only if** `dispatch_cycles
> 0`. So the issue cost is a **single uniform knob** applied identically
to `tl.load`, `tl.dot`, a MATH op, and `tl.composite`.
- That knob is hardcoded to **0** in both live execution paths
(`pe_cpu.py:101` greenlet runner, `:195` legacy replay). ⇒ issuing a
command — *constructing the descriptor and pushing it to the scheduler
queue* — currently costs **0 ns** on PE_CPU.
- `PeCpuOverheadCmd` is consumed as `yield env.timeout(cmd.cycles)` on
PE_CPU (`kernel_runner.py:131-132`) and on PE_SCHEDULER
(`pe_scheduler.py:97-100`); a manual `tl.cycles(n)` also exists
(`tl_context.py:695`).
### Two findings that frame this ADR (from ADR-0060 review)
- **Q1 — composite issue cost.** Constructing + pushing a `tl.composite`
is expected to cost on the order of **~40 ns** of CPU time (descriptor
build + queue push). Today it is **0**. The hook exists; only the value
(and its per-op-type differentiation) is missing.
- **Q2 — scheduler dispatch vs DMA latency.** `PE_SCHEDULER`'s composite
dispatch is **non-blocking**: `_dispatch_composite` generates the tile
plan and enqueues to the feeder, returning immediately
(`pe_scheduler.py:104-121`). The actual **DMA latency is charged on
PE_DMA** as each tile flows through (`drain_ns = compute_drain_ns(path,
nbytes)`, `pe_dma.py:89`), **not** lumped into the scheduler dispatch ⇒
**no double-counting**, and DMA stays on its modelled component
(SPEC §0.1). The scheduler's own plan-generation currently costs **0**
sim time.
### Why uniform-and-zero is wrong for the hybrid
The hybrid's whole argument is that **one** composite descriptor offloads
`N_tiles` worth of GEMM tiling, so the CPU issues `O(1)` coarse commands
instead of `O(N_tiles × ops/tile)` fine ones. With issue cost = 0 (and
uniform), the model cannot show:
- that the primitive path may **fail to saturate** the engines when the
CPU cannot push fast enough (the core ADR-0060 §1 claim), nor
- that a composite **costs more to construct** than a single primitive but
far less than the many primitives it replaces.
A single uniform `dispatch_cycles` cannot express this: `tl.load`
`tl.composite` in real construct cost.
## Decision
### D1. Per-op-type issue-cost table on PE_CPU
Replace the single `dispatch_cycles` scalar with a **cost table keyed by
command type**, charged on PE_CPU at issue (before the command is
dispatched to the scheduler). The greenlet pays this time, which is
exactly what gates how fast the CPU can push work — the saturation lever.
Starting estimates (calibrate later — see review items):
| Issued op | CPU issue cost (construct + push) |
|---|---|
| `tl.composite` (GEMM/MATH descriptor) | ~40 ns (Q1 estimate) |
| `tl.load` / `tl.store` (DMA descriptor) | small (a few ns) |
| `tl.dot` / MATH / IPCQ send/recv | small (a few ns) |
The point is the **ratio**: a composite construct ≫ a primitive issue, but
≪ the sum of the primitive issues it replaces. Exact ns are configurable.
### D2. Execution latency stays on the engines (Q2 — no change)
DMA/GEMM/MATH execution latency remains charged on PE_DMA / PE_GEMM /
PE_MATH as today. The issue cost (D1) is **CPU-side only** and additive; it
does **not** touch `drain_ns`, so there is no double-count. This preserves
SPEC §0.1 (latency on modelled components).
### D3. Scheduler plan-generation cost — start at 0, revisit
`PE_SCHEDULER`'s tile-plan generation costs 0 sim time today. Keep that for
now (the dominant lever is CPU issue cost, D1); expose it later via the
existing `overhead_ns` node attr if a scheduler-side cost proves material.
Marked as a review item, not decided here.
### D4. Configurable values; goldens regenerate
The cost table is configurable (per-topology / node attrs, default to the
D1 estimates). Turning issue cost non-zero changes **every** bench's
latency, so golden latencies are **regenerated** once — a model-fidelity
improvement, same posture as ADR-0062's global lazy-load regression.
## Alternatives
### A1. Keep a single uniform `dispatch_cycles > 0`
Rejected: op construct costs differ by an order of magnitude (`tl.load` vs
`tl.composite`); a uniform value either over-charges primitives or
under-charges composites, and in both cases misrepresents the
composite-vs-primitive trade-off the hybrid depends on.
### A2. Charge the issue cost on PE_SCHEDULER instead of PE_CPU
Rejected: the saturation question is *"can the CPU push descriptors fast
enough to keep the engines busy?"* — that is a **PE_CPU** issue-bandwidth
property. Charging it on the scheduler would not model CPU back-pressure.
### A3. Model DMA program/setup time as a separate fixed per-descriptor cost
Real HW pays a DMA descriptor program cost distinct from transfer time.
Deferred: initially fold the descriptor-program cost into the **issuing
op's** D1 cost (a `tl.load` / composite that triggers DMA). Split it out to
a PE_DMA fixed setup only if calibration shows it matters. Review item.
## Consequences
### Positive
- The hybrid's CPU-offload / saturation win (ADR-0060 §1) becomes
**measurable**, not just structural.
- Composite-vs-primitive and tiling-granularity trade-offs are visible in
latency, enabling the eval narrative.
- More faithful to hardware (issue bandwidth is a real bottleneck).
### Negative
- **All** bench goldens shift → one-time regeneration (D4); CI golden
fixtures update.
- Per-op-type values need calibration; the ~40 ns composite figure is an
estimate, primitives are unspecified — results are only as good as the
numbers (do not over-claim absolute latencies until calibrated).
- Adds a cost-table lookup on the issue path (negligible runtime).
## Open review items (decided autonomously; revise on review)
1. **Calibration source for the per-op-type values.** Composite ≈ 40 ns is
a working estimate; primitive issue costs are placeholders. *Recommend:*
pick values from a documented assumption (instruction-issue + queue
push) and treat absolute latency as provisional until a real reference
exists; keep the **ratios** defensible.
2. **Scheduler plan-gen cost (D3).** *Recommend:* keep 0 initially; expose
via `overhead_ns` if a workload shows scheduler-bound behaviour.
3. **DMA program time (A3).** *Recommend:* fold into the issuing op's cost
first; split to PE_DMA setup only if needed.
4. **Where the table lives.** *Recommend:* a small central cost module
keyed by command type, overridable per topology — not scattered node
attrs — so the values are reviewable in one place.
5. **Regression rollout.** *Recommend:* land lazy-load (ADR-0062) and this
ADR's cost model in one golden-regeneration pass to avoid two churns.
6. **Interaction with the legacy replay path** (`pe_cpu.py:_execute_legacy`)
— ensure both the greenlet and replay paths read the same cost table so
results match. Verify.
## Test Requirements
1. **Composite charges once; primitives charge per-op.** A kernel issuing
one `tl.composite` over `N_tiles` charges one composite issue cost on
PE_CPU; the equivalent `N_tiles × ops` primitive kernel charges the
per-op cost `N_tiles × ops` times. Assert the PE_CPU busy time differs
accordingly.
2. **DMA latency unchanged (Q2).** For a fixed transfer, PE_DMA `drain_ns`
is identical to the pre-ADR value — the issue cost is additive on
PE_CPU, not folded into DMA (no double-count).
3. **Saturation is observable.** With non-zero per-op issue cost, a
many-tile primitive sweep shows GEMM-engine idle (CPU-bound issue)
whereas the composite sweep keeps it busy — the ADR-0060 §1 lever.
4. **Determinism:** identical inputs → identical op_log + latency
(SPEC §0.1).
5. **Path parity:** greenlet and legacy-replay paths produce identical
issue-cost accounting for the same kernel.