Files
kernbench2/docs/adr-proposed/ADR-0064-perf-cpu-issue-cost-model.md
T
ywkang a8c50238c6 gqa(adr): ADR-0064 Rev2 (structural dispatch cost) + ADR-0065/DDD-0065 (flat-ops composite + softmax_merge recipe)
ADR-0064 Revision 2: replace per-op-type calibration table with a
structural formula `FIXED_PER_CMD + cmd.logical_bytes × R`, defaults
anchored at typical composite ≈ 45 ns. Topology yaml override under
`pe_cost_model:` block; logical_bytes property per PE command.

ADR-0065: implement ADR-0060 §5.6 / §8 item 4 carve-out as a flat-ops
CompositeCmd (no head/epilogue structural fields — position + scope
drives placement) + first stateful recipe `softmax_merge` (MATH-only
8-step). RECIPE_DESCRIPTORS lives in TLContext-adjacent module only;
PE_SCHEDULER stays recipe-free and auto-inserts DMAs from operand
`space`. Strict-FIFO RW hazard tracker; ≤1 GEMM per composite
invariant. User-facing `tl.composite(prologue=[...], op=, epilogue=[...])`
API preserved; existing benches unchanged (meaning-preserving refactor).

DDD-0065: implementation-ready phased plan (P0=ADR-0064 Rev2 first,
P1-P6 for ADR-0065), file plan, recipe engine sequence, scheduler
plan-gen algorithm, RW tracker design, test matrix.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-06-10 16:11:20 -07:00

259 lines
10 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# ADR-0064: Structural CPU dispatch cost model (`logical_bytes` + FIXED + R)
## Status
Proposed (Revision 2)
> Supporting ADR for **ADR-0060** (AHBM GQA Fused Attention) and **ADR-0065**
> (flat-ops composite + first stateful recipe). The hybrid decision there
> (GEMMs via `tl.composite`, softmax merge in the kernel) wins by
> **offloading tiling to PE_SCHEDULER so the CPU issues coarse descriptors
> and runs ahead, keeping the engines saturated**. That win is currently
> **invisible in the simulator** because per-op CPU issue cost is zero.
>
> Revision 2 replaces the **op-type calibration table** (the original
> proposal) with a **structural formula** derived from each command's
> `logical_bytes` — no per-op-type calibration needed; new op kinds are
> covered automatically.
## Context
### What exists today
- Every `tl.*` op calls `_emit_dispatch_overhead()` before emitting its
command (`tl_context.py:196-212`), which emits
`PeCpuOverheadCmd(cycles=dispatch_cycles)` **only if** `dispatch_cycles
> 0`. The knob is **uniform** across op kinds and hardcoded to **0**
in both live execution paths (`pe_cpu.py:101` greenlet, `:195` replay).
- ⇒ issuing a command — *constructing the descriptor and pushing it to
the scheduler queue* — currently costs **0 ns** on PE_CPU.
- `PeCpuOverheadCmd` is consumed as `yield env.timeout(cmd.cycles)` on
PE_CPU (`kernel_runner.py:131-132`).
### Why uniform-and-zero is wrong for the hybrid
ADR-0060 §1's argument is that **one** composite descriptor offloads
`N_tiles` worth of GEMM tiling, so the CPU issues `O(1)` coarse commands
instead of `O(N_tiles × ops/tile)` fine ones. With issue cost = 0, the
model cannot show:
- that the primitive path may **fail to saturate** the engines when the
CPU cannot push fast enough, nor
- that a composite **costs more to construct** than a single primitive
but far less than the many primitives it replaces.
### Why per-op-type calibration (Revision 1) was over-shaped
The original proposal had a `cost_table[kind]` keyed by op kind
(`composite`, `load`, `dot`, `math`, …). That required:
- a value per kind (calibration cost ≥ |kinds|),
- a new entry every time a new kind appears,
- and yet the *ratio* it tried to capture — "composite ≫ primitive,
but ≪ the primitives it replaces" — is structurally a function of
**how many fields the command carries**, not of the op kind.
A command's *byte footprint* is the natural proxy: a composite carrying
N OpSpecs has ~N× the bytes of a primitive op with one OpSpec. The
*fixed* part (queue head update, completion register, MMIO-class
latency) is per-command. The two together compose: `FIXED + bytes × R`.
## Decision
### D1. Structural dispatch cost formula
Each PE command going to PE_SCHEDULER incurs PE_CPU dispatch cycles:
```
dispatch_cycles(cmd) = FIXED_PER_CMD + cmd.logical_bytes × R
```
where:
- `FIXED_PER_CMD` (cycles per command) models queue-tail update, MMIO-class
RTT, completion-event registration — fixed per command regardless of size.
- `R` (cycles per byte) models the queue-write bandwidth — bytes of the
command serialized into the scheduler queue.
- `cmd.logical_bytes` (int) is each command's *HW-logical* byte size,
computed from D2 below — not Python's `sys.getsizeof`.
PE_CPU emits `PeCpuOverheadCmd(cycles=dispatch_cycles(cmd))` before
dispatching, exactly as the existing hook (`tl_context.py:_emit_dispatch_
overhead`) — only the cycle value changes.
### D2. `logical_bytes` rule
Each PE command dataclass exposes `logical_bytes: int` (property). The
counting rule (HW-friendly, ignores Python overhead):
| Field kind | Bytes |
|---|---|
| Command framing (cmd-type discriminator + completion id ref) | 4 |
| Opcode (op kind enum) | 1 |
| Enum (scope, etc.) | 1 |
| `TensorHandle` reference (address only — shape/dtype assumed in descriptor table) | 8 |
| Scalar (int/float) | 4 |
| Tuple length marker | 1 |
`CompositeCmd` recursively sums its `ops` and `rw_handles`:
```python
@property
def logical_bytes(self) -> int:
return (
4 # framing
+ 1 + sum(op.logical_bytes for op in self.ops)
+ 1 + 8 * len(self.rw_handles)
)
```
`OpSpec`:
```python
@property
def logical_bytes(self) -> int:
return (
1 + 1 # opcode + scope
+ 1 + 8 * len(self.operands) # named operand handles
+ (8 if self.out is not None else 0) # out handle
+ 1 + sum(4 for _ in self.extra.values()) # extra scalars
)
```
(Identical rule for `DmaReadCmd`, `MathCmd`, etc. — one property per
dataclass, ~3 lines each.)
### D3. Defaults — anchored on a typical composite ≈ 45 ns
Anchor: a typical 1-op DMA→GEMM→DMA composite has `logical_bytes ≈ 52`
(framing 4 + GEMM OpSpec 39 + rw_handles 9). Target dispatch = 45 ns.
On-die producer→consumer queue assumed at 4 bytes/cycle.
```
clock = 1 GHz # 1 cycle = 1 ns
FIXED_PER_CMD = 32 cycles
R = 0.25 cycles/byte
```
Verification: `32 + 52 × 0.25 = 45 cycles ≈ 45 ns`
### D4. Topology config override
Defaults are baked into `pe_cpu.py`. Topology yaml may override under a
`pe_cost_model:` section at the PE node attrs:
```yaml
pe:
attrs:
pe_cost_model:
fixed_per_cmd_cycles: 32
byte_cycles_recip: 0.25
clock_freq_ghz: 1.0
```
Missing keys fall back to defaults. The dispatch formula reads from
`node.attrs["pe_cost_model"]` at PE_CPU init.
### D5. Scope — what does and does not pay
| Path | Pays dispatch cost? |
|---|---|
| PE_CPU → PE_SCHEDULER for any `PeCommand` | **Yes** |
| `PeCpuOverheadCmd` itself (already cycles-explicit) | **No** (formula bypass) |
| Stages auto-generated by PE_SCHEDULER (DMA_READ/WRITE/FETCH/STORE) | **No** (PE_SCHEDULER-internal) |
| Engine compute latency (DMA `drain_ns`, GEMM/MATH `_compute_ns`) | **No change** — stays on engines (SPEC §0.1) |
This preserves the "latency on modelled components" invariant — dispatch
cost is *additional* CPU-side time, not folded into engine times.
### D6. Configurable values; goldens regenerate
Turning issue cost non-zero changes **every** bench's latency. Golden
latencies are **regenerated once** when this ADR lands — same posture as
ADR-0062 D3 lazy-load. After regeneration, the same calibration is in
effect for ADR-0065 opt2 measurement.
## Alternatives
### A1. Keep Revision 1's op-type calibration table
Rejected: calibration cost scales with |kinds|, and the *ratio* the
table tried to capture is structurally a function of cmd size. The
structural formula reaches the same qualitative behaviour with two
calibratable numbers instead of N.
### A2. Byte-only formula (no FIXED term)
Rejected. With FIXED = 0, opt2 (Option Y per ADR-0065) does **not** win
over opt3 — the total *bytes* dispatched per tile are similar (opt3 ≈ 232,
opt2 ≈ 380); the win is entirely in *fewer per-cmd fixed costs*. A
byte-only formula erases the very signal the model needs to expose.
### A3. Charge dispatch on PE_SCHEDULER instead of PE_CPU
Rejected: the saturation question is *"can the CPU push descriptors fast
enough to keep the engines busy?"* — that is a **PE_CPU** issue-bandwidth
property. Charging on the scheduler would not model CPU back-pressure.
### A4. Model DMA program/setup time as a separate fixed per-descriptor cost
Deferred: initially fold the descriptor-program cost into the **issuing
op's** dispatch cost. Split it out to a PE_DMA fixed setup only if
calibration shows it matters.
## Consequences
### Positive
- Hybrid's CPU-offload / saturation win (ADR-0060 §1) becomes
**measurable**, with a structurally honest model (no calibration table).
- Adding new op kinds (e.g., ADR-0065's `softmax_merge` 8-step recipe)
costs zero — they fit the same formula automatically.
- More faithful to hardware (queue-head MMIO RTT + queue-write bandwidth).
### Negative
- **All** bench goldens shift → one-time regeneration (D6); CI golden
fixtures update.
- Two calibration knobs (FIXED, R) need values; defaults are anchored on
a documented assumption — treat absolute latencies as provisional until
a reference exists; keep the **ratios** defensible.
- Adds a small `logical_bytes` property to each PE command dataclass.
## Open review items
1. **Calibration source for FIXED and R.** Defaults from "typical
composite = 45 ns + on-die queue 4 bytes/cycle"; reasonable for an
on-die producer→consumer queue. Revisit when a HW reference appears.
2. **Scheduler plan-gen cost.** Stays 0 — D5 keeps PE_SCHEDULER's
plan-generation outside the dispatch formula. Expose via existing
`overhead_ns` if a workload shows scheduler-bound behaviour.
3. **Where the override lives.** `pe_cost_model:` block under PE node
attrs in topology yaml — keeps all knobs in one place, reviewable.
4. **Path parity.** Both greenlet (`_execute_legacy` and `kernel_runner`)
and replay paths must read the same cost model. Verify.
## Test Requirements
1. **Anchor preservation.** A typical DMA→GEMM→DMA composite (1 op,
`logical_bytes ≈ 52`) dispatches in 45 ns at the default values.
2. **Structural ratio.** opt3 vs opt2 dispatch (per ADR-0065 §verification):
`opt3 / opt2 ≈ 2.4×` at default calibration.
3. **Override path.** Topology yaml `pe_cost_model:` block changes the
per-PE dispatch cost; default is recovered when block is missing.
4. **`PeCpuOverheadCmd` bypass.** Manual `tl.cycles(n)` issues exactly `n`
cycles, not `n + dispatch_cycles(...)`.
5. **No double-count.** PE_DMA `drain_ns`, PE_GEMM/MATH `_compute_ns`
identical to pre-ADR values.
6. **Determinism.** Identical inputs → identical op_log + latency
(SPEC §0.1).
7. **Path parity.** Greenlet and replay paths produce identical
dispatch-cycle accounting for the same kernel.
## Migration
ADR-0064 Revision 2 lands as a single PR with:
- `logical_bytes` property on each `PeCommand` dataclass
- formula application in `pe_cpu.py` dispatch path
- `pe_cost_model:` override read at PE_CPU init
- one-time goldens regeneration
After this lands, ADR-0065 builds on top with no further goldens churn
in existing benches (ADR-0065 is a meaning-preserving refactor of
`CompositeCmd` for the existing path; only opt2 is a new bench).