# ADR-0064: Structural CPU dispatch cost model (`logical_bytes` + FIXED + R) ## Status Proposed (Revision 2) > Supporting ADR for **ADR-0060** (AHBM GQA Fused Attention) and **ADR-0065** > (flat-ops composite + first stateful recipe). The hybrid decision there > (GEMMs via `tl.composite`, softmax merge in the kernel) wins by > **offloading tiling to PE_SCHEDULER so the CPU issues coarse descriptors > and runs ahead, keeping the engines saturated**. That win is currently > **invisible in the simulator** because per-op CPU issue cost is zero. > > Revision 2 replaces the **op-type calibration table** (the original > proposal) with a **structural formula** derived from each command's > `logical_bytes` — no per-op-type calibration needed; new op kinds are > covered automatically. ## Context ### What exists today - Every `tl.*` op calls `_emit_dispatch_overhead()` before emitting its command (`tl_context.py:196-212`), which emits `PeCpuOverheadCmd(cycles=dispatch_cycles)` **only if** `dispatch_cycles > 0`. The knob is **uniform** across op kinds and hardcoded to **0** in both live execution paths (`pe_cpu.py:101` greenlet, `:195` replay). - ⇒ issuing a command — *constructing the descriptor and pushing it to the scheduler queue* — currently costs **0 ns** on PE_CPU. - `PeCpuOverheadCmd` is consumed as `yield env.timeout(cmd.cycles)` on PE_CPU (`kernel_runner.py:131-132`). ### Why uniform-and-zero is wrong for the hybrid ADR-0060 §1's argument is that **one** composite descriptor offloads `N_tiles` worth of GEMM tiling, so the CPU issues `O(1)` coarse commands instead of `O(N_tiles × ops/tile)` fine ones. With issue cost = 0, the model cannot show: - that the primitive path may **fail to saturate** the engines when the CPU cannot push fast enough, nor - that a composite **costs more to construct** than a single primitive but far less than the many primitives it replaces. ### Why per-op-type calibration (Revision 1) was over-shaped The original proposal had a `cost_table[kind]` keyed by op kind (`composite`, `load`, `dot`, `math`, …). That required: - a value per kind (calibration cost ≥ |kinds|), - a new entry every time a new kind appears, - and yet the *ratio* it tried to capture — "composite ≫ primitive, but ≪ the primitives it replaces" — is structurally a function of **how many fields the command carries**, not of the op kind. A command's *byte footprint* is the natural proxy: a composite carrying N OpSpecs has ~N× the bytes of a primitive op with one OpSpec. The *fixed* part (queue head update, completion register, MMIO-class latency) is per-command. The two together compose: `FIXED + bytes × R`. ### What this model actually exposes The **primary** signal is **command-count reduction** through the per-command FIXED cost. The **byte** term is a secondary refinement that prevents pathologically-large composites from looking free. With realistic on-die queue bandwidth (16 B/cycle, D3), FIXED accounts for ≥85% of the dispatch cost differential between opt3 (≈10 cmds/tile) and opt2 (≈2 cmds/tile) for decode opt2. This framing also bounds the model: if a single composite were allowed to grow unboundedly large, the byte term alone would not stop the formula from rewarding ever-bigger fused commands beyond what real HW supports — hence the descriptor-size cap in **D7**. ## Decision ### D1. Structural dispatch cost formula Each PE command going to PE_SCHEDULER incurs PE_CPU dispatch cycles: ``` dispatch_cycles(cmd) = FIXED_PER_CMD + cmd.logical_bytes × R ``` where: - `FIXED_PER_CMD` (cycles per command) models queue-tail update, MMIO-class RTT, completion-event registration — fixed per command regardless of size. - `R` (cycles per byte) models the queue-write bandwidth — bytes of the command serialized into the scheduler queue. - `cmd.logical_bytes` (int) is each command's *HW-logical* byte size, computed from D2 below — not Python's `sys.getsizeof`. PE_CPU emits `PeCpuOverheadCmd(cycles=dispatch_cycles(cmd))` before dispatching, exactly as the existing hook (`tl_context.py:_emit_dispatch_ overhead`) — only the cycle value changes. ### D2. `logical_bytes` rule Each PE command dataclass exposes `logical_bytes: int` (property). The counting rule (HW-friendly, ignores Python overhead): | Field kind | Bytes | |---|---| | Command framing (cmd-type discriminator + completion id ref) | 4 | | Opcode (op kind enum) | 1 | | Enum (scope, etc.) | 1 | | `TensorHandle` reference (address only — shape/dtype assumed in descriptor table) | 8 | | Scalar (int/float) | 4 | | Tuple length marker | 1 | `CompositeCmd` recursively sums its `ops` and `rw_handles`: ```python @property def logical_bytes(self) -> int: return ( 4 # framing + 1 + sum(op.logical_bytes for op in self.ops) + 1 + 8 * len(self.rw_handles) ) ``` `OpSpec`: ```python @property def logical_bytes(self) -> int: return ( 1 + 1 # opcode + scope + 1 + 8 * len(self.operands) # named operand handles + (8 if self.out is not None else 0) # out handle + 1 + sum(_extra_bytes(v) for v in self.extra.values()) ) def _extra_bytes(v) -> int: """Type-aware byte count for OpSpec.extra values.""" if isinstance(v, bool): return 1 if isinstance(v, (int, float)): return 4 if isinstance(v, (tuple, list)): return 1 + 4 * len(v) # shape, axes, … if isinstance(v, str): return 1 # opcode-like tag return 4 # default scalar # Example types in extra: # m, k, n int → 4 each # reduce_axis int → 4 # shape=(64, 64) tuple → 1 + 8 = 9 # factor=1.0 float → 4 ``` (Identical rule for `DmaReadCmd`, `MathCmd`, etc. — one property per dataclass, ~3 lines each.) **Counting rule — per-op summation, no deduplication.** Each `TensorHandle` reference inside an `OpSpec` is counted independently, even when the same handle appears in multiple OpSpecs or in `rw_handles`. The `rw_handles` block is metadata for cross-composite hazard tracking (ADR-0065 D6.3) and is counted separately from operand references in `ops`. There is no deduplication. This matches HW reality: the descriptor encodes each operand slot as an independent address field, and the dispatcher tracks `rw_handles` as a distinct metadata block. Example: `OpSpec(kind="mul_bcast", operands={"src_a": O, "src_b": corr}, out=O)` counts handle `O` *twice* (once for `src_a`, once for `out`); if `O` is also in the enclosing CompositeCmd's `rw_handles`, it is counted a third time. ### D3. Defaults — anchored on a typical composite ≈ 43 ns Anchor: a **single-OpSpec composite for a DMA-staged GEMM path** — one `OpSpec(kind="gemm", ...)`; DMA stages are auto-inserted by PE_SCHEDULER from operand `space` (ADR-0065 D4) and **do not appear in `logical_bytes`** (the kernel does not issue them as separate cmds). Breakdown: ``` framing 4 ops tuple length 1 GEMM OpSpec 40 (opcode 1 + scope 1 + 1 + 2 handles 16 + out 8 + 1 + extra m/k/n 12) rw_handles tuple length 1 rw_handles content 8 (one RW handle for the output) ───────────────────────────── total ~54 bytes ``` Target dispatch = ~43 ns. On-die producer→consumer queue at 16 bytes/cycle (typical on-die descriptor queue width). ``` FIXED_PER_CMD = 40 cycles R = 0.0625 cycles/byte (= 16 bytes/cycle) ``` Verification: `40 + 54 × 0.0625 = 43.375 cycles ≈ 43 ns` ✓ Cycle→ns conversion uses the PE node's existing `clock_freq_ghz` attr (the same one used by PE_MATH `_compute_ns`). The cost-model knobs are **cycle-domain only** — they do not duplicate the clock setting. ### D4. Topology config override Defaults are baked into `pe_cpu.py`. Topology yaml may override under a `pe_cost_model:` section at the PE node attrs (cycle-domain knobs only; clock comes from the PE's existing `clock_freq_ghz`): ```yaml pe: attrs: clock_freq_ghz: 1.0 # existing, used for cycle→ns pe_cost_model: fixed_per_cmd_cycles: 40 byte_cycles_recip: 0.0625 # = 16 bytes/cycle max_composite_logical_bytes: 1024 # D7 — descriptor size cap ``` Missing keys fall back to defaults. The dispatch formula reads from `node.attrs["pe_cost_model"]` at PE_CPU init. ### D5. Scope — what does and does not pay | Path | Pays dispatch cost? | |---|---| | PE_CPU → PE_SCHEDULER for any `PeCommand` | **Yes** | | `PeCpuOverheadCmd` itself (already cycles-explicit) | **No** (formula bypass) | | Stages auto-generated by PE_SCHEDULER (DMA_READ/WRITE/FETCH/STORE) | **No** (PE_SCHEDULER-internal) | | Engine compute latency (DMA `drain_ns`, GEMM/MATH `_compute_ns`) | **No change** — stays on engines (SPEC §0.1) | This preserves the "latency on modelled components" invariant — dispatch cost is *additional* CPU-side time, not folded into engine times. ### D6. Configurable values; goldens regenerate Turning issue cost non-zero changes **every** bench's latency. Golden latencies are **regenerated once** when this ADR lands — same posture as ADR-0062 D3 lazy-load. After regeneration, the same calibration is in effect for ADR-0065 opt2 measurement. ### D7. Composite size cap (deterministic segmentation) Each `CompositeCmd`'s `logical_bytes` is capped at **`MAX_COMPOSITE_LOGICAL_BYTES`** (default **1024 bytes**, overridable per D4). Oversized commands are deterministically **segmented** by the *emitter* (host-side TLContext, ADR-0065 D5) into N consecutive `CompositeCmd`s, each ≤ cap. - Each segment carries its own `completion: CompletionHandle`. - Each segment incurs its own dispatch cost — `total = sum(FIXED + bytes_i × R) = N × FIXED + total_bytes × R`. The FIXED term is paid per segment. - Segments share `rw_handles` where applicable; **strict FIFO** (ADR-0065 D6.3) preserves write-after-write ordering automatically. - The segmentation algorithm is deterministic (greedy by op index in emit order); it is part of the host emitter, not PE_SCHEDULER. Why this is needed. Real hardware imposes hard limits — descriptor queue entry size, scheduler parser buffer, command SRAM, firmware input. Without a cap, the `FIXED + bytes × R` model would reward arbitrarily large fused composites beyond what hardware accepts (e.g., fusing 100 primitive ops into one composite, paying one `FIXED`). A composite of (say) 100 ops at ~30 bytes each = ~3000 bytes > 1024 cap → segmented into 3 ≈ 3 × FIXED, restoring honest accounting. Decode opt2's `#2` composite (10 ops, ~310 bytes) sits comfortably inside the 1024 cap — no segmentation for the GQA workload. ## Alternatives ### A1. Keep Revision 1's op-type calibration table Rejected: calibration cost scales with |kinds|, and the *ratio* the table tried to capture is structurally a function of cmd size. The structural formula reaches the same qualitative behaviour with two calibratable numbers instead of N. ### A2. Byte-only formula (no FIXED term) Rejected. With FIXED = 0, opt2 (Option Y per ADR-0065) does **not** win over opt3 — the total *bytes* dispatched per tile are similar (opt3 ≈ 232, opt2 ≈ 380); the win is entirely in *fewer per-cmd fixed costs*. A byte-only formula erases the very signal the model needs to expose. ### A3. Charge dispatch on PE_SCHEDULER instead of PE_CPU Rejected: the saturation question is *"can the CPU push descriptors fast enough to keep the engines busy?"* — that is a **PE_CPU** issue-bandwidth property. Charging on the scheduler would not model CPU back-pressure. ### A4. Model DMA program/setup time as a separate fixed per-descriptor cost Deferred: initially fold the descriptor-program cost into the **issuing op's** dispatch cost. Split it out to a PE_DMA fixed setup only if calibration shows it matters. ## Consequences ### Positive - Hybrid's CPU-offload / saturation win (ADR-0060 §1) becomes **measurable**, with a structurally honest model (no calibration table). - Adding new op kinds (e.g., ADR-0065's `softmax_merge` 8-step recipe) costs zero — they fit the same formula automatically. - More faithful to hardware (queue-head MMIO RTT + queue-write bandwidth). ### Negative - **All** bench goldens shift → one-time regeneration (D6); CI golden fixtures update. - Two calibration knobs (FIXED, R) need values; defaults are anchored on a documented assumption — treat absolute latencies as provisional until a reference exists; keep the **ratios** defensible. - Adds a small `logical_bytes` property to each PE command dataclass. ## Open review items 1. **Calibration source for FIXED and R.** Defaults from "typical composite ≈ 43 ns + on-die queue 16 bytes/cycle"; reasonable for an on-die descriptor queue. Revisit when a HW reference appears. 2. **Scheduler plan-gen cost vs large composites.** Stays 0 — D5 keeps PE_SCHEDULER's plan-generation outside the dispatch formula. The D7 cap (1024 bytes ≈ 30–35 OpSpecs) bounds the worst case, but a composite near the cap still costs the scheduler the same as a 1-op composite under the current zero-cost model. If a stress test (large-composite microbench) shows scheduler-bound behaviour, expose `overhead_ns` per-op-count. 3. **Where the override lives.** `pe_cost_model:` block under PE node attrs in topology yaml — keeps all knobs in one place, reviewable. Clock comes from the PE's existing `clock_freq_ghz` attr, not duplicated here. 4. **Path parity.** Both greenlet (`_execute_legacy` and `kernel_runner`) and replay paths must read the same cost model. Verify. 5. **Sensitivity of conclusions to R.** opt2 < opt3 must hold across a reasonable range of `R` (queue-bandwidth assumption). Sensitivity sweep is part of Test Requirements (#9). ## Test Requirements 1. **Anchor preservation.** A single-OpSpec composite for a DMA-staged GEMM path (`logical_bytes ≈ 54`) dispatches in ≈43 ns at default values (within ±2 ns). 2. **Structural ratio.** opt3 vs opt2 dispatch (per ADR-0065 §verification): `opt3 / opt2 ≈ 4.0×` at default calibration. 3. **Override path.** Topology yaml `pe_cost_model:` block changes the per-PE dispatch cost; default is recovered when block is missing. 4. **`PeCpuOverheadCmd` bypass.** Manual `tl.cycles(n)` issues exactly `n` cycles, not `n + dispatch_cycles(...)`. 5. **No double-count.** PE_DMA `drain_ns`, PE_GEMM/MATH `_compute_ns` identical to pre-ADR values. 6. **Determinism.** Identical inputs → identical op_log + latency (SPEC §0.1). 7. **Path parity.** Greenlet and replay paths produce identical dispatch-cycle accounting for the same kernel. 8. **Composite size cap (D7).** A recipe that would emit `logical_bytes > MAX_COMPOSITE_LOGICAL_BYTES` is segmented into N consecutive `CompositeCmd`s; total dispatch = sum of segment dispatches; RW ordering across segments preserved by strict FIFO. 9. **Sensitivity sweep.** At `R ∈ {0.25, 0.0625, 0.03125}` cycles/byte (= 4 / 16 / 32 bytes/cycle), the conclusion *opt2 per-tile dispatch < opt3 per-tile dispatch* must hold (ratio monotonically increases as R decreases — FIXED dominates more). ## Migration ADR-0064 Revision 2 lands as a single PR with: - `logical_bytes` property on each `PeCommand` dataclass (type-aware extra-field counting per D2) - formula application in `pe_cpu.py` dispatch path - `pe_cost_model:` override read at PE_CPU init (cycle-domain knobs + `max_composite_logical_bytes`) - **composite size cap (D7)** — TLContext-side segmentation logic with `MAX_COMPOSITE_LOGICAL_BYTES` default 1024; not needed for any existing bench (largest current composite is well under 200 bytes), but the mechanism lands so it is in place when ADR-0065's 10-op decode-opt2 composite (~310 bytes) and larger recipes appear. - one-time goldens regeneration After this lands, ADR-0065 builds on top with no further goldens churn in existing benches (ADR-0065 is a meaning-preserving refactor of `CompositeCmd` for the existing path; only opt2 is a new bench).