# ADR-0064: Per-op-type CPU issue cost model (command construct + dispatch) ## Status Proposed > Supporting ADR for **ADR-0060** (AHBM GQA Fused Attention). The hybrid > decision there (GEMMs via `tl.composite`, softmax merge in the kernel) > wins by **offloading tiling to PE_SCHEDULER so the CPU issues coarse > descriptors and runs ahead, keeping the engines saturated**. That win is > currently **invisible in the simulator** because per-op CPU issue cost is > zero. This ADR makes the issue cost real and **op-type-differentiated**, > so composite-vs-primitive trade-offs (and CPU saturation) are measurable. ## Context ### What exists today - Every `tl.*` op calls `_emit_dispatch_overhead()` before emitting its command (`tl_context.py:123-125, 190, 227, 235, 612, …`), which emits `PeCpuOverheadCmd(cycles=dispatch_cycles)` **only if** `dispatch_cycles > 0`. So the issue cost is a **single uniform knob** applied identically to `tl.load`, `tl.dot`, a MATH op, and `tl.composite`. - That knob is hardcoded to **0** in both live execution paths (`pe_cpu.py:101` greenlet runner, `:195` legacy replay). ⇒ issuing a command — *constructing the descriptor and pushing it to the scheduler queue* — currently costs **0 ns** on PE_CPU. - `PeCpuOverheadCmd` is consumed as `yield env.timeout(cmd.cycles)` on PE_CPU (`kernel_runner.py:131-132`) and on PE_SCHEDULER (`pe_scheduler.py:97-100`); a manual `tl.cycles(n)` also exists (`tl_context.py:695`). ### Two findings that frame this ADR (from ADR-0060 review) - **Q1 — composite issue cost.** Constructing + pushing a `tl.composite` is expected to cost on the order of **~40 ns** of CPU time (descriptor build + queue push). Today it is **0**. The hook exists; only the value (and its per-op-type differentiation) is missing. - **Q2 — scheduler dispatch vs DMA latency.** `PE_SCHEDULER`'s composite dispatch is **non-blocking**: `_dispatch_composite` generates the tile plan and enqueues to the feeder, returning immediately (`pe_scheduler.py:104-121`). The actual **DMA latency is charged on PE_DMA** as each tile flows through (`drain_ns = compute_drain_ns(path, nbytes)`, `pe_dma.py:89`), **not** lumped into the scheduler dispatch ⇒ **no double-counting**, and DMA stays on its modelled component (SPEC §0.1). The scheduler's own plan-generation currently costs **0** sim time. ### Why uniform-and-zero is wrong for the hybrid The hybrid's whole argument is that **one** composite descriptor offloads `N_tiles` worth of GEMM tiling, so the CPU issues `O(1)` coarse commands instead of `O(N_tiles × ops/tile)` fine ones. With issue cost = 0 (and uniform), the model cannot show: - that the primitive path may **fail to saturate** the engines when the CPU cannot push fast enough (the core ADR-0060 §1 claim), nor - that a composite **costs more to construct** than a single primitive but far less than the many primitives it replaces. A single uniform `dispatch_cycles` cannot express this: `tl.load` ≪ `tl.composite` in real construct cost. ## Decision ### D1. Per-op-type issue-cost table on PE_CPU Replace the single `dispatch_cycles` scalar with a **cost table keyed by command type**, charged on PE_CPU at issue (before the command is dispatched to the scheduler). The greenlet pays this time, which is exactly what gates how fast the CPU can push work — the saturation lever. Starting estimates (calibrate later — see review items): | Issued op | CPU issue cost (construct + push) | |---|---| | `tl.composite` (GEMM/MATH descriptor) | ~40 ns (Q1 estimate) | | `tl.load` / `tl.store` (DMA descriptor) | small (a few ns) | | `tl.dot` / MATH / IPCQ send/recv | small (a few ns) | The point is the **ratio**: a composite construct ≫ a primitive issue, but ≪ the sum of the primitive issues it replaces. Exact ns are configurable. ### D2. Execution latency stays on the engines (Q2 — no change) DMA/GEMM/MATH execution latency remains charged on PE_DMA / PE_GEMM / PE_MATH as today. The issue cost (D1) is **CPU-side only** and additive; it does **not** touch `drain_ns`, so there is no double-count. This preserves SPEC §0.1 (latency on modelled components). ### D3. Scheduler plan-generation cost — start at 0, revisit `PE_SCHEDULER`'s tile-plan generation costs 0 sim time today. Keep that for now (the dominant lever is CPU issue cost, D1); expose it later via the existing `overhead_ns` node attr if a scheduler-side cost proves material. Marked as a review item, not decided here. ### D4. Configurable values; goldens regenerate The cost table is configurable (per-topology / node attrs, default to the D1 estimates). Turning issue cost non-zero changes **every** bench's latency, so golden latencies are **regenerated** once — a model-fidelity improvement, same posture as ADR-0062's global lazy-load regression. ## Alternatives ### A1. Keep a single uniform `dispatch_cycles > 0` Rejected: op construct costs differ by an order of magnitude (`tl.load` vs `tl.composite`); a uniform value either over-charges primitives or under-charges composites, and in both cases misrepresents the composite-vs-primitive trade-off the hybrid depends on. ### A2. Charge the issue cost on PE_SCHEDULER instead of PE_CPU Rejected: the saturation question is *"can the CPU push descriptors fast enough to keep the engines busy?"* — that is a **PE_CPU** issue-bandwidth property. Charging it on the scheduler would not model CPU back-pressure. ### A3. Model DMA program/setup time as a separate fixed per-descriptor cost Real HW pays a DMA descriptor program cost distinct from transfer time. Deferred: initially fold the descriptor-program cost into the **issuing op's** D1 cost (a `tl.load` / composite that triggers DMA). Split it out to a PE_DMA fixed setup only if calibration shows it matters. Review item. ## Consequences ### Positive - The hybrid's CPU-offload / saturation win (ADR-0060 §1) becomes **measurable**, not just structural. - Composite-vs-primitive and tiling-granularity trade-offs are visible in latency, enabling the eval narrative. - More faithful to hardware (issue bandwidth is a real bottleneck). ### Negative - **All** bench goldens shift → one-time regeneration (D4); CI golden fixtures update. - Per-op-type values need calibration; the ~40 ns composite figure is an estimate, primitives are unspecified — results are only as good as the numbers (do not over-claim absolute latencies until calibrated). - Adds a cost-table lookup on the issue path (negligible runtime). ## Open review items (decided autonomously; revise on review) 1. **Calibration source for the per-op-type values.** Composite ≈ 40 ns is a working estimate; primitive issue costs are placeholders. *Recommend:* pick values from a documented assumption (instruction-issue + queue push) and treat absolute latency as provisional until a real reference exists; keep the **ratios** defensible. 2. **Scheduler plan-gen cost (D3).** *Recommend:* keep 0 initially; expose via `overhead_ns` if a workload shows scheduler-bound behaviour. 3. **DMA program time (A3).** *Recommend:* fold into the issuing op's cost first; split to PE_DMA setup only if needed. 4. **Where the table lives.** *Recommend:* a small central cost module keyed by command type, overridable per topology — not scattered node attrs — so the values are reviewable in one place. 5. **Regression rollout.** *Recommend:* land lazy-load (ADR-0062) and this ADR's cost model in one golden-regeneration pass to avoid two churns. 6. **Interaction with the legacy replay path** (`pe_cpu.py:_execute_legacy`) — ensure both the greenlet and replay paths read the same cost table so results match. Verify. ## Test Requirements 1. **Composite charges once; primitives charge per-op.** A kernel issuing one `tl.composite` over `N_tiles` charges one composite issue cost on PE_CPU; the equivalent `N_tiles × ops` primitive kernel charges the per-op cost `N_tiles × ops` times. Assert the PE_CPU busy time differs accordingly. 2. **DMA latency unchanged (Q2).** For a fixed transfer, PE_DMA `drain_ns` is identical to the pre-ADR value — the issue cost is additive on PE_CPU, not folded into DMA (no double-count). 3. **Saturation is observable.** With non-zero per-op issue cost, a many-tile primitive sweep shows GEMM-engine idle (CPU-bound issue) whereas the composite sweep keeps it busy — the ADR-0060 §1 lever. 4. **Determinism:** identical inputs → identical op_log + latency (SPEC §0.1). 5. **Path parity:** greenlet and legacy-replay paths produce identical issue-cost accounting for the same kernel.