- ADR-0060: GEMMs (Q.Kt, P.V) via existing tl.composite (scheduler-managed tiling + K/V DMA streaming); softmax merge + IPCQ tree reduction stay in kernel. Front TL;DR pseudocode of the final composite kernel; new section B lists open design items (DDD sync, K pre-transpose, dma_read lever, kernel-vs-scheduler tiling, ring path). - ADR-0062: redefined from a new load_async op to global lazy tl.load (non-blocking + auto-wait on first use; API unchanged; goldens regenerate). - ADR-0064 (new): per-op-type CPU issue cost model (composite ~40ns >> primitive) so the hybrid's CPU-saturation win becomes measurable (currently dispatch_cycles=0 hides it). Cost-model impl deferred. - KO mirrors for ADR-0060/0062/0064 (-ko suffix, adr-proposed). Rationale: non-blocking CompositeCmd offloads tiling to PE_SCHEDULER, decoupling CPU issue-rate from execution so the CPU can saturate the engines; the prior 'composite = no latency benefit' claim was an artifact of dispatch_cycles=0. Docs only; no production code changed. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
8.5 KiB
ADR-0064: Per-op-type CPU issue cost model (command construct + dispatch)
Status
Proposed
Supporting ADR for ADR-0060 (AHBM GQA Fused Attention). The hybrid decision there (GEMMs via
tl.composite, softmax merge in the kernel) wins by offloading tiling to PE_SCHEDULER so the CPU issues coarse descriptors and runs ahead, keeping the engines saturated. That win is currently invisible in the simulator because per-op CPU issue cost is zero. This ADR makes the issue cost real and op-type-differentiated, so composite-vs-primitive trade-offs (and CPU saturation) are measurable.
Context
What exists today
- Every
tl.*op calls_emit_dispatch_overhead()before emitting its command (tl_context.py:123-125, 190, 227, 235, 612, …), which emitsPeCpuOverheadCmd(cycles=dispatch_cycles)only if `dispatch_cycles0
. So the issue cost is a **single uniform knob** applied identically totl.load,tl.dot, a MATH op, andtl.composite`. - That knob is hardcoded to 0 in both live execution paths
(
pe_cpu.py:101greenlet runner,:195legacy replay). ⇒ issuing a command — constructing the descriptor and pushing it to the scheduler queue — currently costs 0 ns on PE_CPU. PeCpuOverheadCmdis consumed asyield env.timeout(cmd.cycles)on PE_CPU (kernel_runner.py:131-132) and on PE_SCHEDULER (pe_scheduler.py:97-100); a manualtl.cycles(n)also exists (tl_context.py:695).
Two findings that frame this ADR (from ADR-0060 review)
- Q1 — composite issue cost. Constructing + pushing a
tl.compositeis expected to cost on the order of ~40 ns of CPU time (descriptor build + queue push). Today it is 0. The hook exists; only the value (and its per-op-type differentiation) is missing. - Q2 — scheduler dispatch vs DMA latency.
PE_SCHEDULER's composite dispatch is non-blocking:_dispatch_compositegenerates the tile plan and enqueues to the feeder, returning immediately (pe_scheduler.py:104-121). The actual DMA latency is charged on PE_DMA as each tile flows through (drain_ns = compute_drain_ns(path, nbytes),pe_dma.py:89), not lumped into the scheduler dispatch ⇒ no double-counting, and DMA stays on its modelled component (SPEC §0.1). The scheduler's own plan-generation currently costs 0 sim time.
Why uniform-and-zero is wrong for the hybrid
The hybrid's whole argument is that one composite descriptor offloads
N_tiles worth of GEMM tiling, so the CPU issues O(1) coarse commands
instead of O(N_tiles × ops/tile) fine ones. With issue cost = 0 (and
uniform), the model cannot show:
- that the primitive path may fail to saturate the engines when the CPU cannot push fast enough (the core ADR-0060 §1 claim), nor
- that a composite costs more to construct than a single primitive but far less than the many primitives it replaces.
A single uniform dispatch_cycles cannot express this: tl.load ≪
tl.composite in real construct cost.
Decision
D1. Per-op-type issue-cost table on PE_CPU
Replace the single dispatch_cycles scalar with a cost table keyed by
command type, charged on PE_CPU at issue (before the command is
dispatched to the scheduler). The greenlet pays this time, which is
exactly what gates how fast the CPU can push work — the saturation lever.
Starting estimates (calibrate later — see review items):
| Issued op | CPU issue cost (construct + push) |
|---|---|
tl.composite (GEMM/MATH descriptor) |
~40 ns (Q1 estimate) |
tl.load / tl.store (DMA descriptor) |
small (a few ns) |
tl.dot / MATH / IPCQ send/recv |
small (a few ns) |
The point is the ratio: a composite construct ≫ a primitive issue, but ≪ the sum of the primitive issues it replaces. Exact ns are configurable.
D2. Execution latency stays on the engines (Q2 — no change)
DMA/GEMM/MATH execution latency remains charged on PE_DMA / PE_GEMM /
PE_MATH as today. The issue cost (D1) is CPU-side only and additive; it
does not touch drain_ns, so there is no double-count. This preserves
SPEC §0.1 (latency on modelled components).
D3. Scheduler plan-generation cost — start at 0, revisit
PE_SCHEDULER's tile-plan generation costs 0 sim time today. Keep that for
now (the dominant lever is CPU issue cost, D1); expose it later via the
existing overhead_ns node attr if a scheduler-side cost proves material.
Marked as a review item, not decided here.
D4. Configurable values; goldens regenerate
The cost table is configurable (per-topology / node attrs, default to the D1 estimates). Turning issue cost non-zero changes every bench's latency, so golden latencies are regenerated once — a model-fidelity improvement, same posture as ADR-0062's global lazy-load regression.
Alternatives
A1. Keep a single uniform dispatch_cycles > 0
Rejected: op construct costs differ by an order of magnitude (tl.load vs
tl.composite); a uniform value either over-charges primitives or
under-charges composites, and in both cases misrepresents the
composite-vs-primitive trade-off the hybrid depends on.
A2. Charge the issue cost on PE_SCHEDULER instead of PE_CPU
Rejected: the saturation question is "can the CPU push descriptors fast enough to keep the engines busy?" — that is a PE_CPU issue-bandwidth property. Charging it on the scheduler would not model CPU back-pressure.
A3. Model DMA program/setup time as a separate fixed per-descriptor cost
Real HW pays a DMA descriptor program cost distinct from transfer time.
Deferred: initially fold the descriptor-program cost into the issuing
op's D1 cost (a tl.load / composite that triggers DMA). Split it out to
a PE_DMA fixed setup only if calibration shows it matters. Review item.
Consequences
Positive
- The hybrid's CPU-offload / saturation win (ADR-0060 §1) becomes measurable, not just structural.
- Composite-vs-primitive and tiling-granularity trade-offs are visible in latency, enabling the eval narrative.
- More faithful to hardware (issue bandwidth is a real bottleneck).
Negative
- All bench goldens shift → one-time regeneration (D4); CI golden fixtures update.
- Per-op-type values need calibration; the ~40 ns composite figure is an estimate, primitives are unspecified — results are only as good as the numbers (do not over-claim absolute latencies until calibrated).
- Adds a cost-table lookup on the issue path (negligible runtime).
Open review items (decided autonomously; revise on review)
- Calibration source for the per-op-type values. Composite ≈ 40 ns is a working estimate; primitive issue costs are placeholders. Recommend: pick values from a documented assumption (instruction-issue + queue push) and treat absolute latency as provisional until a real reference exists; keep the ratios defensible.
- Scheduler plan-gen cost (D3). Recommend: keep 0 initially; expose
via
overhead_nsif a workload shows scheduler-bound behaviour. - DMA program time (A3). Recommend: fold into the issuing op's cost first; split to PE_DMA setup only if needed.
- Where the table lives. Recommend: a small central cost module keyed by command type, overridable per topology — not scattered node attrs — so the values are reviewable in one place.
- Regression rollout. Recommend: land lazy-load (ADR-0062) and this ADR's cost model in one golden-regeneration pass to avoid two churns.
- Interaction with the legacy replay path (
pe_cpu.py:_execute_legacy) — ensure both the greenlet and replay paths read the same cost table so results match. Verify.
Test Requirements
- Composite charges once; primitives charge per-op. A kernel issuing
one
tl.compositeoverN_tilescharges one composite issue cost on PE_CPU; the equivalentN_tiles × opsprimitive kernel charges the per-op costN_tiles × opstimes. Assert the PE_CPU busy time differs accordingly. - DMA latency unchanged (Q2). For a fixed transfer, PE_DMA
drain_nsis identical to the pre-ADR value — the issue cost is additive on PE_CPU, not folded into DMA (no double-count). - Saturation is observable. With non-zero per-op issue cost, a many-tile primitive sweep shows GEMM-engine idle (CPU-bound issue) whereas the composite sweep keeps it busy — the ADR-0060 §1 lever.
- Determinism: identical inputs → identical op_log + latency (SPEC §0.1).
- Path parity: greenlet and legacy-replay paths produce identical issue-cost accounting for the same kernel.