7a9d4ec47b
- ADR-0060: GEMMs (Q.Kt, P.V) via existing tl.composite (scheduler-managed tiling + K/V DMA streaming); softmax merge + IPCQ tree reduction stay in kernel. Front TL;DR pseudocode of the final composite kernel; new section B lists open design items (DDD sync, K pre-transpose, dma_read lever, kernel-vs-scheduler tiling, ring path). - ADR-0062: redefined from a new load_async op to global lazy tl.load (non-blocking + auto-wait on first use; API unchanged; goldens regenerate). - ADR-0064 (new): per-op-type CPU issue cost model (composite ~40ns >> primitive) so the hybrid's CPU-saturation win becomes measurable (currently dispatch_cycles=0 hides it). Cost-model impl deferred. - KO mirrors for ADR-0060/0062/0064 (-ko suffix, adr-proposed). Rationale: non-blocking CompositeCmd offloads tiling to PE_SCHEDULER, decoupling CPU issue-rate from execution so the CPU can saturate the engines; the prior 'composite = no latency benefit' claim was an artifact of dispatch_cycles=0. Docs only; no production code changed. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
180 lines
8.5 KiB
Markdown
180 lines
8.5 KiB
Markdown
# ADR-0064: Per-op-type CPU issue cost model (command construct + dispatch)
|
||
|
||
## Status
|
||
|
||
Proposed
|
||
|
||
> Supporting ADR for **ADR-0060** (AHBM GQA Fused Attention). The hybrid
|
||
> decision there (GEMMs via `tl.composite`, softmax merge in the kernel)
|
||
> wins by **offloading tiling to PE_SCHEDULER so the CPU issues coarse
|
||
> descriptors and runs ahead, keeping the engines saturated**. That win is
|
||
> currently **invisible in the simulator** because per-op CPU issue cost is
|
||
> zero. This ADR makes the issue cost real and **op-type-differentiated**,
|
||
> so composite-vs-primitive trade-offs (and CPU saturation) are measurable.
|
||
|
||
## Context
|
||
|
||
### What exists today
|
||
|
||
- Every `tl.*` op calls `_emit_dispatch_overhead()` before emitting its
|
||
command (`tl_context.py:123-125, 190, 227, 235, 612, …`), which emits
|
||
`PeCpuOverheadCmd(cycles=dispatch_cycles)` **only if** `dispatch_cycles
|
||
> 0`. So the issue cost is a **single uniform knob** applied identically
|
||
to `tl.load`, `tl.dot`, a MATH op, and `tl.composite`.
|
||
- That knob is hardcoded to **0** in both live execution paths
|
||
(`pe_cpu.py:101` greenlet runner, `:195` legacy replay). ⇒ issuing a
|
||
command — *constructing the descriptor and pushing it to the scheduler
|
||
queue* — currently costs **0 ns** on PE_CPU.
|
||
- `PeCpuOverheadCmd` is consumed as `yield env.timeout(cmd.cycles)` on
|
||
PE_CPU (`kernel_runner.py:131-132`) and on PE_SCHEDULER
|
||
(`pe_scheduler.py:97-100`); a manual `tl.cycles(n)` also exists
|
||
(`tl_context.py:695`).
|
||
|
||
### Two findings that frame this ADR (from ADR-0060 review)
|
||
|
||
- **Q1 — composite issue cost.** Constructing + pushing a `tl.composite`
|
||
is expected to cost on the order of **~40 ns** of CPU time (descriptor
|
||
build + queue push). Today it is **0**. The hook exists; only the value
|
||
(and its per-op-type differentiation) is missing.
|
||
- **Q2 — scheduler dispatch vs DMA latency.** `PE_SCHEDULER`'s composite
|
||
dispatch is **non-blocking**: `_dispatch_composite` generates the tile
|
||
plan and enqueues to the feeder, returning immediately
|
||
(`pe_scheduler.py:104-121`). The actual **DMA latency is charged on
|
||
PE_DMA** as each tile flows through (`drain_ns = compute_drain_ns(path,
|
||
nbytes)`, `pe_dma.py:89`), **not** lumped into the scheduler dispatch ⇒
|
||
**no double-counting**, and DMA stays on its modelled component
|
||
(SPEC §0.1). The scheduler's own plan-generation currently costs **0**
|
||
sim time.
|
||
|
||
### Why uniform-and-zero is wrong for the hybrid
|
||
|
||
The hybrid's whole argument is that **one** composite descriptor offloads
|
||
`N_tiles` worth of GEMM tiling, so the CPU issues `O(1)` coarse commands
|
||
instead of `O(N_tiles × ops/tile)` fine ones. With issue cost = 0 (and
|
||
uniform), the model cannot show:
|
||
- that the primitive path may **fail to saturate** the engines when the
|
||
CPU cannot push fast enough (the core ADR-0060 §1 claim), nor
|
||
- that a composite **costs more to construct** than a single primitive but
|
||
far less than the many primitives it replaces.
|
||
|
||
A single uniform `dispatch_cycles` cannot express this: `tl.load` ≪
|
||
`tl.composite` in real construct cost.
|
||
|
||
## Decision
|
||
|
||
### D1. Per-op-type issue-cost table on PE_CPU
|
||
|
||
Replace the single `dispatch_cycles` scalar with a **cost table keyed by
|
||
command type**, charged on PE_CPU at issue (before the command is
|
||
dispatched to the scheduler). The greenlet pays this time, which is
|
||
exactly what gates how fast the CPU can push work — the saturation lever.
|
||
|
||
Starting estimates (calibrate later — see review items):
|
||
|
||
| Issued op | CPU issue cost (construct + push) |
|
||
|---|---|
|
||
| `tl.composite` (GEMM/MATH descriptor) | ~40 ns (Q1 estimate) |
|
||
| `tl.load` / `tl.store` (DMA descriptor) | small (a few ns) |
|
||
| `tl.dot` / MATH / IPCQ send/recv | small (a few ns) |
|
||
|
||
The point is the **ratio**: a composite construct ≫ a primitive issue, but
|
||
≪ the sum of the primitive issues it replaces. Exact ns are configurable.
|
||
|
||
### D2. Execution latency stays on the engines (Q2 — no change)
|
||
|
||
DMA/GEMM/MATH execution latency remains charged on PE_DMA / PE_GEMM /
|
||
PE_MATH as today. The issue cost (D1) is **CPU-side only** and additive; it
|
||
does **not** touch `drain_ns`, so there is no double-count. This preserves
|
||
SPEC §0.1 (latency on modelled components).
|
||
|
||
### D3. Scheduler plan-generation cost — start at 0, revisit
|
||
|
||
`PE_SCHEDULER`'s tile-plan generation costs 0 sim time today. Keep that for
|
||
now (the dominant lever is CPU issue cost, D1); expose it later via the
|
||
existing `overhead_ns` node attr if a scheduler-side cost proves material.
|
||
Marked as a review item, not decided here.
|
||
|
||
### D4. Configurable values; goldens regenerate
|
||
|
||
The cost table is configurable (per-topology / node attrs, default to the
|
||
D1 estimates). Turning issue cost non-zero changes **every** bench's
|
||
latency, so golden latencies are **regenerated** once — a model-fidelity
|
||
improvement, same posture as ADR-0062's global lazy-load regression.
|
||
|
||
## Alternatives
|
||
|
||
### A1. Keep a single uniform `dispatch_cycles > 0`
|
||
|
||
Rejected: op construct costs differ by an order of magnitude (`tl.load` vs
|
||
`tl.composite`); a uniform value either over-charges primitives or
|
||
under-charges composites, and in both cases misrepresents the
|
||
composite-vs-primitive trade-off the hybrid depends on.
|
||
|
||
### A2. Charge the issue cost on PE_SCHEDULER instead of PE_CPU
|
||
|
||
Rejected: the saturation question is *"can the CPU push descriptors fast
|
||
enough to keep the engines busy?"* — that is a **PE_CPU** issue-bandwidth
|
||
property. Charging it on the scheduler would not model CPU back-pressure.
|
||
|
||
### A3. Model DMA program/setup time as a separate fixed per-descriptor cost
|
||
|
||
Real HW pays a DMA descriptor program cost distinct from transfer time.
|
||
Deferred: initially fold the descriptor-program cost into the **issuing
|
||
op's** D1 cost (a `tl.load` / composite that triggers DMA). Split it out to
|
||
a PE_DMA fixed setup only if calibration shows it matters. Review item.
|
||
|
||
## Consequences
|
||
|
||
### Positive
|
||
- The hybrid's CPU-offload / saturation win (ADR-0060 §1) becomes
|
||
**measurable**, not just structural.
|
||
- Composite-vs-primitive and tiling-granularity trade-offs are visible in
|
||
latency, enabling the eval narrative.
|
||
- More faithful to hardware (issue bandwidth is a real bottleneck).
|
||
|
||
### Negative
|
||
- **All** bench goldens shift → one-time regeneration (D4); CI golden
|
||
fixtures update.
|
||
- Per-op-type values need calibration; the ~40 ns composite figure is an
|
||
estimate, primitives are unspecified — results are only as good as the
|
||
numbers (do not over-claim absolute latencies until calibrated).
|
||
- Adds a cost-table lookup on the issue path (negligible runtime).
|
||
|
||
## Open review items (decided autonomously; revise on review)
|
||
|
||
1. **Calibration source for the per-op-type values.** Composite ≈ 40 ns is
|
||
a working estimate; primitive issue costs are placeholders. *Recommend:*
|
||
pick values from a documented assumption (instruction-issue + queue
|
||
push) and treat absolute latency as provisional until a real reference
|
||
exists; keep the **ratios** defensible.
|
||
2. **Scheduler plan-gen cost (D3).** *Recommend:* keep 0 initially; expose
|
||
via `overhead_ns` if a workload shows scheduler-bound behaviour.
|
||
3. **DMA program time (A3).** *Recommend:* fold into the issuing op's cost
|
||
first; split to PE_DMA setup only if needed.
|
||
4. **Where the table lives.** *Recommend:* a small central cost module
|
||
keyed by command type, overridable per topology — not scattered node
|
||
attrs — so the values are reviewable in one place.
|
||
5. **Regression rollout.** *Recommend:* land lazy-load (ADR-0062) and this
|
||
ADR's cost model in one golden-regeneration pass to avoid two churns.
|
||
6. **Interaction with the legacy replay path** (`pe_cpu.py:_execute_legacy`)
|
||
— ensure both the greenlet and replay paths read the same cost table so
|
||
results match. Verify.
|
||
|
||
## Test Requirements
|
||
|
||
1. **Composite charges once; primitives charge per-op.** A kernel issuing
|
||
one `tl.composite` over `N_tiles` charges one composite issue cost on
|
||
PE_CPU; the equivalent `N_tiles × ops` primitive kernel charges the
|
||
per-op cost `N_tiles × ops` times. Assert the PE_CPU busy time differs
|
||
accordingly.
|
||
2. **DMA latency unchanged (Q2).** For a fixed transfer, PE_DMA `drain_ns`
|
||
is identical to the pre-ADR value — the issue cost is additive on
|
||
PE_CPU, not folded into DMA (no double-count).
|
||
3. **Saturation is observable.** With non-zero per-op issue cost, a
|
||
many-tile primitive sweep shows GEMM-engine idle (CPU-bound issue)
|
||
whereas the composite sweep keeps it busy — the ADR-0060 §1 lever.
|
||
4. **Determinism:** identical inputs → identical op_log + latency
|
||
(SPEC §0.1).
|
||
5. **Path parity:** greenlet and legacy-replay paths produce identical
|
||
issue-cost accounting for the same kernel.
|