a8c50238c6
ADR-0064 Revision 2: replace per-op-type calibration table with a structural formula `FIXED_PER_CMD + cmd.logical_bytes × R`, defaults anchored at typical composite ≈ 45 ns. Topology yaml override under `pe_cost_model:` block; logical_bytes property per PE command. ADR-0065: implement ADR-0060 §5.6 / §8 item 4 carve-out as a flat-ops CompositeCmd (no head/epilogue structural fields — position + scope drives placement) + first stateful recipe `softmax_merge` (MATH-only 8-step). RECIPE_DESCRIPTORS lives in TLContext-adjacent module only; PE_SCHEDULER stays recipe-free and auto-inserts DMAs from operand `space`. Strict-FIFO RW hazard tracker; ≤1 GEMM per composite invariant. User-facing `tl.composite(prologue=[...], op=, epilogue=[...])` API preserved; existing benches unchanged (meaning-preserving refactor). DDD-0065: implementation-ready phased plan (P0=ADR-0064 Rev2 first, P1-P6 for ADR-0065), file plan, recipe engine sequence, scheduler plan-gen algorithm, RW tracker design, test matrix. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
259 lines
10 KiB
Markdown
259 lines
10 KiB
Markdown
# ADR-0064: Structural CPU dispatch cost model (`logical_bytes` + FIXED + R)
|
||
|
||
## Status
|
||
|
||
Proposed (Revision 2)
|
||
|
||
> Supporting ADR for **ADR-0060** (AHBM GQA Fused Attention) and **ADR-0065**
|
||
> (flat-ops composite + first stateful recipe). The hybrid decision there
|
||
> (GEMMs via `tl.composite`, softmax merge in the kernel) wins by
|
||
> **offloading tiling to PE_SCHEDULER so the CPU issues coarse descriptors
|
||
> and runs ahead, keeping the engines saturated**. That win is currently
|
||
> **invisible in the simulator** because per-op CPU issue cost is zero.
|
||
>
|
||
> Revision 2 replaces the **op-type calibration table** (the original
|
||
> proposal) with a **structural formula** derived from each command's
|
||
> `logical_bytes` — no per-op-type calibration needed; new op kinds are
|
||
> covered automatically.
|
||
|
||
## Context
|
||
|
||
### What exists today
|
||
|
||
- Every `tl.*` op calls `_emit_dispatch_overhead()` before emitting its
|
||
command (`tl_context.py:196-212`), which emits
|
||
`PeCpuOverheadCmd(cycles=dispatch_cycles)` **only if** `dispatch_cycles
|
||
> 0`. The knob is **uniform** across op kinds and hardcoded to **0**
|
||
in both live execution paths (`pe_cpu.py:101` greenlet, `:195` replay).
|
||
- ⇒ issuing a command — *constructing the descriptor and pushing it to
|
||
the scheduler queue* — currently costs **0 ns** on PE_CPU.
|
||
- `PeCpuOverheadCmd` is consumed as `yield env.timeout(cmd.cycles)` on
|
||
PE_CPU (`kernel_runner.py:131-132`).
|
||
|
||
### Why uniform-and-zero is wrong for the hybrid
|
||
|
||
ADR-0060 §1's argument is that **one** composite descriptor offloads
|
||
`N_tiles` worth of GEMM tiling, so the CPU issues `O(1)` coarse commands
|
||
instead of `O(N_tiles × ops/tile)` fine ones. With issue cost = 0, the
|
||
model cannot show:
|
||
- that the primitive path may **fail to saturate** the engines when the
|
||
CPU cannot push fast enough, nor
|
||
- that a composite **costs more to construct** than a single primitive
|
||
but far less than the many primitives it replaces.
|
||
|
||
### Why per-op-type calibration (Revision 1) was over-shaped
|
||
|
||
The original proposal had a `cost_table[kind]` keyed by op kind
|
||
(`composite`, `load`, `dot`, `math`, …). That required:
|
||
- a value per kind (calibration cost ≥ |kinds|),
|
||
- a new entry every time a new kind appears,
|
||
- and yet the *ratio* it tried to capture — "composite ≫ primitive,
|
||
but ≪ the primitives it replaces" — is structurally a function of
|
||
**how many fields the command carries**, not of the op kind.
|
||
|
||
A command's *byte footprint* is the natural proxy: a composite carrying
|
||
N OpSpecs has ~N× the bytes of a primitive op with one OpSpec. The
|
||
*fixed* part (queue head update, completion register, MMIO-class
|
||
latency) is per-command. The two together compose: `FIXED + bytes × R`.
|
||
|
||
## Decision
|
||
|
||
### D1. Structural dispatch cost formula
|
||
|
||
Each PE command going to PE_SCHEDULER incurs PE_CPU dispatch cycles:
|
||
|
||
```
|
||
dispatch_cycles(cmd) = FIXED_PER_CMD + cmd.logical_bytes × R
|
||
```
|
||
|
||
where:
|
||
- `FIXED_PER_CMD` (cycles per command) models queue-tail update, MMIO-class
|
||
RTT, completion-event registration — fixed per command regardless of size.
|
||
- `R` (cycles per byte) models the queue-write bandwidth — bytes of the
|
||
command serialized into the scheduler queue.
|
||
- `cmd.logical_bytes` (int) is each command's *HW-logical* byte size,
|
||
computed from D2 below — not Python's `sys.getsizeof`.
|
||
|
||
PE_CPU emits `PeCpuOverheadCmd(cycles=dispatch_cycles(cmd))` before
|
||
dispatching, exactly as the existing hook (`tl_context.py:_emit_dispatch_
|
||
overhead`) — only the cycle value changes.
|
||
|
||
### D2. `logical_bytes` rule
|
||
|
||
Each PE command dataclass exposes `logical_bytes: int` (property). The
|
||
counting rule (HW-friendly, ignores Python overhead):
|
||
|
||
| Field kind | Bytes |
|
||
|---|---|
|
||
| Command framing (cmd-type discriminator + completion id ref) | 4 |
|
||
| Opcode (op kind enum) | 1 |
|
||
| Enum (scope, etc.) | 1 |
|
||
| `TensorHandle` reference (address only — shape/dtype assumed in descriptor table) | 8 |
|
||
| Scalar (int/float) | 4 |
|
||
| Tuple length marker | 1 |
|
||
|
||
`CompositeCmd` recursively sums its `ops` and `rw_handles`:
|
||
|
||
```python
|
||
@property
|
||
def logical_bytes(self) -> int:
|
||
return (
|
||
4 # framing
|
||
+ 1 + sum(op.logical_bytes for op in self.ops)
|
||
+ 1 + 8 * len(self.rw_handles)
|
||
)
|
||
```
|
||
|
||
`OpSpec`:
|
||
|
||
```python
|
||
@property
|
||
def logical_bytes(self) -> int:
|
||
return (
|
||
1 + 1 # opcode + scope
|
||
+ 1 + 8 * len(self.operands) # named operand handles
|
||
+ (8 if self.out is not None else 0) # out handle
|
||
+ 1 + sum(4 for _ in self.extra.values()) # extra scalars
|
||
)
|
||
```
|
||
|
||
(Identical rule for `DmaReadCmd`, `MathCmd`, etc. — one property per
|
||
dataclass, ~3 lines each.)
|
||
|
||
### D3. Defaults — anchored on a typical composite ≈ 45 ns
|
||
|
||
Anchor: a typical 1-op DMA→GEMM→DMA composite has `logical_bytes ≈ 52`
|
||
(framing 4 + GEMM OpSpec 39 + rw_handles 9). Target dispatch = 45 ns.
|
||
On-die producer→consumer queue assumed at 4 bytes/cycle.
|
||
|
||
```
|
||
clock = 1 GHz # 1 cycle = 1 ns
|
||
FIXED_PER_CMD = 32 cycles
|
||
R = 0.25 cycles/byte
|
||
```
|
||
|
||
Verification: `32 + 52 × 0.25 = 45 cycles ≈ 45 ns` ✓
|
||
|
||
### D4. Topology config override
|
||
|
||
Defaults are baked into `pe_cpu.py`. Topology yaml may override under a
|
||
`pe_cost_model:` section at the PE node attrs:
|
||
|
||
```yaml
|
||
pe:
|
||
attrs:
|
||
pe_cost_model:
|
||
fixed_per_cmd_cycles: 32
|
||
byte_cycles_recip: 0.25
|
||
clock_freq_ghz: 1.0
|
||
```
|
||
|
||
Missing keys fall back to defaults. The dispatch formula reads from
|
||
`node.attrs["pe_cost_model"]` at PE_CPU init.
|
||
|
||
### D5. Scope — what does and does not pay
|
||
|
||
| Path | Pays dispatch cost? |
|
||
|---|---|
|
||
| PE_CPU → PE_SCHEDULER for any `PeCommand` | **Yes** |
|
||
| `PeCpuOverheadCmd` itself (already cycles-explicit) | **No** (formula bypass) |
|
||
| Stages auto-generated by PE_SCHEDULER (DMA_READ/WRITE/FETCH/STORE) | **No** (PE_SCHEDULER-internal) |
|
||
| Engine compute latency (DMA `drain_ns`, GEMM/MATH `_compute_ns`) | **No change** — stays on engines (SPEC §0.1) |
|
||
|
||
This preserves the "latency on modelled components" invariant — dispatch
|
||
cost is *additional* CPU-side time, not folded into engine times.
|
||
|
||
### D6. Configurable values; goldens regenerate
|
||
|
||
Turning issue cost non-zero changes **every** bench's latency. Golden
|
||
latencies are **regenerated once** when this ADR lands — same posture as
|
||
ADR-0062 D3 lazy-load. After regeneration, the same calibration is in
|
||
effect for ADR-0065 opt2 measurement.
|
||
|
||
## Alternatives
|
||
|
||
### A1. Keep Revision 1's op-type calibration table
|
||
|
||
Rejected: calibration cost scales with |kinds|, and the *ratio* the
|
||
table tried to capture is structurally a function of cmd size. The
|
||
structural formula reaches the same qualitative behaviour with two
|
||
calibratable numbers instead of N.
|
||
|
||
### A2. Byte-only formula (no FIXED term)
|
||
|
||
Rejected. With FIXED = 0, opt2 (Option Y per ADR-0065) does **not** win
|
||
over opt3 — the total *bytes* dispatched per tile are similar (opt3 ≈ 232,
|
||
opt2 ≈ 380); the win is entirely in *fewer per-cmd fixed costs*. A
|
||
byte-only formula erases the very signal the model needs to expose.
|
||
|
||
### A3. Charge dispatch on PE_SCHEDULER instead of PE_CPU
|
||
|
||
Rejected: the saturation question is *"can the CPU push descriptors fast
|
||
enough to keep the engines busy?"* — that is a **PE_CPU** issue-bandwidth
|
||
property. Charging on the scheduler would not model CPU back-pressure.
|
||
|
||
### A4. Model DMA program/setup time as a separate fixed per-descriptor cost
|
||
|
||
Deferred: initially fold the descriptor-program cost into the **issuing
|
||
op's** dispatch cost. Split it out to a PE_DMA fixed setup only if
|
||
calibration shows it matters.
|
||
|
||
## Consequences
|
||
|
||
### Positive
|
||
- Hybrid's CPU-offload / saturation win (ADR-0060 §1) becomes
|
||
**measurable**, with a structurally honest model (no calibration table).
|
||
- Adding new op kinds (e.g., ADR-0065's `softmax_merge` 8-step recipe)
|
||
costs zero — they fit the same formula automatically.
|
||
- More faithful to hardware (queue-head MMIO RTT + queue-write bandwidth).
|
||
|
||
### Negative
|
||
- **All** bench goldens shift → one-time regeneration (D6); CI golden
|
||
fixtures update.
|
||
- Two calibration knobs (FIXED, R) need values; defaults are anchored on
|
||
a documented assumption — treat absolute latencies as provisional until
|
||
a reference exists; keep the **ratios** defensible.
|
||
- Adds a small `logical_bytes` property to each PE command dataclass.
|
||
|
||
## Open review items
|
||
|
||
1. **Calibration source for FIXED and R.** Defaults from "typical
|
||
composite = 45 ns + on-die queue 4 bytes/cycle"; reasonable for an
|
||
on-die producer→consumer queue. Revisit when a HW reference appears.
|
||
2. **Scheduler plan-gen cost.** Stays 0 — D5 keeps PE_SCHEDULER's
|
||
plan-generation outside the dispatch formula. Expose via existing
|
||
`overhead_ns` if a workload shows scheduler-bound behaviour.
|
||
3. **Where the override lives.** `pe_cost_model:` block under PE node
|
||
attrs in topology yaml — keeps all knobs in one place, reviewable.
|
||
4. **Path parity.** Both greenlet (`_execute_legacy` and `kernel_runner`)
|
||
and replay paths must read the same cost model. Verify.
|
||
|
||
## Test Requirements
|
||
|
||
1. **Anchor preservation.** A typical DMA→GEMM→DMA composite (1 op,
|
||
`logical_bytes ≈ 52`) dispatches in 45 ns at the default values.
|
||
2. **Structural ratio.** opt3 vs opt2 dispatch (per ADR-0065 §verification):
|
||
`opt3 / opt2 ≈ 2.4×` at default calibration.
|
||
3. **Override path.** Topology yaml `pe_cost_model:` block changes the
|
||
per-PE dispatch cost; default is recovered when block is missing.
|
||
4. **`PeCpuOverheadCmd` bypass.** Manual `tl.cycles(n)` issues exactly `n`
|
||
cycles, not `n + dispatch_cycles(...)`.
|
||
5. **No double-count.** PE_DMA `drain_ns`, PE_GEMM/MATH `_compute_ns`
|
||
identical to pre-ADR values.
|
||
6. **Determinism.** Identical inputs → identical op_log + latency
|
||
(SPEC §0.1).
|
||
7. **Path parity.** Greenlet and replay paths produce identical
|
||
dispatch-cycle accounting for the same kernel.
|
||
|
||
## Migration
|
||
|
||
ADR-0064 Revision 2 lands as a single PR with:
|
||
- `logical_bytes` property on each `PeCommand` dataclass
|
||
- formula application in `pe_cpu.py` dispatch path
|
||
- `pe_cost_model:` override read at PE_CPU init
|
||
- one-time goldens regeneration
|
||
|
||
After this lands, ADR-0065 builds on top with no further goldens churn
|
||
in existing benches (ADR-0065 is a meaning-preserving refactor of
|
||
`CompositeCmd` for the existing path; only opt2 is a new bench).
|