ADR-0064 Revision 2 second-pass review fixes: - D7 added 1024-byte cap rationale: explicitly framed as a *safe engineering limit* (intentionally above all known composites at ~322 bytes, finite descriptor capacity placeholder), not a measured HW number. Default is the discipline; topology override for real HW. - D7 segment ordering: strict FIFO is the *ordering source*; rw_handles is dependency metadata, not the ordering primitive. Explicit note for the future RW-aware reorder migration path. - Tests rewritten around the formula (D1) instead of specific numbers: - Test #1 → "formula preservation" (parametrized across composites, asserts dispatch == FIXED + bytes × R, not anchor 43 ns). - Test #2 → robust "opt3 > 2 × opt2" qualitative gate; ≈ 4.0× is informative only, not the gate, so calibration changes don't break. - Test #9 → qualitative (opt2 < opt3 at all R), not a numeric ratio. - Anchor description in D3 stays as *informative* (shows the model produces ≈ 43 ns at defaults), but is no longer a test gate. ADR-0065 second-pass review: - D5 step 5 aligned with D6.6 invariant: explicit conflict check before auto-binding (previously written as if auto-bind always happens). Mirrors D6.6 wording: error if both kernel explicit `a` and prologue primary_out are present. - Test #7 made robust: `opt3 > 2 × opt2` gate (was ≈ 4.0×). Default-calibration model expectation (≈ 4.0×) recorded as informative reference, not the gate. DDD-0065 follow-on robust thresholds: - §1.4 success criteria: ratio condition rewritten as `opt3 > 2 × opt2` with model-expected ≈ 4.0× as informative. - §9 P6 phase gate: `opt3 > 2 × opt2` (was 4.0× ± 15%). - §10 dispatch ratio test: assert `opt3 > 2 × opt2`; record observed ratio for performance tracking but do not gate on it. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
18 KiB
ADR-0064: Structural CPU dispatch cost model (logical_bytes + FIXED + R)
Status
Proposed (Revision 2)
Supporting ADR for ADR-0060 (AHBM GQA Fused Attention) and ADR-0065 (flat-ops composite + first stateful recipe). The hybrid decision there (GEMMs via
tl.composite, softmax merge in the kernel) wins by offloading tiling to PE_SCHEDULER so the CPU issues coarse descriptors and runs ahead, keeping the engines saturated. That win is currently invisible in the simulator because per-op CPU issue cost is zero.Revision 2 replaces the op-type calibration table (the original proposal) with a structural formula derived from each command's
logical_bytes— no per-op-type calibration needed; new op kinds are covered automatically.
Context
What exists today
- Every
tl.*op calls_emit_dispatch_overhead()before emitting its command (tl_context.py:196-212), which emitsPeCpuOverheadCmd(cycles=dispatch_cycles)only if `dispatch_cycles0
. The knob is **uniform** across op kinds and hardcoded to **0** in both live execution paths (pe_cpu.py:101greenlet,:195` replay). - ⇒ issuing a command — constructing the descriptor and pushing it to the scheduler queue — currently costs 0 ns on PE_CPU.
PeCpuOverheadCmdis consumed asyield env.timeout(cmd.cycles)on PE_CPU (kernel_runner.py:131-132).
Why uniform-and-zero is wrong for the hybrid
ADR-0060 §1's argument is that one composite descriptor offloads
N_tiles worth of GEMM tiling, so the CPU issues O(1) coarse commands
instead of O(N_tiles × ops/tile) fine ones. With issue cost = 0, the
model cannot show:
- that the primitive path may fail to saturate the engines when the CPU cannot push fast enough, nor
- that a composite costs more to construct than a single primitive but far less than the many primitives it replaces.
Why per-op-type calibration (Revision 1) was over-shaped
The original proposal had a cost_table[kind] keyed by op kind
(composite, load, dot, math, …). That required:
- a value per kind (calibration cost ≥ |kinds|),
- a new entry every time a new kind appears,
- and yet the ratio it tried to capture — "composite ≫ primitive, but ≪ the primitives it replaces" — is structurally a function of how many fields the command carries, not of the op kind.
A command's byte footprint is the natural proxy: a composite carrying
N OpSpecs has ~N× the bytes of a primitive op with one OpSpec. The
fixed part (queue head update, completion register, MMIO-class
latency) is per-command. The two together compose: FIXED + bytes × R.
What this model actually exposes
The primary signal is command-count reduction through the per-command FIXED cost. The byte term is a secondary refinement that prevents pathologically-large composites from looking free. With realistic on-die queue bandwidth (16 B/cycle, D3), FIXED accounts for ≥85% of the dispatch cost differential between opt3 (≈10 cmds/tile) and opt2 (≈2 cmds/tile) for decode opt2.
This framing also bounds the model: if a single composite were allowed to grow unboundedly large, the byte term alone would not stop the formula from rewarding ever-bigger fused commands beyond what real HW supports — hence the descriptor-size cap in D7.
Decision
D1. Structural dispatch cost formula
Each PE command going to PE_SCHEDULER incurs PE_CPU dispatch cycles:
dispatch_cycles(cmd) = FIXED_PER_CMD + cmd.logical_bytes × R
where:
FIXED_PER_CMD(cycles per command) models queue-tail update, MMIO-class RTT, completion-event registration — fixed per command regardless of size.R(cycles per byte) models the queue-write bandwidth — bytes of the command serialized into the scheduler queue.cmd.logical_bytes(int) is each command's HW-logical byte size, computed from D2 below — not Python'ssys.getsizeof.
PE_CPU emits PeCpuOverheadCmd(cycles=dispatch_cycles(cmd)) before
dispatching, exactly as the existing hook (tl_context.py:_emit_dispatch_ overhead) — only the cycle value changes.
D2. logical_bytes rule
Each PE command dataclass exposes logical_bytes: int (property). The
counting rule (HW-friendly, ignores Python overhead):
| Field kind | Bytes |
|---|---|
| Command framing (cmd-type discriminator + completion id ref) | 4 |
| Opcode (op kind enum) | 1 |
| Enum (scope, etc.) | 1 |
TensorHandle reference (address only — shape/dtype assumed in descriptor table) |
8 |
| Scalar (int/float) | 4 |
| Tuple length marker | 1 |
CompositeCmd recursively sums its ops and rw_handles:
@property
def logical_bytes(self) -> int:
return (
4 # framing
+ 1 + sum(op.logical_bytes for op in self.ops)
+ 1 + 8 * len(self.rw_handles)
)
OpSpec:
@property
def logical_bytes(self) -> int:
return (
1 + 1 # opcode + scope
+ 1 + 8 * len(self.operands) # named operand handles
+ (8 if self.out is not None else 0) # out handle
+ 1 + sum(_extra_bytes(v) for v in self.extra.values())
)
def _extra_bytes(v) -> int:
"""Type-aware byte count for OpSpec.extra values."""
if isinstance(v, bool): return 1
if isinstance(v, (int, float)): return 4
if isinstance(v, (tuple, list)): return 1 + 4 * len(v) # shape, axes, …
if isinstance(v, str): return 1 # opcode-like tag
return 4 # default scalar
# Example types in extra:
# m, k, n int → 4 each
# reduce_axis int → 4
# shape=(64, 64) tuple → 1 + 8 = 9
# factor=1.0 float → 4
(Identical rule for DmaReadCmd, MathCmd, etc. — one property per
dataclass, ~3 lines each.)
Counting rule — per-op summation, no deduplication. Each
TensorHandle reference inside an OpSpec is counted independently,
even when the same handle appears in multiple OpSpecs or in
rw_handles. The rw_handles block is metadata for cross-composite
hazard tracking (ADR-0065 D6.3) and is counted separately from
operand references in ops. There is no deduplication. This matches
HW reality: the descriptor encodes each operand slot as an independent
address field, and the dispatcher tracks rw_handles as a distinct
metadata block. Example: OpSpec(kind="mul_bcast", operands={"src_a": O, "src_b": corr}, out=O) counts handle O twice (once for
src_a, once for out); if O is also in the enclosing
CompositeCmd's rw_handles, it is counted a third time.
D3. Defaults — anchored on a typical composite ≈ 43 ns
Anchor: a single-OpSpec composite for a DMA-staged GEMM path —
one OpSpec(kind="gemm", ...); DMA stages are auto-inserted by
PE_SCHEDULER from operand space (ADR-0065 D4) and do not appear in
logical_bytes (the kernel does not issue them as separate cmds).
Breakdown:
framing 4
ops tuple length 1
GEMM OpSpec 40 (opcode 1 + scope 1 + 1 + 2 handles 16
+ out 8 + 1 + extra m/k/n 12)
rw_handles tuple length 1
rw_handles content 8 (one RW handle for the output)
─────────────────────────────
total ~54 bytes
Target dispatch = ~43 ns. On-die producer→consumer queue at 16 bytes/cycle (typical on-die descriptor queue width).
FIXED_PER_CMD = 40 cycles
R = 0.0625 cycles/byte (= 16 bytes/cycle)
Verification: 40 + 54 × 0.0625 = 43.375 cycles ≈ 43 ns ✓
Cycle→ns conversion uses the PE node's existing clock_freq_ghz attr
(the same one used by PE_MATH _compute_ns). The cost-model knobs are
cycle-domain only — they do not duplicate the clock setting.
D4. Topology config override
Defaults are baked into pe_cpu.py. Topology yaml may override under a
pe_cost_model: section at the PE node attrs (cycle-domain knobs only;
clock comes from the PE's existing clock_freq_ghz):
pe:
attrs:
clock_freq_ghz: 1.0 # existing, used for cycle→ns
pe_cost_model:
fixed_per_cmd_cycles: 40
byte_cycles_recip: 0.0625 # = 16 bytes/cycle
max_composite_logical_bytes: 1024 # D7 — descriptor size cap
Missing keys fall back to defaults. The dispatch formula reads from
node.attrs["pe_cost_model"] at PE_CPU init.
D5. Scope — what does and does not pay
| Path | Pays dispatch cost? |
|---|---|
PE_CPU → PE_SCHEDULER for any PeCommand |
Yes |
PeCpuOverheadCmd itself (already cycles-explicit) |
No (formula bypass) |
| Stages auto-generated by PE_SCHEDULER (DMA_READ/WRITE/FETCH/STORE) | No (PE_SCHEDULER-internal) |
Engine compute latency (DMA drain_ns, GEMM/MATH _compute_ns) |
No change — stays on engines (SPEC §0.1) |
This preserves the "latency on modelled components" invariant — dispatch cost is additional CPU-side time, not folded into engine times.
D6. Configurable values; goldens regenerate
Turning issue cost non-zero changes every bench's latency. Golden latencies are regenerated once when this ADR lands — same posture as ADR-0062 D3 lazy-load. After regeneration, the same calibration is in effect for ADR-0065 opt2 measurement.
D7. Composite size cap (deterministic segmentation)
Each CompositeCmd's logical_bytes is capped at
MAX_COMPOSITE_LOGICAL_BYTES (default 1024 bytes, overridable
per D4). Oversized commands are deterministically segmented by the
emitter (host-side TLContext, ADR-0065 D5) into N consecutive
CompositeCmds, each ≤ cap.
- Each segment carries its own
completion: CompletionHandle. - Each segment incurs its own dispatch cost —
total = sum(FIXED + bytes_i × R) = N × FIXED + total_bytes × R. The FIXED term is paid per segment. - Segments share
rw_handleswhere applicable; strict FIFO (ADR-0065 D6.3) preserves write-after-write ordering automatically. - The segmentation algorithm is deterministic (greedy by op index in emit order); it is part of the host emitter, not PE_SCHEDULER.
Why a cap is needed. Real hardware imposes hard limits — descriptor
queue entry size, scheduler parser buffer, command SRAM, firmware
input. Without a cap, the FIXED + bytes × R model would reward
arbitrarily large fused composites beyond what hardware accepts (e.g.,
fusing 100 primitive ops into one composite, paying one FIXED).
Why 1024 bytes specifically. This is a safe engineering limit,
not a measured HW number — intentionally chosen to be well above all
currently known composites (decode opt2's #2 at ~322 bytes is the
largest in the kernbench codebase) while still representing a finite
descriptor capacity that future recipes must respect. The number is
overridable per topology (D4); when a real HW reference appears, the
value should be recalibrated. The role of this default is to make the
cap exist as a discipline, not to fit a specific HW.
Ordering of segments — driven by strict FIFO, not rw_handles.
Segments are emitted into the PE_CPU → PE_SCHEDULER queue in their
emit order. Strict-FIFO dispatch (ADR-0065 D6.3) is the ordering
source: segments execute in the order they enter the queue. The
rw_handles block on each segment is dependency metadata for the
cross-composite hazard tracker — it does not by itself guarantee
inter-segment ordering. If a future scheduler relaxed FIFO (e.g., to
RW-aware reorder, ADR-0065 A4), the segmenter would need to either
(a) merge segments into a single CompositeCmd, or (b) introduce an
explicit completion-handle dependency chain. For Phase 1 this does
not arise: strict FIFO is in effect.
Decode opt2's #2 composite (10 ops, ~322 bytes) sits comfortably
inside the 1024 cap — no segmentation for the GQA workload.
Alternatives
A1. Keep Revision 1's op-type calibration table
Rejected: calibration cost scales with |kinds|, and the ratio the table tried to capture is structurally a function of cmd size. The structural formula reaches the same qualitative behaviour with two calibratable numbers instead of N.
A2. Byte-only formula (no FIXED term)
Rejected. With FIXED = 0, opt2 (Option Y per ADR-0065) does not win over opt3 — the total bytes dispatched per tile are similar (opt3 ≈ 232, opt2 ≈ 380); the win is entirely in fewer per-cmd fixed costs. A byte-only formula erases the very signal the model needs to expose.
A3. Charge dispatch on PE_SCHEDULER instead of PE_CPU
Rejected: the saturation question is "can the CPU push descriptors fast enough to keep the engines busy?" — that is a PE_CPU issue-bandwidth property. Charging on the scheduler would not model CPU back-pressure.
A4. Model DMA program/setup time as a separate fixed per-descriptor cost
Deferred: initially fold the descriptor-program cost into the issuing op's dispatch cost. Split it out to a PE_DMA fixed setup only if calibration shows it matters.
Consequences
Positive
- Hybrid's CPU-offload / saturation win (ADR-0060 §1) becomes measurable, with a structurally honest model (no calibration table).
- Adding new op kinds (e.g., ADR-0065's
softmax_merge8-step recipe) costs zero — they fit the same formula automatically. - More faithful to hardware (queue-head MMIO RTT + queue-write bandwidth).
Negative
- All bench goldens shift → one-time regeneration (D6); CI golden fixtures update.
- Two calibration knobs (FIXED, R) need values; defaults are anchored on a documented assumption — treat absolute latencies as provisional until a reference exists; keep the ratios defensible.
- Adds a small
logical_bytesproperty to each PE command dataclass.
Open review items
- Calibration source for FIXED and R. Defaults from "typical composite ≈ 43 ns + on-die queue 16 bytes/cycle"; reasonable for an on-die descriptor queue. Revisit when a HW reference appears.
- Scheduler plan-gen cost vs large composites. Stays 0 — D5 keeps
PE_SCHEDULER's plan-generation outside the dispatch formula. The
D7 cap (1024 bytes ≈ 30–35 OpSpecs) bounds the worst case, but a
composite near the cap still costs the scheduler the same as a 1-op
composite under the current zero-cost model. If a stress test
(large-composite microbench) shows scheduler-bound behaviour,
expose
overhead_nsper-op-count. - Where the override lives.
pe_cost_model:block under PE node attrs in topology yaml — keeps all knobs in one place, reviewable. Clock comes from the PE's existingclock_freq_ghzattr, not duplicated here. - Path parity. Both greenlet (
_execute_legacyandkernel_runner) and replay paths must read the same cost model. Verify. - Sensitivity of conclusions to R. opt2 < opt3 must hold across a
reasonable range of
R(queue-bandwidth assumption). Sensitivity sweep is part of Test Requirements (#9).
Test Requirements
Tests are written against the formula (D1), not against specific numeric anchors, so they remain valid when calibration changes or when OpSpec/CompositeCmd fields are added.
- Formula preservation. For any
CompositeCmdc, PE_CPU's recorded dispatch overhead equalsFIXED_PER_CMD + c.logical_bytes × R(within ±1 cycle for floor/round-off). Parametrized over several composites: a 1-OpSpec GEMM composite, a 5-op MATH chain, and a 10-op recipe composite. Default-calibration numbers (anchor ≈43 ns for the 1-OpSpec composite) are informative reference, not the test gate — the test gate is the formula equality. - Qualitative ratio (robust). opt3 per-tile PE_CPU dispatch
strictly exceeds opt2 per-tile dispatch by at least a 2× margin —
opt3 > 2 × opt2. The default-calibration model predicts ≈4×; the gate is the loose 2× bound so the test does not break when calibration is moved (e.g., when a HW reference replaces the default). Informative numbers — see ADR-0065 §verification and DDD-0065 §11 for the model expectation. - Override path. Topology yaml
pe_cost_model:block changes the per-PE dispatch cost; default is recovered when block is missing. The formula identity from #1 must hold with the override values. PeCpuOverheadCmdbypass. Manualtl.cycles(n)issues exactlyncycles, notn + dispatch_cycles(...).- No double-count. PE_DMA
drain_ns, PE_GEMM/MATH_compute_nsidentical to pre-ADR values. - Determinism. Identical inputs → identical op_log + latency (SPEC §0.1).
- Path parity. Greenlet and replay paths produce identical dispatch-cycle accounting for the same kernel.
- Composite size cap (D7). A recipe that would emit `logical_bytes
MAX_COMPOSITE_LOGICAL_BYTES
is segmented into N consecutiveCompositeCmds;sum(segment.logical_bytes) == original logical_bytes; total dispatch = sum of segment dispatches (FIXED paid per segment); inter-segment ordering preserved by strict FIFO (not byrw_handles` alone). - Sensitivity (qualitative). At
R ∈ {0.25, 0.0625, 0.03125}cycles/byte,opt3 > opt2at all three points. Direction (ratio monotonically increases as R decreases) is also asserted, but absolute ratio values are not required.
Migration
ADR-0064 Revision 2 lands as a single PR with:
logical_bytesproperty on eachPeCommanddataclass (type-aware extra-field counting per D2)- formula application in
pe_cpu.pydispatch path pe_cost_model:override read at PE_CPU init (cycle-domain knobs +max_composite_logical_bytes)- composite size cap (D7) — TLContext-side segmentation logic with
MAX_COMPOSITE_LOGICAL_BYTESdefault 1024; not needed for any existing bench (largest current composite is well under 200 bytes), but the mechanism lands so it is in place when ADR-0065's 10-op decode-opt2 composite (~310 bytes) and larger recipes appear. - one-time goldens regeneration
After this lands, ADR-0065 builds on top with no further goldens churn
in existing benches (ADR-0065 is a meaning-preserving refactor of
CompositeCmd for the existing path; only opt2 is a new bench).