- ADR-0060: GEMMs (Q.Kt, P.V) via existing tl.composite (scheduler-managed tiling + K/V DMA streaming); softmax merge + IPCQ tree reduction stay in kernel. Front TL;DR pseudocode of the final composite kernel; new section B lists open design items (DDD sync, K pre-transpose, dma_read lever, kernel-vs-scheduler tiling, ring path). - ADR-0062: redefined from a new load_async op to global lazy tl.load (non-blocking + auto-wait on first use; API unchanged; goldens regenerate). - ADR-0064 (new): per-op-type CPU issue cost model (composite ~40ns >> primitive) so the hybrid's CPU-saturation win becomes measurable (currently dispatch_cycles=0 hides it). Cost-model impl deferred. - KO mirrors for ADR-0060/0062/0064 (-ko suffix, adr-proposed). Rationale: non-blocking CompositeCmd offloads tiling to PE_SCHEDULER, decoupling CPU issue-rate from execution so the CPU can saturate the engines; the prior 'composite = no latency benefit' claim was an artifact of dispatch_cycles=0. Docs only; no production code changed. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
7.8 KiB
ADR-0062: Lazy tl.load — non-blocking HBM load with auto-wait on first use
Status
Proposed
Supporting ADR for ADR-0060 (AHBM GQA Fused Attention). Decode and long-context attention are KV-load-bound; load/compute overlap is a dominant lever. This ADR makes
tl.loaditself lazy (non-blocking, with the wait automatically inserted at first use) rather than adding a separateload_asyncop. Today the only async primitive is for IPCQ comms, not for HBM loads.
Context
What overlap requires
FlashAttention streams operands: while the GEMM/MATH engine works on the
current tile, the DMA engine should already be pulling the next operand.
With that overlap a bandwidth-bound kernel runs at roughly
max(compute, dma) per tile instead of compute + dma.
In ADR-0060's hybrid design the two GEMMs are issued as
tl.composite(op="gemm") whose K/V operands are tl.ref (HBM-resident,
streamed per tile by PE_SCHEDULER), so the per-tile K/V prefetch is
handled by the composite scheduler. Lazy tl.load covers the remaining
explicit loads (the Q group, and any non-composite kernel) so those also
overlap the compute that follows instead of stalling the greenlet.
What exists
tl.load(ptr, shape, dtype)is blocking: it emitsDmaReadCmdand the greenlet kernel suspends until PE_DMA signals completion (tl_context.py:177-203; greenlet drivekernel_runner.py:146-153). No twotl.loads can be in flight from one kernel.- An async pattern does exist, but only for IPCQ:
tl.recv_async(dir, ...) -> RecvFuture+ a deferred wait (kernel_runner.py:248-285). It proves the machinery — a non-blocking command that returns a future, resolved later by a wait check (if not future.event.triggered: yield future.event) — works in the greenlet model. - The DMA engine models a read channel as a SimPy resource (capacity 1,
pe_dma.py:45) separate from the write channel, so in-flight reads are representable; they serialise on the single read channel while overlapping compute on PE_GEMM/PE_MATH. Only the kernel-facing API currently serialises load-vs-compute.
tl.load cannot today overlap with the compute that follows it.
Decision
Make tl.load lazy: it issues the DmaReadCmd and returns a handle
immediately (non-blocking); the runtime auto-inserts the wait at the
first point the loaded data is actually consumed. The kernel-facing API
is unchanged — authors keep writing tl.load — so this is a semantics
change, not a new op. It generalises the existing recv_async/wait
machinery (1) to the HBM-load path and (2) from an explicit tl.wait
call to an implicit, dependency-driven wait at first use.
D1. tl surface — unchanged
tl.load(ptr, shape, dtype) keeps its signature and TensorHandle
return. What changes is when it blocks: never at issue, only implicitly
when its result is first read by a consuming op.
D2. Mechanism — non-blocking issue + auto-wait on use
tl.loadposts theDmaReadCmdto PE_DMA without yielding its completion event, and records the pending event on the returned handle (therecv_asyncpattern, applied to loads).- When a consuming op (
tl.dot, a MATH op,tl.store, atl.compositeoperand, …) is dispatched, the runtime checks each input handle for a pending load event and yields it first if not yet triggered — i.e. the wait is inserted automatically at the latest correct point (first use). - The op_log entry is unchanged (
memory/dma_read): asynchrony is a scheduling property, not a new op kind, sodma_read_countand existing op_log consumers keep working.
D3. Scope — global
tl.load is lazy everywhere, not behind an opt-in flag. This is the
faithful model: blocking on every load is a property of a naive kernel;
an efficient kernel (and a real compiler) hoists the load and waits only
at use. Consequence: existing kernels that have independent work between a
tl.load and its first use see lower (faster) latency — a
correctness improvement of the model, not a behaviour regression. Golden
latencies that change must be regenerated; kernels that load then
immediately use see no change (the auto-wait fires at once, identical to
blocking).
D4. Latency / overlap semantics
tl.loadcharges the issue (descriptor push) only; the kernel proceeds. (Per-op issue cost isdispatch_cycles, currently 0; an op-type-differentiated issue cost is tracked separately — ADR-0060 §1, §9.)- The DMA transfer occupies the read channel (capacity 1) for its modelled duration in parallel with whatever compute the kernel issues next; multiple in-flight loads serialise on the channel but overlap compute.
- The auto-inserted wait blocks only if the transfer has not finished.
- Determinism is preserved: the wait point is fixed by program order (first use), and completion is a scheduled event on the modelled DMA channel (SPEC §0.1, R8). No latency subtraction — the overlap is real modelled concurrency.
Modelling assumption. Auto-wait-at-first-use models a well- scheduled kernel (the compiler places the wait at the latest correct point). Real compilers approximate this; some loads cannot be hoisted (register pressure, aliasing). For a performance simulator (SPEC §0) modelling the well-scheduled case is the intended behaviour.
Alternatives
A1. Separate tl.load_async op (earlier proposal)
Add an explicit load_async/wait pair the kernel calls by hand (the
double-buffer dance). Rejected: it enlarges the kernel-facing API and
pushes buffer-lifetime bookkeeping onto every kernel author, when lazy
tl.load gives the same overlap with no API-surface change and a
compiler-style auto-wait.
A2. Per-kernel opt-in lazy load
Keep tl.load blocking by default; make new kernels opt in. Rejected:
splits tl.load semantics into two variants and hides the model
improvement from existing benches; global lazy is cleaner and the
golden-regeneration cost is one-time (D3).
A3. Rely on composite streaming only (no lazy load)
ADR-0060's composites stream their tl.ref K/V operands, so the GEMM
operand DMA already overlaps. But explicit loads (the Q group, and any
non-composite kernel) still stall without lazy tl.load. Insufficient on
its own.
Consequences
Positive
- Load/compute overlap with zero kernel-facing API change; authors
keep writing
tl.load. - Symmetric with
recv_async/wait — low conceptual surface area. - General: any bandwidth-bound kernel prefetches automatically.
Negative
- Existing golden latencies shift (faster) for kernels with independent work between load and use → one-time regeneration (D3).
- Auto-wait requires the runtime to track per-handle pending events and check them at consuming-op dispatch (data-dependency tracking).
- Interacts with scratch lifetime (ADR-0063): an in-flight load's target buffer must not be recycled before its auto-wait fires. The recycling scope must exclude live (un-waited) load buffers.
Test Requirements
- Overlap is real: a kernel that issues
tl.loadthen an independent GEMM of comparable duration completes in ≈max(load, gemm), notload+gemm(end-to-end latency strictly below the serial sum). - Auto-wait correctness: the loaded
TensorHandle, when first consumed, carries the same bytes as today's blockingtl.loadfor the same address (Phase 2). - Two in flight: two
tl.loads to distinct addresses, consumed later, both resolve to correct, independent tensors; their DMAs serialise on the read channel. - op_log compatibility: each
tl.loadstill logs exactly onememory/dma_read;dma_read_countunchanged vs the blocking version. - No-overlap no-change: a kernel that loads then immediately uses has identical latency to the blocking model (auto-wait fires at once).