d282144339
Land the new GQA fused-attention kernels (ADR-0060) for prefill/decode
across long and short context, the TL discipline primitives they depend
on (ADR-0062 lazy load, ADR-0063 scratch_scope + copy_to), and the
per-op-type CPU issue cost model (ADR-0064). Remove the pre-ADR-0060
mesh-attention baseline now that the unified kernels supersede it.
ADR-0060 (long context)
- _gqa_decode.py: M-fold + 2-level chain reduce-to-root (Level-2
intra-CUBE row-then-col + Level-1 inter-CUBE) — root-only output.
- _gqa_prefill.py: head-parallel + Ring KV rotation around C CUBEs,
online-softmax merge per ring step, per-CUBE distributed output.
- Each merge stage wraps in scratch_scope() and persists running
(m, l, O) via copy_to() to lift the 1 MiB scratch ceiling.
ADR-0060 §B.split.2 (short context, kv_per_cube in {1,2,4,8})
- _gqa_decode_short.py / _gqa_prefill_short.py: no cube-SP; each CUBE
owns whole KV heads; PE-parallel heads with intra-group chain
reduce. Prefill has no Ring KV (each head fully resident).
ADR-0062 (lazy tl.load): future-bearing TensorHandle, auto-wait at
first consuming op (dot/MATH/store/send/copy_to/composite).
ADR-0063 (tl.scratch_scope + tl.copy_to): scoped per-tile arena with
copy_to writeback primitive for persistent running state.
ADR-0064 (CPU issue cost model)
- common/cpu_issue_cost.py: per-op-type table (composite=40 ns,
primitives=5 ns); ratios are load-bearing per D1.
- TLContext: issue_cost_table param; _emit_dispatch_overhead(kind)
consults table with dispatch_cycles fallback (ADR-0046 §D6
back-compat).
- Live PE_CPU paths (greenlet + legacy) construct TLContext with
DEFAULT_CPU_ISSUE_COST so saturation lever (ADR-0060 §1) is
measurable end-to-end.
P7 headline bench: milestone-gqa-headline writes per-panel
op_log_summary to 1H_milestone_output/gqa_headline/sweep.json. No
figure renderers yet (deferred).
Removals (pre-ADR-0060 baseline now superseded):
- benches: _attention_mesh_kv.py, _attention_mesh_mlo.py,
_attention_mesh_mlo_2d.py, milestone_gqa_llama70b.py
- tests: test_attention_*, test_mesh_*, test_milestone_gqa_llama70b
- topology: llama70b_4sip.yaml (only consumer was the deleted diag)
- artifacts: 1H_milestone_output/gqa/ (sweep.json + 5 PNGs)
- tests/gqa/ plot helper + test (broken on Windows Tcl/Tkinter)
- ADR-0060/0061 references to deleted file paths cleaned up
(EN + KO kept in sync).
Tests: 124/124 focused regression green (attention + Phase E + TL
discipline + triton_emu + pe_components). Full regression: 764 pass,
2 pre-existing test_bench_registry failures (stale EXPECTED_NAMES
across multiple benches, not introduced here).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
69 lines
1.4 KiB
JSON
69 lines
1.4 KiB
JSON
{
|
|
"version": 1,
|
|
"panels": [
|
|
"single_user_prefill_gqa",
|
|
"multi_user_prefill_gqa",
|
|
"single_user_decode_gqa",
|
|
"multi_user_decode_gqa"
|
|
],
|
|
"config": {
|
|
"T_q_prefill": 4,
|
|
"T_q_decode": 1,
|
|
"S_kv_prefill": 16,
|
|
"h_q_decode": 8,
|
|
"h_kv_decode": 1,
|
|
"d_head": 64
|
|
},
|
|
"rows": [
|
|
{
|
|
"panel": "single_user_prefill_gqa",
|
|
"kind": "prefill",
|
|
"C": 1,
|
|
"S_kv": 16,
|
|
"op_log_summary": {
|
|
"gemm_count": 2,
|
|
"ipcq_copy_count": 0,
|
|
"dma_read_count": 3,
|
|
"dma_write_count": 1
|
|
}
|
|
},
|
|
{
|
|
"panel": "multi_user_prefill_gqa",
|
|
"kind": "prefill",
|
|
"C": 4,
|
|
"S_kv": 16,
|
|
"op_log_summary": {
|
|
"gemm_count": 32,
|
|
"ipcq_copy_count": 24,
|
|
"dma_read_count": 12,
|
|
"dma_write_count": 4
|
|
}
|
|
},
|
|
{
|
|
"panel": "single_user_decode_gqa",
|
|
"kind": "decode",
|
|
"C": 1,
|
|
"P": 8,
|
|
"S_kv": 64,
|
|
"op_log_summary": {
|
|
"gemm_count": 16,
|
|
"ipcq_copy_count": 21,
|
|
"dma_read_count": 24,
|
|
"dma_write_count": 1
|
|
}
|
|
},
|
|
{
|
|
"panel": "multi_user_decode_gqa",
|
|
"kind": "decode",
|
|
"C": 4,
|
|
"P": 8,
|
|
"S_kv": 128,
|
|
"op_log_summary": {
|
|
"gemm_count": 64,
|
|
"ipcq_copy_count": 93,
|
|
"dma_read_count": 96,
|
|
"dma_write_count": 1
|
|
}
|
|
}
|
|
]
|
|
} |