d0db9ff40e
Make reduce_mlo derive its submesh dimensions (sub_w, sub_h, root_col, root_row, root_cube) from C at call time, with peer-existence guards on every inter-cube send/recv so any C ≥ 1 completes cleanly (previously hardcoded to C=8's 4×2 mesh; non-rectangular C like 7 or 12 hit IpcqDeadlock at the root cube's east receive). Byte-equal at C=8 — existing composite digest tests still pass 8/8. Callsites in the four Case-6 kernels (primitive / primitive-tiled / composite / composite_extended) updated to pass C into reduce_mlo and use root_cube_for(C) for the final tl.store gate. Add the multi-model attention bench: sweeps the composite kernel at S_kv=128K across six GQA models (Gemma-2 27B, LLaMA-3 8B, Qwen 2.5 7B, LLaMA-3 70B, Qwen 2.5 72B, Command R+), with per-model topology C = h_q (cubes per KV group = query heads per KV group). Captures per-op-kind occupancy (matmul GEMM+MATH, comm DMA) alongside latency and PE_CPU dispatch. Three-panel plot writes to bench-output and paper-figures dir. Headline result: same-h_q-different-family models land within 0.2 µs (model-agnostic given fixed h_kv=1 per KV group + d_head=128); latency climbs 263 → 492 µs as C climbs 2 → 12, driven by reduce depth (comm/matmul ratio 0.4 → 0.6 across the sweep). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
169 KiB
2775x750px
169 KiB
2775x750px