8176cdf287
Decode-SP and prefill-SP are structurally different and cannot share one kernel (principle: move the smaller thing): - Decode: O=[G,d] tiny, KV cache big -> keep KV statically sharded/resident, G heads replicated (M-fold), move only (m,l,O) via the 2-level reduce (S4). - Prefill: O=[S,d] big -> shard heads (1 query head per CUBE, C=G), rotate KV (Ring KV, S5.5), no (m,l,O) reduce; each CUBE writes its own head. Rewrites TL;DR (two kernels), S0 (head map differs by case), S0.5.4 (output head distribution differs -> downstream out-proj impact), S4 (scoped to decode; S4.1 = intra-CUBE KV-split + PE reduce, the only way decode uses P PEs), S5.1 (decode skeleton), S5.5 (head-parallel Ring KV), S9/S10/S11, and adds SB items (two head mappings, output asymmetry, prefill within-CUBE PE, C=G coupling, reconcile with _attention_mesh_mlo_2d). KO mirror deferred until the design stabilizes (adr-proposed is mirror-exempt). Docs only. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>