ae131aa2af
Both 1-kv (2 users x 64 PE) and 8-kv (16 users x 8 PE) fill all 128 PEs of the SIP; the throughput gap is per-PE HBM efficiency. Splitting a head over 8 PEs leaves each a single tile, so pipeline-fill and the 8-PE softmax reduce dominate -> 46% HBM util; one head per PE streams contiguously with no reduce -> 76%. Throughput ratio (1.6x) tracks the utilization ratio (1.7x): decode is bandwidth-bound, so throughput follows total HBM efficiency, not PE count. 1-kv spends hardware on latency (8 PE/head for a 4.8x speedup = 60% strong-scaling efficiency). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2.9 MiB
2.9 MiB