b7c4c680aa
New section at the bottom of the Chip Roofline tab implementing the three-axis GPU count decision: N_GPUs = max(capacity_floor, KV_headroom, throughput_SLO) × N_replicas Interactive inputs: concurrent users, avg context/user, TPOT SLO. Four KPI cards report each axis + the recommended total, with a 'binds:' delta chip showing which axis is the bottleneck. - Axis A (capacity): PEs to hold weights alone - Axis B (KV): PEs to hold weights + all KV of one replica's users - Axis C (throughput): replicas needed at B_at_SLO users each - Total: max(A,B) × N_replicas Contextual caption explains what to do about the binding axis. If SLO is infeasible (even B=1 exceeds SLO), an error block explains the options (loosen SLO, shorten context, or scale up FLOPs/BW knobs). Plus three collapsed expanders explaining the hybrid deployment pattern hyperscalers use: - Layer 1: one elastic pool + PagedAttention + continuous batching - Layer 2: length-tier routing (standard vs long-context) - Layer 3: disaggregated prefill/decode (DistServe, Splitwise) Pure additions in chip_roofline.py: - max_batch_within_slo(machine, model, s_kv, slo_s) — analytical inverse of step_latency to find the largest per-replica B under SLO - size_deployment(machine, model, n_users, avg_ctx, slo_s) — returns GpuSizingResult with all three axes + binding info 6 new tests cover: capacity scales with model size, KV grows with users/context, binding axis flips with workload shape, tight SLO shrinks max batch, weight-time-exceeds-SLO returns 0, total == per-replica × replicas. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>