9b4aee83da
- Slider caps: B_max = 256, S_kv_max = 1M (matches practical decode
operating ranges).
- Batch axis on every plot capped at 256 to match the slider.
- Removed the two redundant plots (old Plot 1 and old Plot 4). Both
were re-drawing the same weight/compute/KV decomposition as the new
headline 'Cost per token' plot at the top — the only distinguishing
bits (extra 'Memory total' curve, B* marker) have been folded into
the headline plot already, so keeping them was pure clutter.
- Every remaining plot now reacts to the S_kv / B knobs:
* Plot A (step latency): S_kv drives curves, B marker drawn.
* Plot B (cost/token): same.
* Plot 2 (no-knee overlay): S_kv drives the L* ratios shown,
B marker drawn.
* Plot 3 (knee vs context): S_kv AND B markers drawn on both axes.
Net: 4 plots (headline pair + no-knee + knee-vs-context), all live-
interactive against both sliders.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>