f9d0077472
Per-instance cache keyed by (id(adj), start, goal). All 4 adj dicts (_adj, _adj_all, _adj_local, _adj_mcpu_dma) are built in __init__ and never mutated (topology is static per ADR-0006 / SPEC §0.1), so id(adj) is stable for the router's lifetime. Cache is populated on the successful-return paths; RoutingError paths intentionally re-run each call (rare, keeps error semantics unchanged). Motivation: cProfile of Case 4 decode at S_kv=8K showed 1,142 _run_dijkstra_with_dist calls consuming ~1.95s tottime. Paths depend only on (adj, src, dst) so ~99% of those calls are recomputing the same result. Verified: - tests/attention/test_milestone_gqa_decode_long_ctx_4cases.py: 18/18 pass - 128K decode wall: 237.28s -> 233.71s (-1.5%) - Modeled kernel latency: 461.13us (byte-identical before/after) - op_log_len: 3057 (unchanged) Small win but no risk: memoization returns byte-identical results and paths cannot change during a sim run. Larger event-count reductions require touching the per-hop hot path (zero-latency chain collapse etc.). Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>