sccl: drive allreduce tests via torch.distributed; reorganize into tests/sccl/

Convert the multidevice allreduce correctness + latency/buffer-kind sweeps to run through the real PyTorch-distributed path (init_process_group(backend="ahbm") -> mp.spawn -> dist.all_reduce) instead of direct ctx.launch, and reorganize the CCL/allreduce tests into a tests/sccl/ package split one test per file. Production change (required for the distributed path on non-square SIP grids): - AhbmCCLBackend now reads explicit system.sips.w/h from the spec, with a square-only sqrt fallback that raises on ambiguity, instead of silently guessing round(sqrt(count)). This fixes the 2x3 / 3x2 torus + mesh cases, which previously resolved to a wrong 2x2 grid. Mirrors the test helper's _sip_topo_dims precedence (explicit w/h > square fallback > raise). Test reorganization (tests/sccl/): - _allreduce_helpers.py: shared plumbing (distributed driver, config writers, direct-launch run_allreduce parity reference, sweep/buffer-kind constants, plot aggregators, topology-diagram + FSIM-comparison emitters). - test_allreduce_ring_torus_mesh.py: correctness across ring/torus/mesh. - test_distributed_default_topology.py: full distributed path on topology.yaml. - test_plot_latency_sweep.py / test_plot_buffer_kind_sweep.py: sweep rows. - test_plot_topology_diagram.py / test_plot_comparison_fsim.py: plot emitters. - test_intercube_root_center.py: moved in (ADR-0032 center-root latency guard). Also: - Move the FSIM comparison plot generator out of scripts/ into the sccl suite. - Delete superseded test files (test_allreduce_multidevice, test_distributed_lrab_hierarchical_allreduce, test_allreduce_buffer_kind_sweep) and repoint conftest aggregators + the ipcq buffer-kind importers. - Regenerate the allreduce_latency_plots derived artifacts from the full sweep. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-20 22:24:43 -07:00
parent ff7d727ddd
commit b610cb0d9a
22 changed files with 745 additions and 759 deletions
@@ -0,0 +1,58 @@
+"""Allreduce latency sweep (distributed path), xdist-friendly.
+
+Each parametrized case writes one JSON row to the shared staging dir; the
+conftest sessionfinish hook calls ``_aggregate_sweep_plots`` to emit the
+per-topology PNGs + summary.csv after all cases finish.
+"""
+from __future__ import annotations
+
+import json
+
+import pytest
+
+from tests.sccl._allreduce_helpers import (
+    _ELEM_BYTES_F16,
+    _SWEEP_ROWS_DIR,
+    _crit_ns,
+    _run_distributed,
+    _sweep_params,
+    _write_temp_configs,
+)
+
+
+@pytest.mark.parametrize(
+    "algorithm,sip_topology,n_sips,sip_w,sip_h,n_elem", _sweep_params(),
+)
+def test_allreduce_latency_one(
+    tmp_path, monkeypatch, algorithm, sip_topology, n_sips, sip_w, sip_h,
+    n_elem,
+):
+    topo_path, _ = _write_temp_configs(
+        tmp_path, sip_topology, n_sips, algorithm,
+        sip_w=sip_w, sip_h=sip_h,
+        n_elem_override=n_elem,
+    )
+    engine, n_cubes = _run_distributed(
+        tmp_path, monkeypatch, topo_path,
+        f"sweep_{algorithm}_{sip_topology}_{n_elem}", n_elem,
+    )
+
+    crit_ns = _crit_ns(engine)
+
+    bytes_per_sip = n_cubes * n_elem * _ELEM_BYTES_F16
+    bytes_per_pe = n_elem * _ELEM_BYTES_F16
+
+    record = {
+        "algorithm": algorithm,
+        "sip_topology": sip_topology,
+        "n_sips": n_sips,
+        "n_elem": n_elem,
+        "bytes_per_pe": bytes_per_pe,
+        "bytes_per_sip": bytes_per_sip,
+        "latency_ns": crit_ns,
+    }
+
+    _SWEEP_ROWS_DIR.mkdir(parents=True, exist_ok=True)
+    row_path = _SWEEP_ROWS_DIR / f"{sip_topology}_{n_elem}.json"
+    with open(row_path, "w", encoding="utf-8") as f:
+        json.dump(record, f)