- ADR-0060: GEMMs (Q.Kt, P.V) via existing tl.composite (scheduler-managed tiling + K/V DMA streaming); softmax merge + IPCQ tree reduction stay in kernel. Front TL;DR pseudocode of the final composite kernel; new section B lists open design items (DDD sync, K pre-transpose, dma_read lever, kernel-vs-scheduler tiling, ring path). - ADR-0062: redefined from a new load_async op to global lazy tl.load (non-blocking + auto-wait on first use; API unchanged; goldens regenerate). - ADR-0064 (new): per-op-type CPU issue cost model (composite ~40ns >> primitive) so the hybrid's CPU-saturation win becomes measurable (currently dispatch_cycles=0 hides it). Cost-model impl deferred. - KO mirrors for ADR-0060/0062/0064 (-ko suffix, adr-proposed). Rationale: non-blocking CompositeCmd offloads tiling to PE_SCHEDULER, decoupling CPU issue-rate from execution so the CPU can saturate the engines; the prior 'composite = no latency benefit' claim was an artifact of dispatch_cycles=0. Docs only; no production code changed. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
9.2 KiB
ADR-0064: op 종류별 CPU 발행 비용 모델 (command construct + dispatch)
Status
Proposed
ADR-0060(AHBM GQA Fused Attention)의 보조 ADR. 거기서의 하이브리드 결정(GEMM은
tl.composite, softmax 머지는 커널)이 이기는 이유는 타일링을 PE_SCHEDULER에 offload하여 CPU가 coarse descriptor를 내리고 앞서 나가 엔진을 saturate하기 때문이다. 그 이점이 현재 시뮬레이터에서는 per-op CPU 발행 비용이 0이라 보이지 않는다. 본 ADR은 발행 비용을 실재화하고 op 종류별로 차등화하여, composite-vs-primitive 트레이드오프 (그리고 CPU saturation)가 측정 가능하도록 한다.
Context
현재 상태
- 모든
tl.*op은 커맨드를 emit하기 전에_emit_dispatch_overhead()를 호출하며(tl_context.py:123-125, 190, 227, 235, 612, …), 이는dispatch_cycles > 0일 때만PeCpuOverheadCmd(cycles=dispatch_cycles)를 emit한다. 즉 발행 비용은tl.load,tl.dot, MATH op,tl.composite에 동일하게 적용되는 단일 균일 노브다. - 그 노브는 두 실행 경로 모두에서 0으로 하드코딩되어 있다
(
pe_cpu.py:101greenlet runner,:195legacy replay). ⇒ 커맨드 발행 — descriptor를 construct하고 scheduler 큐에 push — 이 현재 PE_CPU에서 0 ns다. PeCpuOverheadCmd는 PE_CPU(kernel_runner.py:131-132)와 PE_SCHEDULER(pe_scheduler.py:97-100)에서yield env.timeout(cmd.cycles)로 소비된다. 수동tl.cycles(n)도 존재한다(tl_context.py:695).
본 ADR을 규정하는 두 가지 발견 (ADR-0060 검토에서)
- Q1 — composite 발행 비용.
tl.composite를 construct + push 하는 비용은 CPU 시간 기준 ~40 ns 정도(descriptor 빌드 + 큐 push)로 예상된다. 현재는 0이다. 훅은 존재하며, 값(그리고 op 종류별 차등)만 빠져 있다. - Q2 — scheduler dispatch vs DMA latency. PE_SCHEDULER의 composite
dispatch는 non-blocking이다:
_dispatch_composite는 tile plan을 생성하고 feeder에 enqueue한 뒤 즉시 반환한다(pe_scheduler.py:104-121). 실제 DMA latency는 타일이 흐를 때 PE_DMA에서 부과되며(drain_ns = compute_drain_ns(path, nbytes),pe_dma.py:89), scheduler dispatch에 덩어리로 들어가지 않는다 ⇒ 중복 계상 없음, DMA는 modeled component에 머문다(SPEC §0.1). scheduler의 plan 생성 자체는 현재 sim 시간 0이다.
균일-그리고-0이 하이브리드에 틀린 이유
하이브리드의 논거 전체는 하나의 composite descriptor가 N_tiles만큼의
GEMM 타일링을 offload하므로, CPU가 O(N_tiles × ops/tile)개의 fine 커맨드
대신 O(1)개의 coarse 커맨드를 낸다는 것이다. 발행 비용 = 0(그리고 균일)이면
모델은 다음을 보여줄 수 없다:
- CPU가 충분히 빨리 못 밀어넣을 때 primitive 경로가 엔진을 saturate 못 할 수 있음(핵심 ADR-0060 §1 주장), 그리고
- composite가 단일 primitive보다 construct 비용이 더 크지만 그것이 대체하는 다수의 primitive보다는 훨씬 작음.
단일 균일 dispatch_cycles로는 이를 표현할 수 없다: 실제 construct 비용에서
tl.load ≪ tl.composite이기 때문이다.
Decision
D1. PE_CPU의 op 종류별 발행 비용 테이블
단일 dispatch_cycles 스칼라를 커맨드 종류별 비용 테이블로 교체하고,
발행 시점(커맨드가 scheduler로 dispatch되기 전)에 PE_CPU에서 부과한다.
greenlet이 이 시간을 지불하며, 이것이 바로 CPU가 얼마나 빨리 work를 밀어넣을
수 있는지를 gate한다 — saturation 레버.
시작 추정치(추후 calibrate — 검토 항목 참조):
| 발행 op | CPU 발행 비용 (construct + push) |
|---|---|
tl.composite (GEMM/MATH descriptor) |
~40 ns (Q1 추정) |
tl.load / tl.store (DMA descriptor) |
작음 (수 ns) |
tl.dot / MATH / IPCQ send/recv |
작음 (수 ns) |
요점은 비율이다: composite construct ≫ primitive 발행, 그러나 그것이 대체하는 primitive 발행들의 합 ≪. 정확한 ns는 설정 가능하다.
D2. 실행 latency는 엔진에 유지 (Q2 — 변경 없음)
DMA/GEMM/MATH 실행 latency는 오늘처럼 PE_DMA / PE_GEMM / PE_MATH에서 부과된다.
발행 비용(D1)은 CPU 측에만 추가되며, drain_ns를 건드리지 않으므로 중복
계상이 없다. 이는 SPEC §0.1(latency는 modeled component에서)을 보존한다.
D3. scheduler plan 생성 비용 — 0에서 시작, 재검토
PE_SCHEDULER의 tile-plan 생성은 오늘 sim 시간 0이다. 일단 유지한다(지배적
레버는 CPU 발행 비용 D1); scheduler 측 비용이 유의미하다고 판명되면 기존
overhead_ns node attr로 노출한다. 여기서 결정하지 않고 검토 항목으로 둔다.
D4. 설정 가능한 값; 골든 재생성
비용 테이블은 설정 가능하다(per-topology / node attrs, 기본값은 D1 추정치). 발행 비용을 0이 아니게 하면 모든 bench latency가 바뀌므로, 골든 latency를 한 번 재생성한다 — ADR-0062의 전역 lazy-load 회귀와 동일한 자세의 모델 충실도 개선이다.
Alternatives
A1. 단일 균일 dispatch_cycles > 0 유지
기각: op construct 비용은 자릿수 단위로 다르다(tl.load vs tl.composite).
균일값은 primitive를 과대 청구하거나 composite를 과소 청구하며, 어느 쪽이든
하이브리드가 의존하는 composite-vs-primitive 트레이드오프를 왜곡한다.
A2. 발행 비용을 PE_CPU 대신 PE_SCHEDULER에 부과
기각: saturation 질문은 *"CPU가 엔진을 바쁘게 유지할 만큼 descriptor를 빨리 밀어넣을 수 있는가?"*이며 — 이는 PE_CPU의 issue-bandwidth 속성이다. scheduler에 부과하면 CPU back-pressure를 모델하지 못한다.
A3. DMA program/setup 시간을 별도의 고정 per-descriptor 비용으로 모델
실제 HW는 전송 시간과 별개로 DMA descriptor program 비용을 지불한다. 연기:
초기에는 descriptor-program 비용을 발행 op의 D1 비용에 합친다(DMA를
유발하는 tl.load / composite). calibration이 유의미함을 보이면 PE_DMA 고정
setup으로 분리한다. 검토 항목.
Consequences
Positive
- 하이브리드의 CPU-offload / saturation 이점(ADR-0060 §1)이 구조적일 뿐 아니라 측정 가능해진다.
- composite-vs-primitive 및 타일링 granularity 트레이드오프가 latency에 보여 eval 서사를 가능케 한다.
- 하드웨어에 더 충실(issue bandwidth는 실제 병목).
Negative
- 모든 bench 골든이 이동 → 한 번의 재생성(D4); CI 골든 fixture 갱신.
- op 종류별 값은 calibration 필요; ~40 ns composite 수치는 추정이고 primitive는 미지정 — 결과는 숫자만큼만 정확하다(calibrate 전까지 절대 latency를 과대 주장 금지).
- 발행 경로에 비용 테이블 lookup 추가(런타임 영향 미미).
Open review items (자율 결정; 검토 시 수정)
- op 종류별 값의 calibration 출처. composite ≈ 40 ns는 작업 추정치; primitive 발행 비용은 placeholder. 권고: 문서화된 가정(instruction-issue + 큐 push)에서 값을 택하고, 실제 reference가 생기기 전까지 절대 latency를 잠정으로 취급; 비율을 방어 가능하게 유지.
- scheduler plan-gen 비용 (D3). 권고: 초기에는 0 유지; scheduler-bound
동작을 보이는 workload가 나오면
overhead_ns로 노출. - DMA program 시간 (A3). 권고: 먼저 발행 op의 비용에 합치고; 필요 시 PE_DMA setup으로 분리.
- 테이블 위치. 권고: 흩어진 node attr이 아니라, 커맨드 종류로 키잉되고 per-topology 오버라이드 가능한 작은 중앙 cost 모듈 — 값을 한 곳에서 검토 가능하도록.
- 회귀 롤아웃. 권고: lazy-load(ADR-0062)와 본 ADR의 cost model을 한 번의 골든 재생성 패스로 함께 도입하여 두 번의 churn을 피함.
- legacy replay 경로와의 상호작용(
pe_cpu.py:_execute_legacy) — greenlet과 replay 경로가 동일한 cost 테이블을 읽어 결과가 일치하도록 보장. 검증 요.
Test Requirements
- composite는 한 번, primitive는 op마다 청구.
N_tiles에 걸쳐 하나의tl.composite를 내는 커널은 PE_CPU에서 composite 발행 비용을 한 번 청구하고; 동등한N_tiles × opsprimitive 커널은 op당 비용을N_tiles × ops번 청구. PE_CPU busy time이 그에 따라 다름을 assert. - DMA latency 불변 (Q2). 고정된 전송에 대해 PE_DMA
drain_ns는 ADR 이전 값과 동일 — 발행 비용은 DMA에 합쳐지지 않고 PE_CPU에 가산(중복 계상 없음). - saturation 관측 가능. per-op 발행 비용이 0이 아니면, many-tile primitive sweep는 GEMM 엔진 idle(CPU-bound 발행)을 보이고 composite sweep는 바쁘게 유지 — ADR-0060 §1 레버.
- 결정성: 동일 입력 → 동일 op_log + latency (SPEC §0.1).
- 경로 일치: greenlet과 legacy-replay 경로가 동일 커널에 대해 동일한 발행 비용 회계를 산출.