tiling.generate_plan_from_ops: scan flat ops for the GEMM (<=1); pre-GEMM KERNEL ops become single-shot prologue MATH stages, post-GEMM ops split by scope (K_TILE/OUTPUT_TILE epilogue + KERNEL post-loop). MATH-only composite reproduces the legacy math-head plan. Prologue/post-loop stages fold into the first/last tile so the feeder + completion counting are untouched (existing benches have neither -> byte-equal op_log). PipelinePlan gains prologue_stages/epilogue_stages. pe_scheduler._generate_plan delegates to generate_plan_from_ops. DMA keeps the existing pinned signal (NOT a space flip); recipe scratch/primary-out handles are pinned=True so the head GEMM's auto-bound a (=P) is consumed in place. ADR-0065 D4/D6.7/Test#5 amended (EN+KO): DMA decision from pinned, not space; prologue recipe ops TCM-only (head op exempt -- it may stream from HBM). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
21 KiB
ADR-0065: 평평한 ops CompositeCmd + 첫 stateful recipe (softmax_merge)
Status
Proposed (verification gated on ADR-0064 Revision 2 land)
ADR-0060 (AHBM GQA Fused Attention) 의 보조 ADR. §5.6 / §8 item 4 의 carve-out 을 정확히 구현: decode opt2 = (#1 기존 Q·Kᵀ GEMM composite) + (#2 softmax + P·V + online-softmax merge 를 담은 단일 composite). ADR-0060 §5.6 의 "ex_composite" / "flash_pv_merge" 명칭은 retire — 본 ADR 이 정식 모양을 확정: 기존
tl.composite입구를 그대로 사용하면서 두 가지 구조적 추가 — (a) 평평한ops튜플, (b) MATH micro-op 시퀀스로 펼쳐지는 첫 stateful recipesoftmax_merge.
Context
ADR-0060 가 이 ADR 에 남긴 것
ADR-0060 §5.6 는 opt3 (software pipelining) 를 지금 ship 으로 권고, opt2 는 — per-tile dispatch 적음, K-before-V DMA priority, ADR-0064 의 cost model 이후에만 측정 가능 — 후속으로 carve out.
§8 item 4 가 carve-out 의 sizing: "재방문 시 두 개 composite 로 분할 —
#1 = Q·Kᵀ (기존 composite + scale), #2 = softmax + P·V + 온라인-softmax
누산기 merge". 본 ADR 이 #2 를 구현.
composite 가 구조 변경이 필요한 이유
현재 CompositeCmd 는 (op, a, b, out_addr, ops=(head, *epi)) 모양 + 암묵 컨벤션:
op가 head 엔진 선택 (gemm | math)ops[0]는 head OpSpec,ops[1:]는 epilogue OpSpec 으로 head 이후 per-output-tile 또는 per-K-tile 에 실행 (scope 결정)
decode opt2 는 MATH op 이 head GEMM 이전 에 실행 (m, l, O 의
online-softmax 갱신 + GEMM 입력 P emit) 이 필요. 현 모양엔 자리 없음.
이전에 검토 후 기각된 경로:
- softmax_merge 를 #1 (Q·Kᵀ) 의 epilogue 로 두기 —
Sj를 읽고 다음 GEMM (#2 P·V) 가 소비할P를 쓰는 구조라 순환. - 단일 mixed-engine recipe
flash_pv_merge가 MATH + GEMM + MATH 흡수 — engine boundary 위반, PE_SCHEDULER 복잡화.
깔끔한 모양: command 포맷에서 head/epilogue 구분 제거. ops 를 평평한
순서 튜플로, 각 op 의 scope + position 이 tile-loop 배치 결정. "prologue"
개념은 사용자 API 의 tl.composite(...) 에 ergonomic 으로 남되 —
HW command 는 보지 않음.
recipe 가 (generic op list 가 아닌) 이유
decode opt2 의 #2 MATH chain 은 8 단계 고정 구조 (online softmax merge).
커널 작성자에게 8 단계 직접 와이어링은 verbose + correctness hazard
(단계 순서 의존). Recipe — TLContext 가 8 단계로 (주소 미리 채워)
펼치는 단일 명명 op (softmax_merge) — 가 커널을 간결히 + 구조를 단일
source 에 보존. recipe 표는 TLContext (compiler analog) 에 위치;
PE_SCHEDULER 는 안 읽음. D5 boundary 참조.
Decision
D1. CompositeCmd 는 평평한 ordered op 리스트
@dataclass(frozen=True)
class CompositeCmd:
completion: CompletionHandle
ops: tuple[OpSpec, ...] # ordered MATH/GEMM ops
rw_handles: tuple[TensorHandle, ...] = () # cross-composite hazard
data_op: bool = True
@property
def logical_bytes(self) -> int: ... # ADR-0064 D2
이전의 (op, a, b, out_addr, out_nbytes, math_op) 필드는 ops 에 흡수.
ops 의 첫 OpSpec 이 kind == "gemm" 이면 head GEMM; 앞의 OpSpec 들이
pre-loop, 뒤의 OpSpec 들이 scope 에 따라 post-loop / per-tile epilogue.
D2. OpSpec 이 named operand 보유
@dataclass(frozen=True)
class OpSpec:
kind: str # "gemm" | "rmax" | ...
scope: Scope # KERNEL | K_TILE | OUTPUT_TILE
operands: dict[str, TensorHandle] # named — D5 참조
extra: dict[str, Any] # scalars, axes, m/k/n 등
out: TensorHandle | None = None # 명시적 write-back handle
@property
def logical_bytes(self) -> int: ...
named operand 로 PE_SCHEDULER 가 (positional 컨벤션 없이) 엔진 port 에
입력 라우팅 (예: GEMM 은 operands["a"], operands["b"]).
D3. Position + scope 가 tile-loop 배치 결정
Semantics 분리. Position 이 phase (tile-loop 수명주기 어디서
실행되는지) 결정; scope 가 그 phase 안에서의 repetition 결정.
Scope.KERNEL 의미는 "CompositeCmd invocation 당 1 회" — kernel
launch 당 1 회 아님. 그 1 회 실행이 op 의 시퀀스 내 position 에서
발생. 같은 KERNEL 값이 OpSpec 이 GEMM op 의 앞/위치/뒤에 있느냐에 따라
phase 가 달라짐.
PE_SCHEDULER 가 cmd.ops 에서 GEMM op 검색 (composite 당 ≤ 1, D6 참조).
인덱스 g:
| OpSpec position | scope | Phase | plan 내 배치 |
|---|---|---|---|
0 .. g-1 |
KERNEL | pre-tile-loop | tile loop 진입 전 1 회 |
g |
KERNEL | head GEMM | extra["m"], extra["k"], extra["n"] 로 tile loop drive |
g+1 .. |
K_TILE | in-accumulation | per K-tile epilogue (K 누산 loop 안) |
g+1 .. |
OUTPUT_TILE | per-output-tile | per (m, n) tile epilogue (K 누산 후) |
g+1 .. |
KERNEL | post-tile-loop | tile loop 종료 후 1 회 |
GEMM op 가 없는 composite (예: MATH-only) 는 모든 op 를 KERNEL-scope 직렬로 처리 — phase 가 "순서대로 single-shot" 으로 collapse.
D4. PE_SCHEDULER 가 operand pinned 플래그 보고 DMA 자동 삽입
PE_SCHEDULER 가 GEMM 의 operands 검사. 각각:
- not
pinned(데이터가 HBM 에 있어 스트리밍 필요) → FETCH/GEMM Stage 앞에 DMA_READ Stage 삽입 pinned(이전tl.load로 이미 TCM 에 staged, 또는 recipe 의 TCM scratch / primary-outP) → DMA Stage 없음, in-place 소비
커널은 composite operand 의 명시적 DMA cmd 를 절대 emit 안 함;
PE_SCHEDULER 가 핸들에서 추론. tl.ref operand 는 not pinned (DMA_READ);
tl.load 결과와 recipe scratch / primary-out 은 pinned (in-place).
as-built 노트 (P3). D4 의 초안은 DMA 결정을
space태그 (hbm→DMA,tcm→in-place) 에서 도출하며 현재space값 (오늘tl.load=hbm,tl.ref=tcm) 을 뒤집어야 했음. 구현은 기존pinned플래그를 DMA 신호로 그대로 유지 —pinned이 이미 "이미 TCM 에 있음, DMA skip" 을 인코딩하고,space뒤집기는 기존 bench Stage 시퀀스를 바꾸면서 기능적 이득이 없음.space통일 +pinned폐기는 후순위 cleanup 으로 연기.
Prologue recipe op 은 TCM-only (Phase 1 한계, TLContext emit 시 강제). prologue recipe (비 GEMM) OpSpec 의 모든 operand 는 반드시 TCM 상주; HBM 핸들을 recipe operand 로 전달하면 emit 시 validation error. Phase 1 에서 DMA staging 을 head op 의 operand 로 제한. head op 자체 (gemm 또는 math) 는 기존 DMA-staged-from-HBM 동작 유지 — D4 의 TCM-only 룰은 head 에 적용 안 됨. 이 invariant 는 D6 #7 에 재기술.
D5. RECIPE_DESCRIPTORS — TLContext 내부, pe_commands.py 가 아님
Recipe 는 새 TLContext-adjacent 모듈
(src/kernbench/triton_emu/tl_recipes.py) 에 위치. PE_SCHEDULER 는
import 안 함. 첫 recipe:
@dataclass(frozen=True)
class RecipeDescriptor:
operands: dict[str, str] # name → "R" | "RW"
primary_out: PrimaryOutSpec | None # 암묵 primary output 의 type/shape 룰
tile_alignment: Literal["single_shot", "tile_aligned"]
internal_scratch_bytes_fn: Callable[..., int]
engine_seq: tuple[EngineOp, ...] # TLContext 가 평평한 OpSpec 으로 펼침
RECIPE_DESCRIPTORS["softmax_merge"] = RecipeDescriptor(
operands={"s": "R", "m": "RW", "l": "RW", "O": "RW"},
primary_out=PrimaryOutSpec(
from_shape="s", from_dtype="s", transform="identity",
),
tile_alignment="single_shot",
internal_scratch_bytes_fn=lambda G, TILE, d, bpe: bpe * (
G + G + G + G * TILE + G # m_loc + m_new + corr + P + l_loc
),
engine_seq=(
EngineOp("MATH", "rmax", src="s", dst="m_loc", reduce_axis=-1),
EngineOp("MATH", "max_elem", src_a="m", src_b="m_loc", dst="m_new"),
EngineOp("MATH", "exp_diff", src_a="m", src_b="m_new", dst="corr"),
EngineOp("MATH", "exp_diff", src_a="s", src_b="m_new", dst="P", bcast_axis=0),
EngineOp("MATH", "rsum", src="P", dst="l_loc", reduce_axis=-1),
EngineOp("MATH", "fma", src_a="l", src_b="corr", src_c="l_loc", dst="l"),
EngineOp("MATH", "mul_bcast", src_a="O", src_b="corr", dst="O", bcast_axis=1),
EngineOp("MATH", "copy", src="m_new", dst="m"),
),
)
TLContext 가 tl.composite(prologue=[{"op": "softmax_merge", ...}], op="gemm", ...) 에서:
- operand R/RW 를
RECIPE_DESCRIPTORS["softmax_merge"]와 검증. - scratch (m_loc, m_new, corr, P, l_loc) 를 scratch_scope helper 로 할당 (ADR-0063).
- primary output P 의 shape 을
s.shape에서 derive (identity). engine_seq를 8 개의 평평한 MATH OpSpec 으로 펼침 (모든 주소/크기 채움). 각 OpSpec 의scope = KERNEL.- Auto-bind (충돌 검사, D6.6). head GEMM 이
operands["a"]를 명시하지 않은 경우에만operands["a"] = P_handleauto-bind. 커널이a를 명시적으로 제공했다면 validation error — 모호한 primary-output 대체. CompositeCmd(ops=(8 MATH + 1 GEMM + epilogue), rw_handles=(m, l, O))emit.
RECIPE_DESCRIPTORS 는 HW 경로 어디에도 안 나타남. PE_SCHEDULER 는 평평한 ops 리스트만 봄.
D6. Invariants
- GEMM 수.
0 ≤ count(op.kind == "gemm" in cmd.ops) ≤ 1. GEMM 0 개면 MATH-only composite. - softmax_merge 가 V operand 없음.
RECIPE_DESCRIPTORS[ "softmax_merge"].operands에 HBM-resident handle 없음 → prologue 동안 DMA 삽입 없음 → V DMA 는 GEMM head 시작 시에만 발생 → decode opt2 의 #1 / #2 간 K-before-V DMA priority 자연 강제. - Cross-composite strict FIFO. PE_SCHEDULER 가 in-flight
rw_handles추적. 신규 composite 의rw_handles(또는 op 입력) 가 in-flight 와 교차하면 모든 이전 composite 완료까지 대기 — out-of-order reorder 없음. - legacy caller 후방 호환. 사용자 API
tl.composite(op="gemm", a, b, epilogue=[...])보존. TLContext 가 내부적으로 평평한 opsCompositeCmd로 lowering. - PE_SCHEDULER recipe-free. PE_SCHEDULER 는 RECIPE_DESCRIPTORS import 안 함, "softmax_merge" 의 존재 모름, recipe 별 코드 분기 없음.
- 암묵 operand 대체 없음 (auto-bind 충돌). prologue recipe 가
primary_out선언하고 head op 의 auto-bind 대상 operand (예: GEMMa) 도 커널이 명시적 으로 제공하면, TLContext 가 emit 시 validation error. 커널은 recipe 가 바인딩하게 두거나 (operand 생략) 혹은 명시적으로 지정 (prologue 생략, 또는primary_out없는 recipe 사용) 해야 함. "어느 값이 이기느냐" 의 모호함 방지. - Prologue MATH operand TCM-only. D4 의 재기술: prologue recipe (비 GEMM) OpSpec 의 모든 operand 는 TCM 상주. head op (gemm 또는 math) 은 예외 — HBM 스트리밍 허용. Phase 1 에서 TLContext emit 시 강제.
D7. Boundary 요약 (compiler vs scheduler vs engine)
HOST DEVICE
TLContext (compiler analog) PE_SCHEDULER (scheduler/dispatcher)
- RECIPE_DESC 와 검증 - 평평한 ops 리스트 스캔
- primary_out shape derive - GEMM 식별 (≤ 1)
- scratch 할당 - position-기반 stage 배치
- recipe → 평평한 OpSpec 펼침 - space 보고 DMA stage 자동 삽입
- rw_handles 계산 - strict-FIFO RW 해저드 추적
- CompositeCmd emit - Stage 리스트 emit
imports: RECIPE_DESCRIPTORS imports: recipe 관련 없음
ENGINES (PE_MATH / PE_GEMM / PE_DMA)
- Stage.params 읽음 (op_kind, addresses, n_elements)
- 기존 latency 모델 변화 없음
imports: recipe 관련 없음
Alternatives
A1. tile 당 3 composite (prologue 개념 없이)
#2 를 softmax_merge + gemm(P·V) + add(O) 의 3 개로 분할. 장점:
flat-ops 재구조화 없음 (add 가 기존 epilogue 경로 재사용). 단점: tile
당 dispatch 3 회 (대신 2) — ADR-0060 §8 item 4 sizing ("두 개 로
분할") 과 어긋남; ~32 ns × N_tiles 의 fixed-cost 절감 소실. 기각.
A2. Mixed-engine recipe (flash_pv_merge)
MATH + GEMM + MATH 흡수하는 단일 recipe (ADR-0060 §5.6 원안 표현). 장점: 가장 적은 dispatch (1 composite). 단점: PE_SCHEDULER 가 engine-경계 recipe 알아야 함, "엔진은 Stage.op_kind 만" boundary 위반, recipe 표가 ISA-급이 되어 HW 인터페이스 오염. 기각.
A3. Generic op-list DAG (tl.ex_composite([...]))
사용자 정의 임의 op 시퀀스 + 명명 슬롯 + 데이터플로 분석. 장점: 최대 유연성. 단점: 단일 user (softmax_merge) 가 generic DAG 언어를 정당화 못함; PE_SCHEDULER 에 데이터플로 분석 필요; ADR-0060 §8 item 4 의 "inner loop 흡수" 트랩 재등장. 기각 (Simplicity First).
A4. RW-aware scheduler reorder (strict FIFO 가 아님)
PE_SCHEDULER 가 rw_handles 교차 없으면 후속 composite 가 in-flight 를
추월 허용 (예: 다음 tile 의 #1 가 현재 tile 의 #2 와 겹침 — #1 은
m/l/O 안 만짐). 장점: 더 많은 overlap. 단점: scheduler state 증가, 이득
한정 (다음 #1 은 어차피 K DMA 대기). 연기 — strict FIFO 먼저;
RW-aware reorder 는 후속 ADR 가능.
A5. softmax_merge 를 #1 (Q·Kᵀ) 의 epilogue 로 embedding
오답: softmax_merge 의 출력 P 는 #2 의 P·V GEMM 의 입력, #1 의
epilogue 이후 실행. embedding 하면 #2 가 자기 입력을 #1 의 epilogue
scratch 에서 읽는 형태 — 순환 tile-loop 의존성. 기각 (incorrect).
Consequences
Positive
- ADR-0060 §5.6 / §8 item 4 carve-out 정확 구현: opt2 = tile 당 2 composite, K-before-V DMA priority 자연.
- composite contract 균일화: scope-기반 배치의 평평한 ops 리스트. HW cmd 에서 "head" 와 "epilogue" 구조 구분 없음.
- DMA 가 operand
space에서 자동 추론; 커널 표면 단순화. - HW 인터페이스 (
CompositeCmd모양) 가 recipe 증가에도 작게 유지 — recipe 펼침이 host-side. - 새 fused 패턴 (linear+gelu+norm, rmsnorm+linear 등) 은
RECIPE_DESCRIPTORSentry 추가만으로 가능; PE_SCHEDULER 변경 없음.
Negative
- CompositeCmd 구조 변경 — 모든 현재 caller 의 refactor. 의미 보존 (D6.4); ADR-0064 Revision 2 land 후 기존 골든 불변.
OpSpec.operands가 positionaltuple[Any, ...]에서dict[str, TensorHandle]로 변경 — 기존 epilogue lowering 접촉.- 새 TLContext 코드: recipe 펼침 + scratch slot 할당 + primary_out auto-bind. ~80–120 LOC.
- PE_SCHEDULER 의
_generate_plan에 position-기반 스캔 + DMA 자동 삽입- strict-FIFO RW tracker. ~60–100 LOC.
- strict-FIFO RW 해저드 추적이 안전하지만 보수적. 두 composite 가
rw_handle을 공유할 때 신규는 모든 이전 겹치는 composite 가 완전히 완료 될 때까지 대기 — 원칙적으로 overlap 가능한 경우에도 (예: 다음-tile #1 = Q·Kᵀ 가 m/l/O 안 만지므로 현재-tile 의 #2 와 overlap 가능하지만, 같은 핸들을 만지는 후속 path 도 FIFO 로 직렬화). 더 똑똑한 scheduler 대비 composite-level overlap 을 과소 노출. RW-aware reorder (A4) 가 연기된 개선.
Open review items
- identity 를 넘는 recipe shape derivation 룰. 미래 recipe 가
non-identity transform (예: 전치된 primary_out) 필요할 수 있음.
PrimaryOutSpec.transform필드가 forward-compat; 필요 시 새 transform 추가. - Recipe scratch allocator 의 ADR-0063
scratch_scope통합. 초기 배선: TLContext 가 활성scratch_scope가 있으면 그 안에서 할당, 없으면 kernel-scope persistent slot. tl_recipes.py위치.triton_emu/아래 새 모듈. recipe 지식을 compiler analog (TLContext) 와 함께 위치.- GEMM
out=TensorHandle. 기존out_ptr: int와 함께 새 형태. 둘 다 허용; handle 형태가 신규 코드에 권장 (handle identity 로 strict-FIFO RW 추적 가능).
Test Requirements
- CompositeCmd 평평화 refactor — 의미 보존. 기존 모든 bench 의
op_log 가 refactor 전후로 byte-equal (legacy API
tl.composite( op="gemm", a, b, epilogue=[...])가 같은 Stage 시퀀스로 lowering). - softmax_merge recipe lowering. TLContext 호출
tl.composite( prologue=[{"op": "softmax_merge", "s": Sj, "m": m, "l": l, "O": O}], op="gemm", b=V_ref, out=O, epilogue=[{"op": "add", "other": O}])가 정확히 10 ops (8 MATH + 1 GEMM + 1 MATH(add)) +rw_handles == (m, l, O)의CompositeCmd생성. - K-before-V DMA priority invariant. decode opt2 의 op_log 에서 #2 의 MATH prologue 동안 V 관련 DMA 없음 (#2 의 V DMA 는 GEMM head 시작 시에만).
- Strict-FIFO RW 직렬화.
rw_handle공유하는 두 연속 composite 가 dispatch 순서대로 완료; stage 가 interleave 안 함. pinned에서 DMA 자동 삽입. not-pinnedoperand (예:b=tl.ref(...)) 의 GEMM composite 가 DMA_READ Stage emit;pinnedoperand (예: recipe 의 TCM scratch / primary-out, 또는tl.load결과) 는 emit 안 함. (as-built D4 노트: DMA 결정은space태그가 아니라 기존pinned플래그 사용.)- opt2 가 opt3 와 수치 동등. data mode 에서 opt2 의 최종
(m, l, O)가 opt3 와 fp tolerance 안. - opt2 dispatch ratio (ADR-0064 Rev2 이후) — robust. opt3 vs opt2
per-tile PE_CPU dispatch cycles 가
opt3 > 2 × opt2만족 (정성적 gate, calibration-독립). Default-calibration 모델 기대치는 ≈ 4.0× (FIXED=40 cycles, R=0.0625 cycles/byte) — DDD-0065 §11 의 모델-유도 숫자 참조; 테스트 gate 는 느슨한> 2×한계라 calibration 변경에도 살아남음. Ratio 가 FIXED-dominated — command-count 감소가 1차 신호임을 반영. - GEMM-count invariant. GEMM OpSpec 두 개 가진 composite 가 TLContext emit 시 validation error.
- Composite 크기 cap 안에 (ADR-0064 D7). Decode opt2 의
#2composite (10 ops, ~322 logical bytes — operand 핸들의 per-op 합산) 가 defaultMAX_COMPOSITE_LOGICAL_BYTES=1024안에 편안히 — segmentation 없음. recipe lowering 테스트가cmd.logical_bytes < 1024단언. - Auto-bind 충돌.
tl.composite(op="gemm", a=A_handle, prologue=[{"op": "softmax_merge", ...}], ...)처럼 softmax_merge 가primary_out선언하는데a도 명시되면 emit 시 validation error (D6.6). - MATH operand TCM invariant.
tl.composite(op="math", a=tl.ref(addr, shape), ...)(HBM-resident 핸들을 MATH operand 로 전달) 가 emit 시 validation error (D6.7).
Dependencies
- ADR-0060 §5.6, §8 item 4 — 본 ADR 이 구현하는 carve-out.
- ADR-0064 Revision 2 — dispatch cost model; 먼저 land, opt2 의 fewer-issues win 을 측정 가능하게 함.
- ADR-0063
scratch_scope— recipe intermediate 가 안에 할당;m, l, O는 persistent arena 로 밖에 할당 (DDD-0060 §6.2 참조). - ADR-0046
tl_contextcontract —prologue=[...]kwarg,out=TensorHandle, RECIPE_DESCRIPTORS lookup 으로 확장. - ADR-0042 tile plan generators — (재설계 아닌) 확장 — 평평한 ops 리스트 처리 + DMA 자동 삽입.
- ADR-0014 PE pipeline — boundary 보존: scheduler 가 recipe-free, 엔진이 op_kind-opaque.
Migration
구현 순서 노트. 아래의 ADR-0064-우선 순서가 원래 계획이었음. 실제 구현은 1↔2 단계를 뒤집어 — ADR-0065 Phase 1 (평평화
CompositeCmd) 이 ADR-0064 Rev2 보다 먼저 land. 이유:logical_bytes(ADR-0064 D2) 가 평평화 형태 위에서 자연스럽게 정의되므로, refactor 를 먼저 하면 버려지는 transitionallogical_bytes를 피함. 각 단계는 독립적으로 안전: P1 은 기존 cost 경로 하에서 의미 보존 refactor (op_log byte-equal); ADR-0064 Rev2 가 이미 평평화된 형태 위에 구조적 공식을 교체. 최종 상태는 어느 쪽이든 동일.
Land 순서 (원래 계획; as-built 1↔2 뒤집힘은 위 노트 참조):
- ADR-0064 Revision 2 (별도 PR): 구조적 dispatch cost +
logical_bytesproperty + topology config override + 일회성 골든 재생성. - ADR-0065 Phase 1 (본 ADR, PR a):
CompositeCmd평평화 refactor + legacy-API lowering. refactor only; 골든 불변. - ADR-0065 Phase 2 (본 ADR, PR b):
tl_recipes.py+softmax_mergerecipe + PE_SCHEDULER position-스캔 + DMA 자동 삽입 + strict-FIFO RW tracker. - ADR-0065 Phase 3 (본 ADR, PR c):
_gqa_decode_long.pyopt2 variant + 수치 동등 테스트 + opt3 대비 dispatch-ratio 측정.
상세 구현 계획: DDD-0065 참조.