Files
kernbench2/docs/adr-proposed/ADR-0065-prog-flat-ops-composite-softmax-merge-ko.md
T
ywkang 55f025c4b1 gqa(adr-0065): P3 — flat-ops PE_SCHEDULER plan (position-scan + prologue stages)
tiling.generate_plan_from_ops: scan flat ops for the GEMM (<=1); pre-GEMM KERNEL ops become single-shot prologue MATH stages, post-GEMM ops split by scope (K_TILE/OUTPUT_TILE epilogue + KERNEL post-loop). MATH-only composite reproduces the legacy math-head plan. Prologue/post-loop stages fold into the first/last tile so the feeder + completion counting are untouched (existing benches have neither -> byte-equal op_log).

PipelinePlan gains prologue_stages/epilogue_stages. pe_scheduler._generate_plan delegates to generate_plan_from_ops. DMA keeps the existing pinned signal (NOT a space flip); recipe scratch/primary-out handles are pinned=True so the head GEMM's auto-bound a (=P) is consumed in place.

ADR-0065 D4/D6.7/Test#5 amended (EN+KO): DMA decision from pinned, not space; prologue recipe ops TCM-only (head op exempt -- it may stream from HBM).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-10 20:24:33 -07:00

406 lines
21 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# ADR-0065: 평평한 ops `CompositeCmd` + 첫 stateful recipe (`softmax_merge`)
## Status
Proposed (verification gated on ADR-0064 Revision 2 land)
> **ADR-0060** (AHBM GQA Fused Attention) 의 보조 ADR. §5.6 / §8 item 4 의
> carve-out 을 정확히 구현: decode opt2 = (#1 기존 Q·Kᵀ GEMM composite) +
> (#2 softmax + P·V + online-softmax merge 를 담은 단일 composite). ADR-0060
> §5.6 의 "ex_composite" / "flash_pv_merge" 명칭은 *retire* — 본 ADR 이
> 정식 모양을 확정: 기존 `tl.composite` 입구를 그대로 사용하면서 두 가지
> 구조적 추가 — (a) 평평한 `ops` 튜플, (b) MATH micro-op 시퀀스로
> 펼쳐지는 첫 stateful recipe `softmax_merge`.
## Context
### ADR-0060 가 이 ADR 에 남긴 것
ADR-0060 §5.6 는 opt3 (software pipelining) 를 *지금 ship* 으로 권고,
opt2 는 — per-tile dispatch 적음, K-before-V DMA priority,
**ADR-0064 의 cost model 이후에만 측정 가능** — 후속으로 carve out.
§8 item 4 가 carve-out 의 sizing: "재방문 시 **두 개** composite 로 분할 —
`#1` = Q·Kᵀ (기존 composite + `scale`), `#2` = softmax + P·V + 온라인-softmax
누산기 merge". 본 ADR 이 `#2` 를 구현.
### composite 가 구조 변경이 필요한 이유
현재 `CompositeCmd``(op, a, b, out_addr, ops=(head, *epi))` 모양 + 암묵 컨벤션:
- `op` 가 head 엔진 선택 (gemm | math)
- `ops[0]` 는 head OpSpec, `ops[1:]` 는 epilogue OpSpec 으로 head *이후*
per-output-tile 또는 per-K-tile 에 실행 (scope 결정)
decode opt2 는 MATH op 이 head GEMM **이전** 에 실행 (`m, l, O`
online-softmax 갱신 + GEMM 입력 `P` emit) 이 필요. 현 모양엔 자리 없음.
이전에 검토 후 기각된 경로:
- softmax_merge 를 #1 (Q·Kᵀ) 의 *epilogue* 로 두기 — `Sj` 를 읽고 *다음*
GEMM (#2 P·V) 가 소비할 `P` 를 쓰는 구조라 순환.
- 단일 mixed-engine recipe `flash_pv_merge` 가 MATH + GEMM + MATH 흡수 —
engine boundary 위반, PE_SCHEDULER 복잡화.
깔끔한 모양: **command 포맷에서 head/epilogue 구분 제거**. `ops` 를 평평한
순서 튜플로, 각 op 의 `scope` + position 이 tile-loop 배치 결정. "prologue"
개념은 *사용자 API*`tl.composite(...)` 에 ergonomic 으로 남되 —
**HW command 는 보지 않음**.
### recipe 가 (generic op list 가 아닌) 이유
decode opt2 의 `#2` MATH chain 은 8 단계 고정 구조 (online softmax merge).
커널 작성자에게 8 단계 직접 와이어링은 verbose + correctness hazard
(단계 순서 의존). *Recipe* — TLContext 가 8 단계로 (주소 미리 채워)
펼치는 단일 명명 op (`softmax_merge`) — 가 커널을 간결히 + 구조를 단일
source 에 보존. recipe 표는 TLContext (compiler analog) 에 위치;
PE_SCHEDULER 는 안 읽음. D5 boundary 참조.
## Decision
### D1. `CompositeCmd` 는 평평한 ordered op 리스트
```python
@dataclass(frozen=True)
class CompositeCmd:
completion: CompletionHandle
ops: tuple[OpSpec, ...] # ordered MATH/GEMM ops
rw_handles: tuple[TensorHandle, ...] = () # cross-composite hazard
data_op: bool = True
@property
def logical_bytes(self) -> int: ... # ADR-0064 D2
```
이전의 `(op, a, b, out_addr, out_nbytes, math_op)` 필드는 `ops` 에 흡수.
`ops` 의 첫 OpSpec 이 `kind == "gemm"` 이면 head GEMM; 앞의 OpSpec 들이
pre-loop, 뒤의 OpSpec 들이 `scope` 에 따라 post-loop / per-tile epilogue.
### D2. `OpSpec` 이 named operand 보유
```python
@dataclass(frozen=True)
class OpSpec:
kind: str # "gemm" | "rmax" | ...
scope: Scope # KERNEL | K_TILE | OUTPUT_TILE
operands: dict[str, TensorHandle] # named — D5 참조
extra: dict[str, Any] # scalars, axes, m/k/n 등
out: TensorHandle | None = None # 명시적 write-back handle
@property
def logical_bytes(self) -> int: ...
```
named operand 로 PE_SCHEDULER 가 (positional 컨벤션 없이) 엔진 port 에
입력 라우팅 (예: GEMM 은 `operands["a"]`, `operands["b"]`).
### D3. Position + scope 가 tile-loop 배치 결정
**Semantics 분리.** *Position***phase** (tile-loop 수명주기 어디서
실행되는지) 결정; *scope* 가 그 phase 안에서의 **repetition** 결정.
`Scope.KERNEL` 의미는 "**`CompositeCmd` invocation 당 1 회**" — kernel
launch 당 1 회 아님. 그 1 회 실행이 op 의 시퀀스 내 *position* 에서
발생. 같은 KERNEL 값이 OpSpec 이 GEMM op 의 앞/위치/뒤에 있느냐에 따라
*phase* 가 달라짐.
PE_SCHEDULER 가 `cmd.ops` 에서 GEMM op 검색 (composite 당 ≤ 1, D6 참조).
인덱스 `g`:
| OpSpec position | scope | Phase | plan 내 배치 |
|---|---|---|---|
| `0 .. g-1` | KERNEL | pre-tile-loop | tile loop 진입 전 1 회 |
| `g` | KERNEL | head GEMM | `extra["m"], extra["k"], extra["n"]` 로 tile loop drive |
| `g+1 .. ` | K_TILE | in-accumulation | per K-tile epilogue (K 누산 loop 안) |
| `g+1 .. ` | OUTPUT_TILE | per-output-tile | per (m, n) tile epilogue (K 누산 후) |
| `g+1 .. ` | KERNEL | post-tile-loop | tile loop 종료 후 1 회 |
GEMM op 가 없는 composite (예: MATH-only) 는 모든 op 를 KERNEL-scope
직렬로 처리 — phase 가 "순서대로 single-shot" 으로 collapse.
### D4. PE_SCHEDULER 가 operand `pinned` 플래그 보고 DMA 자동 삽입
PE_SCHEDULER 가 GEMM 의 `operands` 검사. 각각:
- **not `pinned`** (데이터가 HBM 에 있어 스트리밍 필요) → FETCH/GEMM Stage
앞에 DMA_READ Stage 삽입
- **`pinned`** (이전 `tl.load` 로 이미 TCM 에 staged, 또는 recipe 의 TCM
scratch / primary-out `P`) → DMA Stage 없음, in-place 소비
커널은 composite operand 의 명시적 DMA cmd 를 절대 emit 안 함;
PE_SCHEDULER 가 핸들에서 추론. `tl.ref` operand 는 not pinned (DMA_READ);
`tl.load` 결과와 recipe scratch / primary-out 은 pinned (in-place).
> **as-built 노트 (P3).** D4 의 초안은 DMA 결정을 `space` 태그
> (`hbm`→DMA, `tcm`→in-place) 에서 도출하며 현재 `space` 값
> (오늘 `tl.load`=`hbm`, `tl.ref`=`tcm`) 을 뒤집어야 했음. 구현은 기존
> **`pinned` 플래그**를 DMA 신호로 그대로 유지 — `pinned` 이 이미
> "이미 TCM 에 있음, DMA skip" 을 인코딩하고, `space` 뒤집기는 기존 bench
> Stage 시퀀스를 바꾸면서 기능적 이득이 없음. `space` 통일 + `pinned`
> 폐기는 후순위 cleanup 으로 연기.
**Prologue recipe op 은 TCM-only (Phase 1 한계, TLContext emit 시 강제).**
*prologue recipe* (비 GEMM) OpSpec 의 모든 operand 는 **반드시** TCM 상주;
HBM 핸들을 recipe operand 로 전달하면 emit 시 validation error. Phase 1
에서 DMA staging 을 head op 의 operand 로 제한. **head op 자체** (gemm
또는 math) 는 기존 DMA-staged-from-HBM 동작 유지 — D4 의 TCM-only 룰은
head 에 적용 **안 됨**. 이 invariant 는 D6 #7 에 재기술.
### D5. RECIPE_DESCRIPTORS — TLContext 내부, `pe_commands.py` 가 아님
Recipe 는 새 TLContext-adjacent 모듈
(`src/kernbench/triton_emu/tl_recipes.py`) 에 위치. PE_SCHEDULER 는
import **안 함**. 첫 recipe:
```python
@dataclass(frozen=True)
class RecipeDescriptor:
operands: dict[str, str] # name → "R" | "RW"
primary_out: PrimaryOutSpec | None # 암묵 primary output 의 type/shape 룰
tile_alignment: Literal["single_shot", "tile_aligned"]
internal_scratch_bytes_fn: Callable[..., int]
engine_seq: tuple[EngineOp, ...] # TLContext 가 평평한 OpSpec 으로 펼침
RECIPE_DESCRIPTORS["softmax_merge"] = RecipeDescriptor(
operands={"s": "R", "m": "RW", "l": "RW", "O": "RW"},
primary_out=PrimaryOutSpec(
from_shape="s", from_dtype="s", transform="identity",
),
tile_alignment="single_shot",
internal_scratch_bytes_fn=lambda G, TILE, d, bpe: bpe * (
G + G + G + G * TILE + G # m_loc + m_new + corr + P + l_loc
),
engine_seq=(
EngineOp("MATH", "rmax", src="s", dst="m_loc", reduce_axis=-1),
EngineOp("MATH", "max_elem", src_a="m", src_b="m_loc", dst="m_new"),
EngineOp("MATH", "exp_diff", src_a="m", src_b="m_new", dst="corr"),
EngineOp("MATH", "exp_diff", src_a="s", src_b="m_new", dst="P", bcast_axis=0),
EngineOp("MATH", "rsum", src="P", dst="l_loc", reduce_axis=-1),
EngineOp("MATH", "fma", src_a="l", src_b="corr", src_c="l_loc", dst="l"),
EngineOp("MATH", "mul_bcast", src_a="O", src_b="corr", dst="O", bcast_axis=1),
EngineOp("MATH", "copy", src="m_new", dst="m"),
),
)
```
TLContext 가 `tl.composite(prologue=[{"op": "softmax_merge", ...}],
op="gemm", ...)` 에서:
1. operand R/RW 를 `RECIPE_DESCRIPTORS["softmax_merge"]` 와 검증.
2. scratch (m_loc, m_new, corr, P, l_loc) 를 scratch_scope helper 로 할당 (ADR-0063).
3. primary output P 의 shape 을 `s.shape` 에서 derive (identity).
4. `engine_seq` 를 8 개의 평평한 MATH OpSpec 으로 펼침 (모든 주소/크기 채움).
각 OpSpec 의 `scope = KERNEL`.
5. **Auto-bind (충돌 검사, D6.6).** head GEMM 이 `operands["a"]`
*명시하지 않은* 경우에만 `operands["a"] = P_handle` auto-bind.
커널이 `a` 를 명시적으로 제공했다면 validation error — 모호한
primary-output 대체.
6. `CompositeCmd(ops=(8 MATH + 1 GEMM + epilogue), rw_handles=(m, l, O))` emit.
**RECIPE_DESCRIPTORS 는 HW 경로 어디에도 안 나타남**. PE_SCHEDULER 는
평평한 ops 리스트만 봄.
### D6. Invariants
1. **GEMM 수.** `0 ≤ count(op.kind == "gemm" in cmd.ops) ≤ 1`. GEMM 0 개면
MATH-only composite.
2. **softmax_merge 가 V operand 없음.** `RECIPE_DESCRIPTORS[
"softmax_merge"].operands` 에 HBM-resident handle 없음 → prologue 동안
DMA 삽입 없음 → V DMA 는 GEMM head 시작 시에만 발생 → decode opt2 의
#1 / #2 간 K-before-V DMA priority 자연 강제.
3. **Cross-composite strict FIFO.** PE_SCHEDULER 가 in-flight `rw_handles`
추적. 신규 composite 의 `rw_handles` (또는 op 입력) 가 in-flight 와
교차하면 모든 이전 composite 완료까지 대기 — out-of-order reorder 없음.
4. **legacy caller 후방 호환.** 사용자 API `tl.composite(op="gemm", a, b,
epilogue=[...])` 보존. TLContext 가 내부적으로 평평한 ops `CompositeCmd`
로 lowering.
5. **PE_SCHEDULER recipe-free.** PE_SCHEDULER 는 RECIPE_DESCRIPTORS import
안 함, "softmax_merge" 의 존재 모름, recipe 별 코드 분기 없음.
6. **암묵 operand 대체 없음 (auto-bind 충돌).** prologue recipe 가
`primary_out` 선언하고 head op 의 auto-bind 대상 operand (예: GEMM
`a`) 도 *커널이 명시적* 으로 제공하면, TLContext 가 emit 시 validation
error. 커널은 recipe 가 바인딩하게 두거나 (operand 생략) **혹은**
명시적으로 지정 (prologue 생략, 또는 `primary_out` 없는 recipe 사용)
해야 함. "어느 값이 이기느냐" 의 모호함 방지.
7. **Prologue MATH operand TCM-only.** D4 의 재기술: *prologue recipe*
(비 GEMM) OpSpec 의 모든 operand 는 TCM 상주. head op (gemm 또는 math)
은 예외 — HBM 스트리밍 허용. Phase 1 에서 TLContext emit 시 강제.
### D7. Boundary 요약 (compiler vs scheduler vs engine)
```
HOST DEVICE
TLContext (compiler analog) PE_SCHEDULER (scheduler/dispatcher)
- RECIPE_DESC 와 검증 - 평평한 ops 리스트 스캔
- primary_out shape derive - GEMM 식별 (≤ 1)
- scratch 할당 - position-기반 stage 배치
- recipe → 평평한 OpSpec 펼침 - space 보고 DMA stage 자동 삽입
- rw_handles 계산 - strict-FIFO RW 해저드 추적
- CompositeCmd emit - Stage 리스트 emit
imports: RECIPE_DESCRIPTORS imports: recipe 관련 없음
ENGINES (PE_MATH / PE_GEMM / PE_DMA)
- Stage.params 읽음 (op_kind, addresses, n_elements)
- 기존 latency 모델 변화 없음
imports: recipe 관련 없음
```
## Alternatives
### A1. tile 당 3 composite (prologue 개념 없이)
#2 를 `softmax_merge` + `gemm(P·V)` + `add(O)` 의 3 개로 분할. 장점:
flat-ops 재구조화 없음 (`add` 가 기존 epilogue 경로 재사용). 단점: tile
당 dispatch 3 회 (대신 2) — ADR-0060 §8 item 4 sizing ("**두 개** 로
분할") 과 어긋남; ~32 ns × N_tiles 의 fixed-cost 절감 소실. **기각.**
### A2. Mixed-engine recipe (`flash_pv_merge`)
MATH + GEMM + MATH 흡수하는 단일 recipe (ADR-0060 §5.6 원안 표현). 장점:
가장 적은 dispatch (1 composite). 단점: PE_SCHEDULER 가 engine-경계
recipe 알아야 함, "엔진은 Stage.op_kind 만" boundary 위반, recipe 표가
ISA-급이 되어 HW 인터페이스 오염. **기각.**
### A3. Generic op-list DAG (`tl.ex_composite([...])`)
사용자 정의 임의 op 시퀀스 + 명명 슬롯 + 데이터플로 분석. 장점: 최대
유연성. 단점: 단일 user (softmax_merge) 가 generic DAG 언어를 정당화 못함;
PE_SCHEDULER 에 데이터플로 분석 필요; ADR-0060 §8 item 4 의 "inner loop
흡수" 트랩 재등장. **기각 (Simplicity First).**
### A4. RW-aware scheduler reorder (strict FIFO 가 아님)
PE_SCHEDULER 가 `rw_handles` 교차 없으면 후속 composite 가 in-flight 를
추월 허용 (예: 다음 tile 의 #1 가 현재 tile 의 #2 와 겹침 — #1 은
m/l/O 안 만짐). 장점: 더 많은 overlap. 단점: scheduler state 증가, 이득
한정 (다음 #1 은 어차피 K DMA 대기). **연기** — strict FIFO 먼저;
RW-aware reorder 는 후속 ADR 가능.
### A5. softmax_merge 를 #1 (Q·Kᵀ) 의 epilogue 로 embedding
오답: softmax_merge 의 출력 `P` 는 #2 의 P·V GEMM 의 입력, #1 의
epilogue *이후* 실행. embedding 하면 #2 가 자기 입력을 #1 의 epilogue
scratch 에서 읽는 형태 — 순환 tile-loop 의존성. **기각 (incorrect).**
## Consequences
### Positive
- ADR-0060 §5.6 / §8 item 4 carve-out 정확 구현: opt2 = tile 당 2
composite, K-before-V DMA priority 자연.
- composite contract 균일화: scope-기반 배치의 평평한 ops 리스트. HW cmd
에서 "head" 와 "epilogue" 구조 구분 없음.
- DMA 가 operand `space` 에서 자동 추론; 커널 표면 단순화.
- HW 인터페이스 (`CompositeCmd` 모양) 가 recipe 증가에도 작게 유지 —
recipe 펼침이 host-side.
- 새 fused 패턴 (linear+gelu+norm, rmsnorm+linear 등) 은 `RECIPE_DESCRIPTORS`
entry 추가만으로 가능; PE_SCHEDULER 변경 없음.
### Negative
- CompositeCmd 구조 변경 — 모든 현재 caller 의 refactor. 의미 보존
(D6.4); ADR-0064 Revision 2 land 후 기존 골든 불변.
- `OpSpec.operands` 가 positional `tuple[Any, ...]` 에서
`dict[str, TensorHandle]` 로 변경 — 기존 epilogue lowering 접촉.
- 새 TLContext 코드: recipe 펼침 + scratch slot 할당 + primary_out
auto-bind. ~80120 LOC.
- PE_SCHEDULER 의 `_generate_plan` 에 position-기반 스캔 + DMA 자동 삽입
+ strict-FIFO RW tracker. ~60100 LOC.
- **strict-FIFO RW 해저드 추적이 안전하지만 보수적.** 두 composite 가
`rw_handle` 을 공유할 때 신규는 모든 이전 겹치는 composite 가 *완전히
완료* 될 때까지 대기 — 원칙적으로 overlap 가능한 경우에도
(예: 다음-tile #1 = Q·Kᵀ 가 m/l/O 안 만지므로 현재-tile 의 #2 와
overlap 가능하지만, 같은 핸들을 만지는 후속 path 도 FIFO 로 직렬화).
더 똑똑한 scheduler 대비 **composite-level overlap 을 과소 노출**.
RW-aware reorder (A4) 가 연기된 개선.
## Open review items
1. **identity 를 넘는 recipe shape derivation 룰.** 미래 recipe 가
non-identity transform (예: 전치된 primary_out) 필요할 수 있음.
`PrimaryOutSpec.transform` 필드가 forward-compat; 필요 시 새 transform 추가.
2. **Recipe scratch allocator** 의 ADR-0063 `scratch_scope` 통합. 초기
배선: TLContext 가 활성 `scratch_scope` 가 있으면 그 안에서 할당,
없으면 kernel-scope persistent slot.
3. **`tl_recipes.py` 위치.** `triton_emu/` 아래 새 모듈. recipe 지식을
compiler analog (TLContext) 와 함께 위치.
4. **GEMM `out=TensorHandle`.** 기존 `out_ptr: int` 와 함께 새 형태.
둘 다 허용; handle 형태가 신규 코드에 권장 (handle identity 로
strict-FIFO RW 추적 가능).
## Test Requirements
1. **CompositeCmd 평평화 refactor — 의미 보존.** 기존 모든 bench 의
op_log 가 refactor 전후로 byte-equal (legacy API `tl.composite(
op="gemm", a, b, epilogue=[...])` 가 같은 Stage 시퀀스로 lowering).
2. **softmax_merge recipe lowering.** TLContext 호출 `tl.composite(
prologue=[{"op": "softmax_merge", "s": Sj, "m": m, "l": l, "O": O}],
op="gemm", b=V_ref, out=O, epilogue=[{"op": "add", "other": O}])` 가
정확히 10 ops (8 MATH + 1 GEMM + 1 MATH(add)) + `rw_handles == (m, l, O)`
의 `CompositeCmd` 생성.
3. **K-before-V DMA priority invariant.** decode opt2 의 op_log 에서
#2 의 MATH prologue 동안 V 관련 DMA 없음 (#2 의 V DMA 는 GEMM head
시작 시에만).
4. **Strict-FIFO RW 직렬화.** `rw_handle` 공유하는 두 연속 composite 가
dispatch 순서대로 완료; stage 가 interleave 안 함.
5. **`pinned` 에서 DMA 자동 삽입.** not-`pinned` operand (예: `b=tl.ref(...)`)
의 GEMM composite 가 DMA_READ Stage emit; `pinned` operand (예: recipe
의 TCM scratch / primary-out, 또는 `tl.load` 결과) 는 emit 안 함.
(as-built D4 노트: DMA 결정은 `space` 태그가 아니라 기존 `pinned`
플래그 사용.)
6. **opt2 가 opt3 와 수치 동등.** data mode 에서 opt2 의 최종
`(m, l, O)` 가 opt3 와 fp tolerance 안.
7. **opt2 dispatch ratio (ADR-0064 Rev2 이후) — robust.** opt3 vs opt2
per-tile PE_CPU dispatch cycles 가 `opt3 > 2 × opt2` 만족
(정성적 gate, calibration-독립). Default-calibration 모델 기대치는
≈ 4.0× (FIXED=40 cycles, R=0.0625 cycles/byte) — DDD-0065 §11 의
모델-유도 숫자 참조; 테스트 gate 는 느슨한 `> 2×` 한계라
calibration 변경에도 살아남음. Ratio 가 FIXED-dominated —
command-count 감소가 1차 신호임을 반영.
8. **GEMM-count invariant.** GEMM OpSpec 두 개 가진 composite 가
TLContext emit 시 validation error.
9. **Composite 크기 cap 안에 (ADR-0064 D7).** Decode opt2 의 `#2`
composite (10 ops, ~322 logical bytes — operand 핸들의 per-op 합산)
가 default `MAX_COMPOSITE_LOGICAL_BYTES=1024` 안에 편안히 —
segmentation 없음. recipe lowering 테스트가 `cmd.logical_bytes
< 1024` 단언.
10. **Auto-bind 충돌.** `tl.composite(op="gemm", a=A_handle,
prologue=[{"op": "softmax_merge", ...}], ...)` 처럼 softmax_merge 가
`primary_out` 선언하는데 `a` 도 명시되면 emit 시 validation error
(D6.6).
11. **MATH operand TCM invariant.** `tl.composite(op="math",
a=tl.ref(addr, shape), ...)` (HBM-resident 핸들을 MATH operand 로
전달) 가 emit 시 validation error (D6.7).
## Dependencies
- **ADR-0060** §5.6, §8 item 4 — 본 ADR 이 구현하는 carve-out.
- **ADR-0064 Revision 2** — dispatch cost model; 먼저 land, opt2 의
fewer-issues win 을 측정 가능하게 함.
- **ADR-0063** `scratch_scope` — recipe intermediate 가 안에 할당;
`m, l, O` 는 persistent arena 로 밖에 할당 (DDD-0060 §6.2 참조).
- **ADR-0046** `tl_context` contract — `prologue=[...]` kwarg,
`out=TensorHandle`, RECIPE_DESCRIPTORS lookup 으로 확장.
- **ADR-0042** tile plan generators — (재설계 아닌) 확장 — 평평한 ops
리스트 처리 + DMA 자동 삽입.
- **ADR-0014** PE pipeline — boundary 보존: scheduler 가 recipe-free,
엔진이 op_kind-opaque.
## Migration
> **구현 순서 노트.** 아래의 ADR-0064-우선 순서가 원래 계획이었음.
> 실제 구현은 **1↔2 단계를 뒤집어** — ADR-0065 Phase 1 (평평화
> `CompositeCmd`) 이 ADR-0064 Rev2 *보다 먼저* land. 이유: `logical_bytes`
> (ADR-0064 D2) 가 평평화 형태 위에서 자연스럽게 정의되므로, refactor 를
> 먼저 하면 버려지는 transitional `logical_bytes` 를 피함. 각 단계는
> 독립적으로 안전: P1 은 기존 cost 경로 하에서 의미 보존 refactor
> (op_log byte-equal); ADR-0064 Rev2 가 이미 평평화된 형태 위에 구조적
> 공식을 교체. 최종 상태는 어느 쪽이든 동일.
Land 순서 (원래 계획; as-built 1↔2 뒤집힘은 위 노트 참조):
1. **ADR-0064 Revision 2** (별도 PR): 구조적 dispatch cost +
`logical_bytes` property + topology config override + 일회성 골든 재생성.
2. **ADR-0065 Phase 1** (본 ADR, PR a): `CompositeCmd` 평평화 refactor +
legacy-API lowering. refactor only; 골든 불변.
3. **ADR-0065 Phase 2** (본 ADR, PR b): `tl_recipes.py` + `softmax_merge`
recipe + PE_SCHEDULER position-스캔 + DMA 자동 삽입 + strict-FIFO
RW tracker.
4. **ADR-0065 Phase 3** (본 ADR, PR c): `_gqa_decode_long.py` opt2
variant + 수치 동등 테스트 + opt3 대비 dispatch-ratio 측정.
상세 구현 계획: **DDD-0065** 참조.