a8d04750e6
ADR-0065 review fixes: - D3 semantics clarified: position determines *phase* (pre-loop / head / in-accumulation / per-output-tile / post-loop), scope determines *repetition*. KERNEL = "once per CompositeCmd invocation", not "once per kernel launch" — removes ambiguity around softmax_merge's KERNEL scope at different positions before/after GEMM. - D4 MATH-TCM invariant: Phase 1 limits DMA auto-insertion to GEMM operands and head outputs; MATH operand with space="hbm" raises validation error at TLContext emit. Catches kernel misuse early; a MATH-with-HBM path is future work (explicit tl.load + MATH composite). - D6 #6 NEW: auto-bind conflict — explicit GEMM `a=` + prologue primary_out simultaneously → validation error (prevents "which value wins" ambiguity). - D6 #7 NEW: MATH operand TCM-only restatement. - Consequences/Negative: strict-FIFO conservatism note — safe but may under-expose composite-level overlap when handle-sharing composites could in principle run in parallel. - Test Requirements #10, #11: auto-bind conflict + MATH-HBM validation error tests. ADR-0064 Revision 2 clarification: - D2 counting rule: per-op summation, no deduplication. Same handle appearing in multiple OpSpecs or in rw_handles is counted independently — matches HW reality (each descriptor field is a distinct address slot, rw_handles is separate metadata block). Example walked through (mul_bcast with O in src_a + out + rw_handles). DDD-0065 follow-on: - §3.2: new test file test_tl_composite_validation.py covers auto-bind conflict, MATH-TCM, GEMM-count invariants. - §4.3: phase vs repetition split aligned with ADR-0065 D3. - §6 lowering pipeline: step 5 adds auto-bind conflict check; step 8 NEW — operand space validation; step 9 (size cap); step 10 (emit). - §8 strict-FIFO conservatism note: explains why next-tile #1 cross-tile pipelining still works under FIFO (no rw overlap), and where the conservatism would hurt (future handle-sharing recipes). - §10 verification: invariant-guard tests (auto-bind conflict, MATH TCM, per-op summation accuracy). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
381 lines
19 KiB
Markdown
381 lines
19 KiB
Markdown
# ADR-0065: 평평한 ops `CompositeCmd` + 첫 stateful recipe (`softmax_merge`)
|
||
|
||
## Status
|
||
|
||
Proposed (verification gated on ADR-0064 Revision 2 land)
|
||
|
||
> **ADR-0060** (AHBM GQA Fused Attention) 의 보조 ADR. §5.6 / §8 item 4 의
|
||
> carve-out 을 정확히 구현: decode opt2 = (#1 기존 Q·Kᵀ GEMM composite) +
|
||
> (#2 softmax + P·V + online-softmax merge 를 담은 단일 composite). ADR-0060
|
||
> §5.6 의 "ex_composite" / "flash_pv_merge" 명칭은 *retire* — 본 ADR 이
|
||
> 정식 모양을 확정: 기존 `tl.composite` 입구를 그대로 사용하면서 두 가지
|
||
> 구조적 추가 — (a) 평평한 `ops` 튜플, (b) MATH micro-op 시퀀스로
|
||
> 펼쳐지는 첫 stateful recipe `softmax_merge`.
|
||
|
||
## Context
|
||
|
||
### ADR-0060 가 이 ADR 에 남긴 것
|
||
|
||
ADR-0060 §5.6 는 opt3 (software pipelining) 를 *지금 ship* 으로 권고,
|
||
opt2 는 — per-tile dispatch 적음, K-before-V DMA priority,
|
||
**ADR-0064 의 cost model 이후에만 측정 가능** — 후속으로 carve out.
|
||
|
||
§8 item 4 가 carve-out 의 sizing: "재방문 시 **두 개** composite 로 분할 —
|
||
`#1` = Q·Kᵀ (기존 composite + `scale`), `#2` = softmax + P·V + 온라인-softmax
|
||
누산기 merge". 본 ADR 이 `#2` 를 구현.
|
||
|
||
### composite 가 구조 변경이 필요한 이유
|
||
|
||
현재 `CompositeCmd` 는 `(op, a, b, out_addr, ops=(head, *epi))` 모양 + 암묵 컨벤션:
|
||
- `op` 가 head 엔진 선택 (gemm | math)
|
||
- `ops[0]` 는 head OpSpec, `ops[1:]` 는 epilogue OpSpec 으로 head *이후*
|
||
per-output-tile 또는 per-K-tile 에 실행 (scope 결정)
|
||
|
||
decode opt2 는 MATH op 이 head GEMM **이전** 에 실행 (`m, l, O` 의
|
||
online-softmax 갱신 + GEMM 입력 `P` emit) 이 필요. 현 모양엔 자리 없음.
|
||
|
||
이전에 검토 후 기각된 경로:
|
||
- softmax_merge 를 #1 (Q·Kᵀ) 의 *epilogue* 로 두기 — `Sj` 를 읽고 *다음*
|
||
GEMM (#2 P·V) 가 소비할 `P` 를 쓰는 구조라 순환.
|
||
- 단일 mixed-engine recipe `flash_pv_merge` 가 MATH + GEMM + MATH 흡수 —
|
||
engine boundary 위반, PE_SCHEDULER 복잡화.
|
||
|
||
깔끔한 모양: **command 포맷에서 head/epilogue 구분 제거**. `ops` 를 평평한
|
||
순서 튜플로, 각 op 의 `scope` + position 이 tile-loop 배치 결정. "prologue"
|
||
개념은 *사용자 API* 의 `tl.composite(...)` 에 ergonomic 으로 남되 —
|
||
**HW command 는 보지 않음**.
|
||
|
||
### recipe 가 (generic op list 가 아닌) 이유
|
||
|
||
decode opt2 의 `#2` MATH chain 은 8 단계 고정 구조 (online softmax merge).
|
||
커널 작성자에게 8 단계 직접 와이어링은 verbose + correctness hazard
|
||
(단계 순서 의존). *Recipe* — TLContext 가 8 단계로 (주소 미리 채워)
|
||
펼치는 단일 명명 op (`softmax_merge`) — 가 커널을 간결히 + 구조를 단일
|
||
source 에 보존. recipe 표는 TLContext (compiler analog) 에 위치;
|
||
PE_SCHEDULER 는 안 읽음. D5 boundary 참조.
|
||
|
||
## Decision
|
||
|
||
### D1. `CompositeCmd` 는 평평한 ordered op 리스트
|
||
|
||
```python
|
||
@dataclass(frozen=True)
|
||
class CompositeCmd:
|
||
completion: CompletionHandle
|
||
ops: tuple[OpSpec, ...] # ordered MATH/GEMM ops
|
||
rw_handles: tuple[TensorHandle, ...] = () # cross-composite hazard
|
||
data_op: bool = True
|
||
|
||
@property
|
||
def logical_bytes(self) -> int: ... # ADR-0064 D2
|
||
```
|
||
|
||
이전의 `(op, a, b, out_addr, out_nbytes, math_op)` 필드는 `ops` 에 흡수.
|
||
`ops` 의 첫 OpSpec 이 `kind == "gemm"` 이면 head GEMM; 앞의 OpSpec 들이
|
||
pre-loop, 뒤의 OpSpec 들이 `scope` 에 따라 post-loop / per-tile epilogue.
|
||
|
||
### D2. `OpSpec` 이 named operand 보유
|
||
|
||
```python
|
||
@dataclass(frozen=True)
|
||
class OpSpec:
|
||
kind: str # "gemm" | "rmax" | ...
|
||
scope: Scope # KERNEL | K_TILE | OUTPUT_TILE
|
||
operands: dict[str, TensorHandle] # named — D5 참조
|
||
extra: dict[str, Any] # scalars, axes, m/k/n 등
|
||
out: TensorHandle | None = None # 명시적 write-back handle
|
||
|
||
@property
|
||
def logical_bytes(self) -> int: ...
|
||
```
|
||
|
||
named operand 로 PE_SCHEDULER 가 (positional 컨벤션 없이) 엔진 port 에
|
||
입력 라우팅 (예: GEMM 은 `operands["a"]`, `operands["b"]`).
|
||
|
||
### D3. Position + scope 가 tile-loop 배치 결정
|
||
|
||
**Semantics 분리.** *Position* 이 **phase** (tile-loop 수명주기 어디서
|
||
실행되는지) 결정; *scope* 가 그 phase 안에서의 **repetition** 결정.
|
||
`Scope.KERNEL` 의미는 "**`CompositeCmd` invocation 당 1 회**" — kernel
|
||
launch 당 1 회 아님. 그 1 회 실행이 op 의 시퀀스 내 *position* 에서
|
||
발생. 같은 KERNEL 값이 OpSpec 이 GEMM op 의 앞/위치/뒤에 있느냐에 따라
|
||
*phase* 가 달라짐.
|
||
|
||
PE_SCHEDULER 가 `cmd.ops` 에서 GEMM op 검색 (composite 당 ≤ 1, D6 참조).
|
||
인덱스 `g`:
|
||
|
||
| OpSpec position | scope | Phase | plan 내 배치 |
|
||
|---|---|---|---|
|
||
| `0 .. g-1` | KERNEL | pre-tile-loop | tile loop 진입 전 1 회 |
|
||
| `g` | KERNEL | head GEMM | `extra["m"], extra["k"], extra["n"]` 로 tile loop drive |
|
||
| `g+1 .. ` | K_TILE | in-accumulation | per K-tile epilogue (K 누산 loop 안) |
|
||
| `g+1 .. ` | OUTPUT_TILE | per-output-tile | per (m, n) tile epilogue (K 누산 후) |
|
||
| `g+1 .. ` | KERNEL | post-tile-loop | tile loop 종료 후 1 회 |
|
||
|
||
GEMM op 가 없는 composite (예: MATH-only) 는 모든 op 를 KERNEL-scope
|
||
직렬로 처리 — phase 가 "순서대로 single-shot" 으로 collapse.
|
||
|
||
### D4. PE_SCHEDULER 가 operand `space` 보고 DMA 자동 삽입
|
||
|
||
PE_SCHEDULER 가 GEMM 의 `operands` 와 `out` 검사. 각각:
|
||
- `space == "hbm"` → FETCH/GEMM Stage 앞에 DMA_READ Stage 삽입 (`out` 은 DMA_WRITE)
|
||
- `space == "tcm"` → DMA Stage 없음, in-place 소비
|
||
|
||
오늘 이미 `a_pinned`/`b_pinned` 로 부분 모델링; D4 가 룰을 균일화 +
|
||
가시화: **커널은 composite operand 의 명시적 DMA cmd 를 절대 emit 안 함;
|
||
PE_SCHEDULER 가 핸들에서 추론**. `tl.ref(addr, shape)` → `space="hbm"`;
|
||
`tl.zeros` / `tl.full` / `tl.load` 결과 → `space="tcm"`.
|
||
|
||
**비 GEMM op 은 TCM-only (Phase 1 한계, TLContext emit 시 강제).**
|
||
모든 non-GEMM (MATH) OpSpec 의 모든 operand 는 **반드시**
|
||
`space == "tcm"`. 커널이 `space == "hbm"` 핸들을 MATH operand 로 전달하면
|
||
TLContext 가 emit 시 validation error. Phase 1 에서 DMA 자동 삽입을
|
||
GEMM operand + head 출력 핸들로 제한; MATH-with-HBM 경로는 미래 확장
|
||
(명시적 `tl.load` + MATH composite 분리, 또는 새 "DMA-as-prologue" recipe
|
||
variant). 이 invariant 는 D6 #7 에 재기술.
|
||
|
||
### D5. RECIPE_DESCRIPTORS — TLContext 내부, `pe_commands.py` 가 아님
|
||
|
||
Recipe 는 새 TLContext-adjacent 모듈
|
||
(`src/kernbench/triton_emu/tl_recipes.py`) 에 위치. PE_SCHEDULER 는
|
||
import **안 함**. 첫 recipe:
|
||
|
||
```python
|
||
@dataclass(frozen=True)
|
||
class RecipeDescriptor:
|
||
operands: dict[str, str] # name → "R" | "RW"
|
||
primary_out: PrimaryOutSpec | None # 암묵 primary output 의 type/shape 룰
|
||
tile_alignment: Literal["single_shot", "tile_aligned"]
|
||
internal_scratch_bytes_fn: Callable[..., int]
|
||
engine_seq: tuple[EngineOp, ...] # TLContext 가 평평한 OpSpec 으로 펼침
|
||
|
||
|
||
RECIPE_DESCRIPTORS["softmax_merge"] = RecipeDescriptor(
|
||
operands={"s": "R", "m": "RW", "l": "RW", "O": "RW"},
|
||
primary_out=PrimaryOutSpec(
|
||
from_shape="s", from_dtype="s", transform="identity",
|
||
),
|
||
tile_alignment="single_shot",
|
||
internal_scratch_bytes_fn=lambda G, TILE, d, bpe: bpe * (
|
||
G + G + G + G * TILE + G # m_loc + m_new + corr + P + l_loc
|
||
),
|
||
engine_seq=(
|
||
EngineOp("MATH", "rmax", src="s", dst="m_loc", reduce_axis=-1),
|
||
EngineOp("MATH", "max_elem", src_a="m", src_b="m_loc", dst="m_new"),
|
||
EngineOp("MATH", "exp_diff", src_a="m", src_b="m_new", dst="corr"),
|
||
EngineOp("MATH", "exp_diff", src_a="s", src_b="m_new", dst="P", bcast_axis=0),
|
||
EngineOp("MATH", "rsum", src="P", dst="l_loc", reduce_axis=-1),
|
||
EngineOp("MATH", "fma", src_a="l", src_b="corr", src_c="l_loc", dst="l"),
|
||
EngineOp("MATH", "mul_bcast", src_a="O", src_b="corr", dst="O", bcast_axis=1),
|
||
EngineOp("MATH", "copy", src="m_new", dst="m"),
|
||
),
|
||
)
|
||
```
|
||
|
||
TLContext 가 `tl.composite(prologue=[{"op": "softmax_merge", ...}],
|
||
op="gemm", ...)` 에서:
|
||
1. operand R/RW 를 `RECIPE_DESCRIPTORS["softmax_merge"]` 와 검증.
|
||
2. scratch (m_loc, m_new, corr, P, l_loc) 를 scratch_scope helper 로 할당 (ADR-0063).
|
||
3. primary output P 의 shape 을 `s.shape` 에서 derive (identity).
|
||
4. `engine_seq` 를 8 개의 평평한 MATH OpSpec 으로 펼침 (모든 주소/크기 채움).
|
||
각 OpSpec 의 `scope = KERNEL`.
|
||
5. head GEMM 의 `operands["a"] = P_handle` auto-bind.
|
||
6. `CompositeCmd(ops=(8 MATH + 1 GEMM + epilogue), rw_handles=(m, l, O))` emit.
|
||
|
||
**RECIPE_DESCRIPTORS 는 HW 경로 어디에도 안 나타남**. PE_SCHEDULER 는
|
||
평평한 ops 리스트만 봄.
|
||
|
||
### D6. Invariants
|
||
|
||
1. **GEMM 수.** `0 ≤ count(op.kind == "gemm" in cmd.ops) ≤ 1`. GEMM 0 개면
|
||
MATH-only composite.
|
||
2. **softmax_merge 가 V operand 없음.** `RECIPE_DESCRIPTORS[
|
||
"softmax_merge"].operands` 에 HBM-resident handle 없음 → prologue 동안
|
||
DMA 삽입 없음 → V DMA 는 GEMM head 시작 시에만 발생 → decode opt2 의
|
||
#1 / #2 간 K-before-V DMA priority 자연 강제.
|
||
3. **Cross-composite strict FIFO.** PE_SCHEDULER 가 in-flight `rw_handles`
|
||
추적. 신규 composite 의 `rw_handles` (또는 op 입력) 가 in-flight 와
|
||
교차하면 모든 이전 composite 완료까지 대기 — out-of-order reorder 없음.
|
||
4. **legacy caller 후방 호환.** 사용자 API `tl.composite(op="gemm", a, b,
|
||
epilogue=[...])` 보존. TLContext 가 내부적으로 평평한 ops `CompositeCmd`
|
||
로 lowering.
|
||
5. **PE_SCHEDULER recipe-free.** PE_SCHEDULER 는 RECIPE_DESCRIPTORS import
|
||
안 함, "softmax_merge" 의 존재 모름, recipe 별 코드 분기 없음.
|
||
6. **암묵 operand 대체 없음 (auto-bind 충돌).** prologue recipe 가
|
||
`primary_out` 선언하고 head op 의 auto-bind 대상 operand (예: GEMM
|
||
`a`) 도 *커널이 명시적* 으로 제공하면, TLContext 가 emit 시 validation
|
||
error. 커널은 recipe 가 바인딩하게 두거나 (operand 생략) **혹은**
|
||
명시적으로 지정 (prologue 생략, 또는 `primary_out` 없는 recipe 사용)
|
||
해야 함. "어느 값이 이기느냐" 의 모호함 방지.
|
||
7. **MATH operand TCM-only.** D4 의 재기술: 모든 비 GEMM OpSpec 의 모든
|
||
operand 는 `space == "tcm"`. Phase 1 에서 TLContext emit 시 강제.
|
||
|
||
### D7. Boundary 요약 (compiler vs scheduler vs engine)
|
||
|
||
```
|
||
HOST DEVICE
|
||
TLContext (compiler analog) PE_SCHEDULER (scheduler/dispatcher)
|
||
- RECIPE_DESC 와 검증 - 평평한 ops 리스트 스캔
|
||
- primary_out shape derive - GEMM 식별 (≤ 1)
|
||
- scratch 할당 - position-기반 stage 배치
|
||
- recipe → 평평한 OpSpec 펼침 - space 보고 DMA stage 자동 삽입
|
||
- rw_handles 계산 - strict-FIFO RW 해저드 추적
|
||
- CompositeCmd emit - Stage 리스트 emit
|
||
imports: RECIPE_DESCRIPTORS imports: recipe 관련 없음
|
||
|
||
ENGINES (PE_MATH / PE_GEMM / PE_DMA)
|
||
- Stage.params 읽음 (op_kind, addresses, n_elements)
|
||
- 기존 latency 모델 변화 없음
|
||
imports: recipe 관련 없음
|
||
```
|
||
|
||
## Alternatives
|
||
|
||
### A1. tile 당 3 composite (prologue 개념 없이)
|
||
|
||
#2 를 `softmax_merge` + `gemm(P·V)` + `add(O)` 의 3 개로 분할. 장점:
|
||
flat-ops 재구조화 없음 (`add` 가 기존 epilogue 경로 재사용). 단점: tile
|
||
당 dispatch 3 회 (대신 2) — ADR-0060 §8 item 4 sizing ("**두 개** 로
|
||
분할") 과 어긋남; ~32 ns × N_tiles 의 fixed-cost 절감 소실. **기각.**
|
||
|
||
### A2. Mixed-engine recipe (`flash_pv_merge`)
|
||
|
||
MATH + GEMM + MATH 흡수하는 단일 recipe (ADR-0060 §5.6 원안 표현). 장점:
|
||
가장 적은 dispatch (1 composite). 단점: PE_SCHEDULER 가 engine-경계
|
||
recipe 알아야 함, "엔진은 Stage.op_kind 만" boundary 위반, recipe 표가
|
||
ISA-급이 되어 HW 인터페이스 오염. **기각.**
|
||
|
||
### A3. Generic op-list DAG (`tl.ex_composite([...])`)
|
||
|
||
사용자 정의 임의 op 시퀀스 + 명명 슬롯 + 데이터플로 분석. 장점: 최대
|
||
유연성. 단점: 단일 user (softmax_merge) 가 generic DAG 언어를 정당화 못함;
|
||
PE_SCHEDULER 에 데이터플로 분석 필요; ADR-0060 §8 item 4 의 "inner loop
|
||
흡수" 트랩 재등장. **기각 (Simplicity First).**
|
||
|
||
### A4. RW-aware scheduler reorder (strict FIFO 가 아님)
|
||
|
||
PE_SCHEDULER 가 `rw_handles` 교차 없으면 후속 composite 가 in-flight 를
|
||
추월 허용 (예: 다음 tile 의 #1 가 현재 tile 의 #2 와 겹침 — #1 은
|
||
m/l/O 안 만짐). 장점: 더 많은 overlap. 단점: scheduler state 증가, 이득
|
||
한정 (다음 #1 은 어차피 K DMA 대기). **연기** — strict FIFO 먼저;
|
||
RW-aware reorder 는 후속 ADR 가능.
|
||
|
||
### A5. softmax_merge 를 #1 (Q·Kᵀ) 의 epilogue 로 embedding
|
||
|
||
오답: softmax_merge 의 출력 `P` 는 #2 의 P·V GEMM 의 입력, #1 의
|
||
epilogue *이후* 실행. embedding 하면 #2 가 자기 입력을 #1 의 epilogue
|
||
scratch 에서 읽는 형태 — 순환 tile-loop 의존성. **기각 (incorrect).**
|
||
|
||
## Consequences
|
||
|
||
### Positive
|
||
- ADR-0060 §5.6 / §8 item 4 carve-out 정확 구현: opt2 = tile 당 2
|
||
composite, K-before-V DMA priority 자연.
|
||
- composite contract 균일화: scope-기반 배치의 평평한 ops 리스트. HW cmd
|
||
에서 "head" 와 "epilogue" 구조 구분 없음.
|
||
- DMA 가 operand `space` 에서 자동 추론; 커널 표면 단순화.
|
||
- HW 인터페이스 (`CompositeCmd` 모양) 가 recipe 증가에도 작게 유지 —
|
||
recipe 펼침이 host-side.
|
||
- 새 fused 패턴 (linear+gelu+norm, rmsnorm+linear 등) 은 `RECIPE_DESCRIPTORS`
|
||
entry 추가만으로 가능; PE_SCHEDULER 변경 없음.
|
||
|
||
### Negative
|
||
- CompositeCmd 구조 변경 — 모든 현재 caller 의 refactor. 의미 보존
|
||
(D6.4); ADR-0064 Revision 2 land 후 기존 골든 불변.
|
||
- `OpSpec.operands` 가 positional `tuple[Any, ...]` 에서
|
||
`dict[str, TensorHandle]` 로 변경 — 기존 epilogue lowering 접촉.
|
||
- 새 TLContext 코드: recipe 펼침 + scratch slot 할당 + primary_out
|
||
auto-bind. ~80–120 LOC.
|
||
- PE_SCHEDULER 의 `_generate_plan` 에 position-기반 스캔 + DMA 자동 삽입
|
||
+ strict-FIFO RW tracker. ~60–100 LOC.
|
||
- **strict-FIFO RW 해저드 추적이 안전하지만 보수적.** 두 composite 가
|
||
`rw_handle` 을 공유할 때 신규는 모든 이전 겹치는 composite 가 *완전히
|
||
완료* 될 때까지 대기 — 원칙적으로 overlap 가능한 경우에도
|
||
(예: 다음-tile #1 = Q·Kᵀ 가 m/l/O 안 만지므로 현재-tile 의 #2 와
|
||
overlap 가능하지만, 같은 핸들을 만지는 후속 path 도 FIFO 로 직렬화).
|
||
더 똑똑한 scheduler 대비 **composite-level overlap 을 과소 노출**.
|
||
RW-aware reorder (A4) 가 연기된 개선.
|
||
|
||
## Open review items
|
||
|
||
1. **identity 를 넘는 recipe shape derivation 룰.** 미래 recipe 가
|
||
non-identity transform (예: 전치된 primary_out) 필요할 수 있음.
|
||
`PrimaryOutSpec.transform` 필드가 forward-compat; 필요 시 새 transform 추가.
|
||
2. **Recipe scratch allocator** 의 ADR-0063 `scratch_scope` 통합. 초기
|
||
배선: TLContext 가 활성 `scratch_scope` 가 있으면 그 안에서 할당,
|
||
없으면 kernel-scope persistent slot.
|
||
3. **`tl_recipes.py` 위치.** `triton_emu/` 아래 새 모듈. recipe 지식을
|
||
compiler analog (TLContext) 와 함께 위치.
|
||
4. **GEMM `out=TensorHandle`.** 기존 `out_ptr: int` 와 함께 새 형태.
|
||
둘 다 허용; handle 형태가 신규 코드에 권장 (handle identity 로
|
||
strict-FIFO RW 추적 가능).
|
||
|
||
## Test Requirements
|
||
|
||
1. **CompositeCmd 평평화 refactor — 의미 보존.** 기존 모든 bench 의
|
||
op_log 가 refactor 전후로 byte-equal (legacy API `tl.composite(
|
||
op="gemm", a, b, epilogue=[...])` 가 같은 Stage 시퀀스로 lowering).
|
||
2. **softmax_merge recipe lowering.** TLContext 호출 `tl.composite(
|
||
prologue=[{"op": "softmax_merge", "s": Sj, "m": m, "l": l, "O": O}],
|
||
op="gemm", b=V_ref, out=O, epilogue=[{"op": "add", "other": O}])` 가
|
||
정확히 10 ops (8 MATH + 1 GEMM + 1 MATH(add)) + `rw_handles == (m, l, O)`
|
||
의 `CompositeCmd` 생성.
|
||
3. **K-before-V DMA priority invariant.** decode opt2 의 op_log 에서
|
||
#2 의 MATH prologue 동안 V 관련 DMA 없음 (#2 의 V DMA 는 GEMM head
|
||
시작 시에만).
|
||
4. **Strict-FIFO RW 직렬화.** `rw_handle` 공유하는 두 연속 composite 가
|
||
dispatch 순서대로 완료; stage 가 interleave 안 함.
|
||
5. **`space` 에서 DMA 자동 삽입.** `b=tl.ref(...)` (space=hbm) 의 GEMM
|
||
composite 가 DMA_READ Stage emit; `b` 가 `tl.zeros(...)` (space=tcm) 면
|
||
emit 안 함. 출력 핸들 `space=hbm` 이면 DMA_WRITE emit; `space=tcm` 이면
|
||
emit 안 함.
|
||
6. **opt2 가 opt3 와 수치 동등.** data mode 에서 opt2 의 최종
|
||
`(m, l, O)` 가 opt3 와 fp tolerance 안.
|
||
7. **opt2 dispatch ratio (ADR-0064 Rev2 이후).** opt3 vs opt2 per-tile
|
||
PE_CPU dispatch cycles ratio 가 default calibration (FIXED=40 cycles,
|
||
R=0.0625 cycles/byte) 에서 ≈ 4.0×. Ratio 가 FIXED-dominated —
|
||
command-count 감소가 1차 신호임을 반영.
|
||
8. **GEMM-count invariant.** GEMM OpSpec 두 개 가진 composite 가
|
||
TLContext emit 시 validation error.
|
||
9. **Composite 크기 cap 안에 (ADR-0064 D7).** Decode opt2 의 `#2`
|
||
composite (10 ops, ~322 logical bytes — operand 핸들의 per-op 합산)
|
||
가 default `MAX_COMPOSITE_LOGICAL_BYTES=1024` 안에 편안히 —
|
||
segmentation 없음. recipe lowering 테스트가 `cmd.logical_bytes
|
||
< 1024` 단언.
|
||
10. **Auto-bind 충돌.** `tl.composite(op="gemm", a=A_handle,
|
||
prologue=[{"op": "softmax_merge", ...}], ...)` 처럼 softmax_merge 가
|
||
`primary_out` 선언하는데 `a` 도 명시되면 emit 시 validation error
|
||
(D6.6).
|
||
11. **MATH operand TCM invariant.** `tl.composite(op="math",
|
||
a=tl.ref(addr, shape), ...)` (HBM-resident 핸들을 MATH operand 로
|
||
전달) 가 emit 시 validation error (D6.7).
|
||
|
||
## Dependencies
|
||
|
||
- **ADR-0060** §5.6, §8 item 4 — 본 ADR 이 구현하는 carve-out.
|
||
- **ADR-0064 Revision 2** — dispatch cost model; 먼저 land, opt2 의
|
||
fewer-issues win 을 측정 가능하게 함.
|
||
- **ADR-0063** `scratch_scope` — recipe intermediate 가 안에 할당;
|
||
`m, l, O` 는 persistent arena 로 밖에 할당 (DDD-0060 §6.2 참조).
|
||
- **ADR-0046** `tl_context` contract — `prologue=[...]` kwarg,
|
||
`out=TensorHandle`, RECIPE_DESCRIPTORS lookup 으로 확장.
|
||
- **ADR-0042** tile plan generators — (재설계 아닌) 확장 — 평평한 ops
|
||
리스트 처리 + DMA 자동 삽입.
|
||
- **ADR-0014** PE pipeline — boundary 보존: scheduler 가 recipe-free,
|
||
엔진이 op_kind-opaque.
|
||
|
||
## Migration
|
||
|
||
Land 순서:
|
||
1. **ADR-0064 Revision 2** (별도 PR): 구조적 dispatch cost +
|
||
`logical_bytes` property + topology config override + 일회성 골든 재생성.
|
||
2. **ADR-0065 Phase 1** (본 ADR, PR a): `CompositeCmd` 평평화 refactor +
|
||
legacy-API lowering. refactor only; 골든 불변.
|
||
3. **ADR-0065 Phase 2** (본 ADR, PR b): `tl_recipes.py` + `softmax_merge`
|
||
recipe + PE_SCHEDULER position-스캔 + DMA 자동 삽입 + strict-FIFO
|
||
RW tracker.
|
||
4. **ADR-0065 Phase 3** (본 ADR, PR c): `_gqa_decode_long.py` opt2
|
||
variant + 수치 동등 테스트 + opt3 대비 dispatch-ratio 측정.
|
||
|
||
상세 구현 계획: **DDD-0065** 참조.
|