# Bài 10 — Đọc biến thể bằng recipe và bằng chứng

**Đích học:** nhận diện một model từ luồng target/loss/update; phân biệt phương pháp có thật, implementation có thật, checkpoint có thật và kết quả có thể tái lập; không suy forensic utility từ tên JEPA hoặc benchmark ASR. Cần bài [2](02-TAXONOMY-TARGET-LOSS.md), [6](06-JEPA-TRAINING-STEP.md), [9](09-ENCODER-ERROR-SCORING.md).

## 1. “JEPA” chưa điền đủ một recipe

Với một paper mới, điền bảy ô trước khi đọc bảng thắng/thua: **student thấy gì; target chứa gì; target tính từ đâu; loss tính ở đâu; gradient đi đâu; tham số target cập nhật ra sao; feature nào giữ downstream**. Hai paper cùng tên họ có thể khác cả target, loss và teacher.

| Nguồn primary | Student / target | Loss và cập nhật | Bài học phân biệt |
|---|---|---|---|
| [I-JEPA, Assran et al.](https://arxiv.org/html/2301.08243v3) | Visible image context → predictor; full-image EMA encoder → target vectors | Paper squared L2; code target LayerNorm + Smooth-L1, gradient context/predictor, teacher EMA | Full-view contextualized target khác patch-local target |
| [V-JEPA, Bardes et al.](https://arxiv.org/html/2404.08471v1) | Video context → predictor; full-video EMA features | Paper **L1** trên target features; teacher stop-gradient/EMA | Ngay họ continuous JEPA cũng không có một loss MSE duy nhất |
| [A-JEPA, Fei et al.](https://arxiv.org/html/2311.15830v3) | Audio context → predictor; EMA target vectors | Latent L2; curriculum block mask → time-frequency mask | Paper riêng với Audio-JEPA của Tuncay; curriculum là geometry chứ không chỉ ratio |
| [Audio-JEPA, Tuncay et al.](https://arxiv.org/pdf/2507.02915v2) | Visible log-mel patches → predictor; full-spectrogram EMA features | Paper squared L2; released code target standardization + unit-normalized loss | Cần ghi paper/config/checkpoint riêng, xem bài 6 |
| [GMM-Anchored JEPA](https://arxiv.org/html/2602.09040v1) | Student masked/augmented speech; clean EMA latent target **và** frozen log-mel GMM posterior | Continuous regression + auxiliary KL anchor giảm trọng số 1→.01 | Hybrid hai loại target; GMM này không cập nhật online |
| [S-JEPA](https://arxiv.org/html/2606.19398v1) | Masked speech → predictor/head; GMM posterior | **Single KL**, pha 1 frozen MFCC GMM; pha 2 online GMM nhận EMA features | EMA cung cấp features cho clustering; không phải teacher vector để regression |
| [GLaS-JEPA](https://arxiv.org/html/2609.37798v1) | Full/masked views qua **cùng encoder hiện hành** + token-wise projector | Masked MSE với stop-gradient target + full-view SIGReg; không EMA teacher riêng | Loss prediction và regularization có gradient paths khác nhau |
| [BEST-RQ-2](https://arxiv.org/html/2606.30700v1) | Visible-only ViT → predictor/classifier; fixed random-quantizer patch IDs | Masked CE; gradient encoder/predictor/classifier; R/codebook fixed | Decomposition context/predictor không bắt buộc continuous target hoặc teacher EMA |

Đọc phương pháp ở các phiên bản được dẫn, không ghép tên giống nhau thành cùng model. Bảng là bản đồ; [hồ sơ recipe](14-HO-SO-RECIPE.md) ghi dữ liệu, head giữ lại, normalization và giới hạn nguồn.

## 2. Hai loại GMM target phải tách riêng

GMM cho posterior $q_k(m)=\pi_k\mathcal N(m;\mu_k,\Sigma_k)/\sum_l\pi_l\mathcal N(m;\mu_l,\Sigma_l)$. Một điểm gần ranh giới có thể có $q=(.5,.5)$ thay vì hard ID. Với logits student $p$, KL $q\|p$ hoặc soft CE có cùng gradient nếu $q$ cố định; chúng khác entropy hằng số của target, xem bài 2.

**GMM-Anchored:** một GMM fit trên log-mel rồi giữ cố định, posterior là anchor bổ sung cho latent JEPA loss. Trọng số anchor giảm trong training. Target acoustic không tự là phoneme hay forensic cue. [Method §§3.1–3.4](https://arxiv.org/html/2602.09040v1#S3).

**S-JEPA:** pha 1 GMM MFCC 39 chiều, K=100 cố định; pha 2 K=500 trên EMA encoder features, cập nhật GMM từ minibatch sufficient statistics và chọn layer theo effective rank. Predictor/head match posterior bằng KL, không có latent-regression term thứ hai. EMA decay pha 2 luân phiên .999/.9999. Set positions không bất biến suốt training: pha 1 dùng masked + visible; pha 2 chuyển từ cấu hình đó sang masked-only và tắt augmentation. Phần architecture viết visible logits không supervised ở mô tả masked-only; khi ghi recipe phải kèm stage ở §3.3 và Appendix C. [S-JEPA §§3.1–3.4](https://arxiv.org/html/2606.19398v1#S3).

Điểm dễ nhầm: “không pretrained teacher distillation” không đồng nghĩa không có EMA encoder. S-JEPA vẫn dùng EMA features để tạo vocabulary online; chỉ không hồi quy trực tiếp features đó hoặc kế thừa một teacher pretrained ngoài pipeline.

## 3. GLaS-JEPA: stop-gradient không chặn regularizer

Full path tạo $z=p_\alpha(f_\theta(h))$, masked path tạo $\hat z=p_\alpha(f_\theta(m(h)))$. Projector tuyến tính từng token có 128 chiều, không trộn các vị trí; encoder làm contextual prediction, không temporal predictor riêng. Objective paper:

$$L=(1-\lambda)L_{pred}+\lambda L_{SIGReg},\quad\lambda=.01.$$

$L_{pred}$ ở masked positions dùng $\operatorname{sg}(z)$ nên gradient qua masked path. SIGReg dùng full-view $z$ **không detach** để regularize distribution; gradient qua full path. Hai paths cùng tham số nên cùng encoder nhận hai contributions. “Không teacher EMA” không phải “không target”: target vẫn là full-view current representation. Cụm “without engineered targets” cần hiểu theo định nghĩa paper. [GLaS-JEPA §§3.1–3.3](https://arxiv.org/html/2609.37798v1#S3).

Population SIGReg có thể là fixed-time across utterances hoặc marginal across times/utterances. Thay population đổi ý nghĩa penalty, xem phản ví dụ position-only ở bài 8. Các giả định của lý thuyết Gaussian/identifiability được paper dẫn không tự đúng với speech; không dùng một theorem của bài khác để khẳng định forensic signal được bảo toàn.

## 4. BEST-RQ-2: architecture split với categorical target

Trace: patch input gốc → random projection/codebook cố định → ID; visible patches → ViT encoder; context features + mask tokens/vị trí → predictor → classifier logits; CE chỉ masked positions. Không teacher EMA, không learned quantizer. Predictor bị bỏ downstream. BEST-RQ gốc dùng in-place corruption trong sequence encoder; BEST-RQ-2 bỏ masked tokens khỏi context encoder rồi thêm mask tokens ở predictor. [BEST-RQ-2 §§2.1–2.6](https://arxiv.org/html/2606.30700v1#S2).

Paper có BEST-RQ (ViT) để đối chiếu architecture/patchification với decomposition hai bước. Đối chứng này hỗ trợ một câu hỏi cơ chế trong protocol của paper, chưa đủ chuyển kết luận sang spoof detection. Không nhập target IDs ngẫu nhiên vào cùng nhóm “speech units từ clustering” chỉ vì đều dùng CE.

## 5. Chuỗi bằng chứng có nhiều nấc

| Nấc | Điều đã xác nhận | Chưa xác nhận |
|---|---|---|
| arXiv metadata và abstract | Title, authors, version, claim tác giả nêu | Cơ chế chi tiết, implementation đúng |
| Method equations/algorithm | Target, mask, loss, updates được mô tả | Code/checkpoint thực hiện y hệt |
| Pinned code/config | Cách một revision wired functions/schedules | Revision đó tạo checkpoint hoặc reproduces table |
| Model card/file metadata | Resource tồn tại, file lịch sử ra sao | Exact training recipe, weights integrity hoặc score thực tế |
| Run manifest/log + artifact hash | Liên kết một run với config/checkpoint cụ thể | Kết quả robust ngoài protocol đó |
| Independent evaluation | Kết quả trong một protocol được chạy lại | Mọi domain/task khác |

Với Audio-JEPA, source code và later card có năm 2026; checkpoint file có commit tháng 5/2025. Không dùng latest config để gán recipe cho checkpoint cũ. Với GLaS-JEPA chưa xác minh official repo trong lượt này. Với BEST-RQ-2 paper liên kết resource nhưng chưa audit implementation/checkpoint. [Crosswalk và provenance](research/jepa-code-doi-chieu.md).

## 6. “Hơn baseline” trả lời câu hỏi nào?

Một checkpoint A pretrained trên general audio và B pretrained trên speech, khác architecture/frontend/compute, rồi cùng train downstream head: đo **hai hệ thống checkpoint + adaptation**. Nó không isolate ảnh hưởng objective A vs B. Giữ head giống nhau giảm một số confound, không làm các yếu tố pretraining tự bằng nhau.

Toy decomposition của một improvement quan sát $\Delta$:

$$\Delta=\Delta_{objective}+\Delta_{data}+\Delta_{architecture}+\Delta_{budget}+\Delta_{frontend}+\Delta_{interaction}.$$

Đây là ký hiệu để nhắc các yếu tố, **không** là mô hình additive đã được nhận dạng từ dữ liệu. Một tổng quan sát không cho từng term; interactions có thể không additive. Cần đối chứng tương ứng trước khi viết causal claim.

GLaS-JEPA tự ghi baselines khác architecture/budget, nên bảng contextual comparison không isolate SSL objective. S-JEPA đo ASR/emotion/slot-filling; Audio-JEPA/BEST-RQ-2 đo general-audio tasks. Những kết quả đó giúp biết loại information đọc được trong các protocol cụ thể, chưa là bằng chứng unseen-spoof robustness. [GLaS-JEPA §5.1](https://arxiv.org/html/2609.37798v1#S5), [S-JEPA §4](https://arxiv.org/html/2606.19398v1#S4).

## 7. Bài tập tổng hợp — nhận diện recipe không nhìn tên

1. Model có R/codebook fixed, visible-only encoder, predictor và CE masked. Nó có bắt buộc teacher EMA không? Thuộc nhóm target nào?
2. Model có EMA encoder nhưng head dự đoán GMM posterior bằng KL. Có thể gọi loss continuous latent regression không?
3. Một full-view path nhận zero gradient từ MSE nhưng nhận gradient từ SIGReg. Câu “full-view network không học” sai ở đâu?
4. Repo có checkpoint tháng 5/2025 và config thêm tháng 7/2026. Có thể khẳng định checkpoint dùng WD trong config mới không?
5. Encoder ASR tốt hơn trên LibriSpeech có đủ bằng chứng chọn nó cho unseen generator không?
6. Hai checkpoints có cùng downstream recipe nhưng khác pretraining data. Claim “JEPA objective tốt hơn MAE objective” cần giới hạn thế nào?

<details>
<summary>Đáp án và lý do</summary>

1. Không. Đây là categorical fixed-random target, architecture encoder–predictor split; tương ứng BEST-RQ-2 ở source đã đọc.
2. Không. EMA features làm input cho clustering; target optimization là distribution, loss KL.
3. Chặn một term không chặn term khác. Shared parameters vẫn nhận gradient full-view regularization.
4. Không. Cần run/config provenance cùng checkpoint; mốc file sau không chứng minh lịch sử run trước.
5. Không. ASR và forensic cue khác nhau, protocol generalization cũng khác.
6. Có thể nói checkpoint systems khác nhau dưới downstream protocol này; chưa isolate causal effect của objective.

</details>

**Bài đánh giá cuối lớp:** chọn một recipe trong [hồ sơ 14](14-HO-SO-RECIPE.md), che tên, tự vẽ forward/gradient/update và nêu một shortcut, một phản ví dụ cho claim downstream, một bằng chứng còn thiếu. Làm xong mới đánh dấu “đã hiểu”; việc tài liệu được viết xong chưa tính là người học đã học.
