# Hồ sơ 19 recipe — target, gradient, update và output

Đây là phiếu đối chiếu phương pháp, không phải leaderboard hoặc danh sách đề xuất model. **Data/budget** ghi corpus/regime đã đọc; số giờ unique không bằng tổng samples seen. Chỗ chưa chốt số bước/batch/compute được đánh dấu, không suy từ tên checkpoint. Ngoài I-JEPA/Audio-JEPA code được ghim dưới đây, chi tiết là **paper recipe**, chưa audit implementation hoặc load weights.

Mẫu tự điền khi đọc bài mới: student view → target content/source/contextualization → main/aux loss và positions → gradient recipients → target update → cơ chế tránh nghiệm vô nghĩa → output giữ lại → data/budget/source/unknown. Loss reduction theo vector norm khác mean theo coordinate; các bài 2–6 tính sự khác biệt.

## 1. Predictive speech SSL

### CPC

- **Student:** waveform → local convolutional encoder $z_t$ → autoregressive context $c_t$ chỉ từ quá khứ. Không cần masked input trong recipe audio gốc.
- **Target:** future continuous $z_{t+k}$ của cùng encoder; score bilinear theo horizon, đối chiếu với negatives. Future target là feature local, không full-input bidirectional teacher.
- **Loss:** InfoNCE chọn future positive giữa candidates, aggregate horizons/times; không coi đây là MSE reconstruction.
- **Gradient/update:** local encoder, autoregressive model và prediction score parameters nhận gradient, gồm feature dùng làm keys; không target teacher EMA riêng.
- **Nghiệm/giới hạn:** negatives khuyến khích discrimination; sampling quyết định signal có thể dùng. Equal scores cho loss $\log K$, chưa có định lý task relevance.
- **Output:** encoder/context features cho audio probes; không score genuine/spoof mặc định.
- **Data/budget/source:** LibriSpeech trong audio setup [CPC §§2–3.1](https://arxiv.org/pdf/1807.03748); paper còn vision/language/RL experiments. Không chốt một số giờ/budget cho mọi modality.

### wav2vec 2.0

- **Student:** CNN waveform latents, span mask trước context Transformer; full-length masked latent sequence, không visible-only patch encoder.
- **Target:** quantized **unmasked** CNN latents, codebooks và selection logits được học; grouped Gumbel-Softmax/straight-through. Target không phải offline frozen ID hoặc EMA teacher.
- **Loss:** masked contrastive CE; negatives lấy từ các masked timesteps khác trong cùng utterance ở recipe paper. Auxiliary diversity maximizes entropy của average code-use distribution; một số recipe có feature activation L2/gradient scaling.
- **Gradient/update:** CNN, context Transformer, target quantizer/selection/codebook có gradient qua differentiable/ST path; diversity tác động quantizer. Không suy “target nên detach” từ tên vai trò.
- **Nghiệm/giới hạn:** negatives và code diversity tạo incentives khác nhau; codebook perplexity cao chưa chứng minh acoustic forensic features tốt.
- **Output:** context features, sau đó task head/CTC hoặc adaptation; quantizer/loss không phải downstream spoof score.
- **Data/budget/source:** paper dùng các regimes LibriSpeech 960 h / Libri-Light khoảng 60k h, Base/Large và số updates khác nhau; không gán cùng budget cho cả hai. [wav2vec 2.0 §§2–4](https://arxiv.org/pdf/2006.11477v3).

### HuBERT

- **Student:** waveform → CNN features bị mask spans → contextual Transformer.
- **Target:** offline hard k-means IDs, đầu từ MFCC, lượt sau từ intermediate features của model trước. IDs frozen trong mỗi training iteration; target không cần trùng nhãn phoneme.
- **Loss:** paper định nghĩa masked/unmasked mixture; main setting dùng masked loss ($\alpha=1$). Projection và learned classification codeword embeddings tạo logits; không contrastive negatives cho main loss.
- **Gradient/update:** student encoder/Transformer/projection/classification embeddings nhận gradient; k-means/IDs không nhận gradient từ CE. Reclustering offline giữa iterations là update pipeline riêng, không EMA.
- **Nghiệm/giới hạn:** fixed nontrivial targets và contextual prediction; label consistency quan trọng. Constant/poor targets vẫn có thể tạo task vô ích.
- **Output:** speech encoder features downstream; offline clustering head không bắt buộc giữ.
- **Data/budget/source:** Base LibriSpeech 960 h, larger regimes Libri-Light khoảng 60k h; iterations/budget là phần recipe, không chỉ unique hours. [HuBERT §II](https://arxiv.org/pdf/2106.07447v1).

### WavLM

- **Student:** masked speech sequence có noise/utterance mixing ở subset samples, Transformer có gated relative-position bias.
- **Target:** offline units từ **primary/original speech**, không phải mixture waveform hoặc nhãn người nói thứ hai; clustering từ speech features theo setup.
- **Loss:** masked-unit classification/denoising CE. “Denoise” ở đây là recover unit target, không waveform MSE. Relative-position bias là architecture, không auxiliary loss chống collapse.
- **Gradient/update:** student CNN/Transformer/classifier nhận gradient; pseudo-label clustering fixed trong lượt. Không EMA teacher của kiểu data2vec.
- **Nghiệm/giới hạn:** target ổn định + denoising task tạo incentive giữ primary speech; không chứng minh invariance với mọi noise/channel hoặc spoof.
- **Output:** speech encoder/features cho nhiều tasks; task head riêng.
- **Data/budget/source:** Base 960 h LibriSpeech; Base+/Large 94k h = 60k Libri-Light +10k GigaSpeech +24k VoxPopuli theo paper. Không so objective mà bỏ khác biệt data/scale. [WavLM §IV và setup](https://arxiv.org/pdf/2110.13900v5).

### BEST-RQ

- **Student:** acoustic features bị thay bằng small Gaussian noise ở masked spans; sequence encoder giữ time structure.
- **Target:** random projection và random codebook fixed, normalize projected vector/codebook rồi nearest-code assignment của original features. Hard categorical ID, không learned speech unit.
- **Loss:** CE ở masked positions; không negatives contrastive, không learned-quantizer diversity loss bắt buộc trong paper recipe.
- **Gradient/update:** speech encoder/classifier nhận gradient; projection, codebook và argmin target không học qua CE; không EMA teacher.
- **Nghiệm/giới hạn:** fixed mapping tránh co-adaptation của target nhưng không bảo đảm semantic/balanced targets. Nếu target IDs đều một code, constant classification có thể thắng toy task.
- **Output:** pretrained sequence encoder; target quantizer/classifier không cần downstream.
- **Data/budget/source:** Libri-Light pretraining → LibriSpeech fine-tuning; paper còn multilingual/streaming regimes, không đồng nhất compute. [BEST-RQ §§3–4](https://proceedings.mlr.press/v162/chiu22a/chiu22a.pdf).

### BEST-RQ-2

- **Student:** 2D log-mel patches; **remove** masked patches khỏi ViT context encoder; predictor nhận visible features + mask token/positions, rồi classifier.
- **Target:** fixed random projection/codebook trên original patches, hard IDs. Paper setting K=8192, projection dim=16; không teacher features/EMA hoặc learned quantizer.
- **Loss:** CE chỉ masked patches, main categorical loss; không variance/covariance auxiliary được mô tả.
- **Gradient/update:** encoder, predictor, classifier và participating learned embeddings nhận gradient; R/codebook fixed. Token removal khác BEST-RQ original noise-in-place và one-stage BEST-RQ (ViT) comparator.
- **Nghiệm/giới hạn:** target diversity/predictability vẫn cần; split architecture không chứng minh tránh mọi collapse hoặc task transfer.
- **Output:** encoder final-block token features; predictor/head bỏ, downstream benchmark chọn pooling/probe.
- **Data/budget/source:** khoảng 1.9M filtered AudioSet 10 s clips; 200k updates, comparable compute reported, **không strictly matched examples processed**. [BEST-RQ-2 §§2–4](https://arxiv.org/html/2606.30700v1). Resource links trong paper đã ghi nhận, chưa audit code/load checkpoint.

## 2. Reconstruction và continuous contextual target

### AudioMAE

- **Student:** log-mel → patches; encoder chỉ visible tokens, decoder nhận visible encoded tokens và learned mask tokens/positions.
- **Target:** original spectrogram patch values, có tùy chọn per-patch normalization; target là input space, không teacher embedding. Decoder local-window attention là architecture riêng của audio method.
- **Loss:** masked patch reconstruction MSE; averaging coordinate/vector/reduction phải đọc implementation trước khi so absolute values.
- **Gradient/update:** encoder/decoder/learned embedding nhận gradient; input values không là learned target branch; không EMA teacher.
- **Nghiệm/giới hạn:** fixed input targets, bottleneck/masking; decoder conditional mean có thể smooth trong MSE. Điều đó chưa chứng minh encoder mất artifact, xem counterexample bài 4.
- **Output:** encoder features, decoder bỏ; fine-tune/probe tùy downstream task.
- **Data/budget/source:** AudioSet-scale pretraining; main random mask 80%, native 10 s/16 kHz, 1,024 time ×128 mel, patch 16×16 →512 tokens theo setup/Appendix B. [AudioMAE §§3–4](https://arxiv.org/html/2207.06405v3). Không audit training budget/checkpoint history ở đây.

### data2vec (bản gốc, speech)

- **Student:** masked speech features qua contextual Transformer.
- **Target:** full-input EMA Transformer, selected top layers được normalize rồi average; speech dùng per-sample/per-feature normalization theo sequence (instance norm), lấy FFN output trước last residual theo method. Feature/positional encoders được shared; không sao chép mọi tham số frontend thành teacher EMA độc lập.
- **Loss:** generic formulation Smooth-L1; **speech experimental setup dùng simple L2**, K=8 upper layers. Target full input contextualized khác quantized local targets.
- **Gradient/update:** masked student có gradient; target Transformer under no-grad, EMA từ student. Shared frontend có gradient từ student, không direct gradient từ target regression branch.
- **Nghiệm/giới hạn:** target normalization/EMA/contextualization là training recipe, không variance penalty hoặc định lý anti-collapse. Modality-specific normalization không được tráo cho nhau.
- **Output:** contextual encoder features; predictor/head/task adaptation theo modality; không decoder reconstruct waveform.
- **Data/budget/source:** speech setup LibriSpeech 960 h; EMA .999→.9999 trong 30k updates đầu. Đây là bản paper gốc, không nhập data2vec 2.0 efficiency recipe. [data2vec §§3–4.2](https://arxiv.org/pdf/2202.03555v3).

## 3. Matching views và collapse regularization

Các paper dưới đây có bằng chứng chính trên ảnh/ImageNet. Chúng dạy cơ chế chung; **không phải experiments deepfake audio**.

### BYOL

- **Student/target:** hai augmented views. Online encoder + projector + predictor match target encoder + projector của view kia; target normalized continuous vector. Target không có online predictor copy để regression theo cùng cách.
- **Loss/positions:** symmetric normalized squared distance giữa two-view pooled embeddings; không masked audio positions hoặc negatives.
- **Gradient/update:** gradient online encoder/projector/predictor; target detached; encoder/projector target EMA. Momentum tăng theo schedule, không một fixed tau cho cả run.
- **Nghiệm/giới hạn:** predictor, stop-gradient, normalization/teacher dynamics và augmentation trong recipe thực nghiệm; constant vectors vẫn là counterexample của objective alone.
- **Output/data/source:** online encoder bỏ projector/predictor khi evaluate; ImageNet setup, không chốt mọi ablation budget. [BYOL §3](https://arxiv.org/pdf/2006.07733v3).

### SimSiam

- **Student/target:** shared encoder/projector của hai augmented views; predictor online từng direction, target là view kia bị detach trong **term đó**.
- **Loss/positions:** symmetric negative cosine, pooled image features; không negatives hoặc masked targets.
- **Gradient/update:** encoder/projector/predictor cập nhật bằng gradient từ online paths của cả hai terms; không separate EMA target weights. Đừng nói “branch kia không bao giờ nhận gradient”: roles đảo trong symmetric term.
- **Nghiệm/giới hạn:** stop-gradient + predictor là asymmetry được ablate trong paper; không universal proof cho mọi loss/data.
- **Output/data/source:** encoder downstream, heads bỏ; ImageNet. [SimSiam §3/Algorithm 1](https://arxiv.org/pdf/2011.10566v1).

### VICReg

- **Student/target:** two augmented views qua shared encoder/expander; continuous paired vectors, không detached EMA teacher.
- **Loss:** alignment MSE giữa views + variance floor từng coordinate trên batch của mỗi view + covariance off-diagonal penalty mỗi view. Covariance dùng centered samples và $n-1$.
- **Gradient/update:** cả view paths/encoder/expander nhận gradient từ các terms; không teacher update riêng.
- **Nghiệm/giới hạn:** constant output bị variance penalty, covariance term alone không đủ; exact constant point có thể có zero gradient với smoothed std. Penalty value và khả năng rời stationary point là hai câu hỏi khác nhau.
- **Output/data/source:** encoder bỏ expander khi downstream; ImageNet, coefficients/budget tùy setup, không tự gán cho audio. [VICReg §4.1](https://arxiv.org/pdf/2105.04906v3).

### Barlow Twins

- **Student/target:** paired augmented views, shared encoder/projector, batch-normalized coordinates.
- **Loss:** squared diagonal error của cross-correlation $R$ so với 1, cộng $\lambda$ squared off-diagonals; mean cross-view alignment và redundancy khác VICReg covariance từng view.
- **Gradient/update:** cả paths và shared parameters học bằng gradient; không EMA/detached teacher riêng.
- **Nghiệm/giới hạn:** identity cross-correlation thúc đẩy non-redundant aligned features. At constant batch, ideal correlation denominator bằng 0; numerical epsilon không tạo task information.
- **Output/data/source:** encoder bỏ projector; ImageNet. [Barlow Twins §2.1](https://arxiv.org/pdf/2103.03230v3).

## 4. JEPA-family recipes

### I-JEPA

- **Student/target:** image visible context → encoder/predictor với target positions; full-image EMA encoder → contextualized target output rồi gather. Không teacher chỉ nhìn target crop.
- **Loss:** paper squared L2 ở target blocks; official train code target feature LayerNorm rồi Smooth-L1 mean. Multi-block masks không là random ratio toàn grid.
- **Gradient/update:** context encoder/predictor gradient; target no-grad và ngoài optimizer; source EMA sau optimizer update. Paper EMA .996→1; exact schedule/config cần artifact cụ thể.
- **Nghiệm/giới hạn:** asymmetry, spatial task geometry và teacher dynamics; paper/code chênh một số loss/schedule values. Không universal anti-collapse proof từ EMA.
- **Output/data/source:** encoder features (evaluation convention cần ghi context/target); ImageNet-1K, paper main batch 2048 và ViT configs khác nhau. [Paper](https://arxiv.org/html/2301.08243v3), [pinned code](https://github.com/facebookresearch/ijepa/tree/52c1ae95d05f743e000e8f10a1f3a79b10cff048), [crosswalk](research/jepa-code-doi-chieu.md).

### Audio-JEPA

- **Student/target:** native input time256 ×mel128, patch16×16 →16 time ×8 mel =128 tokens; context encoder chỉ visible patches; predictor position-conditioned, width384/6 blocks; teacher full spectrogram, output target gather dưới no-grad. Mask ratio .4–.6, random permutation split tại pinned code.
- **Loss:** paper squared vector L2; code `norm_pix_loss=true` chuẩn hóa target từng token theo coordinates, `norm_mse` L2-normalize prediction/target rồi mean $2-2\cos$. Không raw MSE identical.
- **Gradient/update:** context encoder/predictor optimizer; teacher EMA callback .996→1 ở batch end. Batch-hook clock khác optimizer `global_step` khi accumulation/limits; run setting chưa xác minh.
- **Nghiệm/giới hạn:** asymmetry/masking/teacher dynamics, không explicit variance term ở recipe chính. Latest scheduler peak LR1e-3/WD1e-6 khác paper3e-4/.05; chưa truy recipe của public weights.
- **Output/data/source:** paper downstream frozen **target encoder**, predictor bỏ; manuscript SOICT dùng **context encoder**. AudioSet-2M/general audio, paper reported batch256/100k updates; không tái lập budget. [Paper v2](https://arxiv.org/pdf/2507.02915v2), [pinned code](https://github.com/LudovicTuncay/Audio-JEPA/tree/ddd97ee9b88c00572f59bf2a00eb42120e6a4dea), [bài 6](06-JEPA-TRAINING-STEP.md).

### V-JEPA

- **Student/target:** video visible context → predictor; full-video EMA target features selected ở masked spatiotemporal positions, positions điều kiện prediction.
- **Loss/gradient/update:** paper L1 feature regression; gradient context/predictor, target stop-gradient và EMA.
- **Nghiệm/giới hạn:** masking/asymmetry/teacher dynamics; feature prediction alone chưa chứng minh action planning hoặc thế giới vật lý được identify.
- **Output/data/source:** frozen visual/video encoder cho downstream; video datasets/pretraining mixtures và scale riêng của paper, không audit số clip/steps trong pass này. [V-JEPA §3](https://arxiv.org/html/2404.08471v1).

### A-JEPA

- **Student/target:** audio patch context → encoder/predictor; EMA target latent vectors. Curriculum chuyển từ block masking sang time-frequency-aware masks.
- **Loss/gradient/update:** masked latent L2; context/predictor gradient, target EMA/detach. Không dùng downstream attention-mask regularization để mô tả thay pretraining objective.
- **Nghiệm/giới hạn:** curriculum thay task geometry; không đồng nhất paper/model với Audio-JEPA Tuncay hoặc giả định mọi masking curriculum tăng ratio.
- **Output/data/source:** audio encoder downstream, AudioSet setting; không audit checkpoint/budget trong pass này. [A-JEPA §3](https://arxiv.org/html/2311.15830v3).

### GMM-Anchored JEPA

- **Student/target:** masked/augmented speech student; clean full-view EMA continuous targets **và** frozen GMM soft posteriors từ log-mel.
- **Loss:** masked continuous regression cộng auxiliary KL $q_{GMM}\|p_{head}$; anchor weight linear decay1→.01, residual anchor vẫn hiện diện.
- **Gradient/update:** student encoder/predictor và cluster head gradient; teacher EMA/no direct gradient; GMM fit một lần, frozen trong training.
- **Nghiệm/giới hạn:** fixed acoustic anchor chống co-adaptation trong setting reported. Paper không ablate soft vs hard GMM nên “softness là nguyên nhân improvement” vẫn là hypothesis; “JEPA luôn collapse nếu không anchor” vượt evidence.
- **Output/data/source:** encoder/features, training auxiliaries bỏ theo downstream recipe; khoảng50k h Libri-Light large subset +English Granary reported, variants Conformer/Transformer. [Method/setup/limitations](https://arxiv.org/html/2602.09040v1). Author repo tồn tại, chưa audit run/provenance toàn pipeline.

### S-JEPA

- **Student/target:** waveform encoder + masked predictor/cluster head; target là **posterior distribution**, pha1 fixed MFCC GMM K100, pha2 GMM K500 từ EMA features, adaptive layer theo effective rank.
- **Loss:** single KL $q\|p$, **không continuous latent regression**. Stage-specific positions: pha1 masked+visible; pha2 ban đầu như pha1 rồi masked-only và augmentation off.
- **Gradient/update:** encoder/predictor/head gradient; GMM không gradient, pha2 cập nhật statistics online; EMA encoder supplies features, decay luân phiên .999/.9999. Pha1 không dùng EMA encoder làm target.
- **Nghiệm/giới hạn:** fixed starting anchors rồi slowly evolving clusters; không pretrained external teacher. Effective-rank layer rule là proxy, không causal proof forensic usefulness.
- **Output/data/source:** encoder giữ; predictor/head/GMM bỏ. Paper51.8M/6 layers, khoảng83k h Libri-Light +English Granary; repo default không được gán là checkpoint paper. [S-JEPA §§3–4](https://arxiv.org/html/2606.19398v1).

### GLaS-JEPA

- **Student/target:** cùng current encoder + shared token-wise linear projector128; full view tạo target, masked spans tạo prediction. Không separate temporal predictor, không separate EMA teacher.
- **Loss:** $(1-.01)L_{pred}+.01L_{SIGReg}$; masked MSE dùng detached target; SIGReg trên full-view representations nhận gradient. Main population Full Marginal across batch/time; controlled ablation có fixed-time và shuffled marginal.
- **Gradient/update:** prediction gradient qua masked path; distribution penalty qua full path; shared encoder/projector học từ cả hai. Targets đổi theo current weights, không EMA update riêng.
- **Nghiệm/giới hạn:** explicit Gaussian-distribution regularization là anti-collapse incentive; finite correlated speech không tự thỏa IID/Gaussian theory. Không tuyên bố scaling mọi cỡ model hoặc forensic utility.
- **Output/data/source:** encoder giữ, projector bỏ; LibriSpeech960 h,57M encoder,220k AdamW updates reported. Paper comparisons khác architecture/budget. [GLaS-JEPA §§3–5](https://arxiv.org/html/2609.37798v1). Official code chưa xác minh trong lượt này.

## 5. Dùng hồ sơ để kiểm hiểu

Che tên model rồi đọc student/target/loss/update. Nếu chỉ từ các ô đó nhận ra được nhóm và vẽ đúng gradient, bạn đang hiểu recipe. Nếu chỉ nhận ra nhờ tên hoặc “có masking”, quay lại bài 2–6. **Giờ data, loss thấp và rank đẹp không tự là bằng chứng cue forensic còn được giữ.**
