Hồ sơ 19 recipe — target, gradient, update và output#
Đây là phiếu đối chiếu phương pháp, không phải leaderboard hoặc danh sách đề xuất model. Data/budget ghi corpus/regime đã đọc; số giờ unique không bằng tổng samples seen. Chỗ chưa chốt số bước/batch/compute được đánh dấu, không suy từ tên checkpoint. Ngoài I-JEPA/Audio-JEPA code được ghim dưới đây, chi tiết là paper recipe, chưa audit implementation hoặc load weights.
Mẫu tự điền khi đọc bài mới: student view → target content/source/contextualization → main/aux loss và positions → gradient recipients → target update → cơ chế tránh nghiệm vô nghĩa → output giữ lại → data/budget/source/unknown. Loss reduction theo vector norm khác mean theo coordinate; các bài 2–6 tính sự khác biệt.
1. Predictive speech SSL#
CPC#
- Student: waveform → local convolutional encoder → autoregressive context chỉ từ quá khứ. Không cần masked input trong recipe audio gốc.
- Target: future continuous của cùng encoder; score bilinear theo horizon, đối chiếu với negatives. Future target là feature local, không full-input bidirectional teacher.
- Loss: InfoNCE chọn future positive giữa candidates, aggregate horizons/times; không coi đây là MSE reconstruction.
- Gradient/update: local encoder, autoregressive model và prediction score parameters nhận gradient, gồm feature dùng làm keys; không target teacher EMA riêng.
- Nghiệm/giới hạn: negatives khuyến khích discrimination; sampling quyết định signal có thể dùng. Equal scores cho loss , chưa có định lý task relevance.
- Output: encoder/context features cho audio probes; không score genuine/spoof mặc định.
- Data/budget/source: LibriSpeech trong audio setup CPC §§2–3.1; paper còn vision/language/RL experiments. Không chốt một số giờ/budget cho mọi modality.
wav2vec 2.0#
- Student: CNN waveform latents, span mask trước context Transformer; full-length masked latent sequence, không visible-only patch encoder.
- Target: quantized unmasked CNN latents, codebooks và selection logits được học; grouped Gumbel-Softmax/straight-through. Target không phải offline frozen ID hoặc EMA teacher.
- Loss: masked contrastive CE; negatives lấy từ các masked timesteps khác trong cùng utterance ở recipe paper. Auxiliary diversity maximizes entropy của average code-use distribution; một số recipe có feature activation L2/gradient scaling.
- Gradient/update: CNN, context Transformer, target quantizer/selection/codebook có gradient qua differentiable/ST path; diversity tác động quantizer. Không suy “target nên detach” từ tên vai trò.
- Nghiệm/giới hạn: negatives và code diversity tạo incentives khác nhau; codebook perplexity cao chưa chứng minh acoustic forensic features tốt.
- Output: context features, sau đó task head/CTC hoặc adaptation; quantizer/loss không phải downstream spoof score.
- Data/budget/source: paper dùng các regimes LibriSpeech 960 h / Libri-Light khoảng 60k h, Base/Large và số updates khác nhau; không gán cùng budget cho cả hai. wav2vec 2.0 §§2–4.
HuBERT#
- Student: waveform → CNN features bị mask spans → contextual Transformer.
- Target: offline hard k-means IDs, đầu từ MFCC, lượt sau từ intermediate features của model trước. IDs frozen trong mỗi training iteration; target không cần trùng nhãn phoneme.
- Loss: paper định nghĩa masked/unmasked mixture; main setting dùng masked loss (). Projection và learned classification codeword embeddings tạo logits; không contrastive negatives cho main loss.
- Gradient/update: student encoder/Transformer/projection/classification embeddings nhận gradient; k-means/IDs không nhận gradient từ CE. Reclustering offline giữa iterations là update pipeline riêng, không EMA.
- Nghiệm/giới hạn: fixed nontrivial targets và contextual prediction; label consistency quan trọng. Constant/poor targets vẫn có thể tạo task vô ích.
- Output: speech encoder features downstream; offline clustering head không bắt buộc giữ.
- Data/budget/source: Base LibriSpeech 960 h, larger regimes Libri-Light khoảng 60k h; iterations/budget là phần recipe, không chỉ unique hours. HuBERT §II.
WavLM#
- Student: masked speech sequence có noise/utterance mixing ở subset samples, Transformer có gated relative-position bias.
- Target: offline units từ primary/original speech, không phải mixture waveform hoặc nhãn người nói thứ hai; clustering từ speech features theo setup.
- Loss: masked-unit classification/denoising CE. “Denoise” ở đây là recover unit target, không waveform MSE. Relative-position bias là architecture, không auxiliary loss chống collapse.
- Gradient/update: student CNN/Transformer/classifier nhận gradient; pseudo-label clustering fixed trong lượt. Không EMA teacher của kiểu data2vec.
- Nghiệm/giới hạn: target ổn định + denoising task tạo incentive giữ primary speech; không chứng minh invariance với mọi noise/channel hoặc spoof.
- Output: speech encoder/features cho nhiều tasks; task head riêng.
- Data/budget/source: Base 960 h LibriSpeech; Base+/Large 94k h = 60k Libri-Light +10k GigaSpeech +24k VoxPopuli theo paper. Không so objective mà bỏ khác biệt data/scale. WavLM §IV và setup.
BEST-RQ#
- Student: acoustic features bị thay bằng small Gaussian noise ở masked spans; sequence encoder giữ time structure.
- Target: random projection và random codebook fixed, normalize projected vector/codebook rồi nearest-code assignment của original features. Hard categorical ID, không learned speech unit.
- Loss: CE ở masked positions; không negatives contrastive, không learned-quantizer diversity loss bắt buộc trong paper recipe.
- Gradient/update: speech encoder/classifier nhận gradient; projection, codebook và argmin target không học qua CE; không EMA teacher.
- Nghiệm/giới hạn: fixed mapping tránh co-adaptation của target nhưng không bảo đảm semantic/balanced targets. Nếu target IDs đều một code, constant classification có thể thắng toy task.
- Output: pretrained sequence encoder; target quantizer/classifier không cần downstream.
- Data/budget/source: Libri-Light pretraining → LibriSpeech fine-tuning; paper còn multilingual/streaming regimes, không đồng nhất compute. BEST-RQ §§3–4.
BEST-RQ-2#
- Student: 2D log-mel patches; remove masked patches khỏi ViT context encoder; predictor nhận visible features + mask token/positions, rồi classifier.
- Target: fixed random projection/codebook trên original patches, hard IDs. Paper setting K=8192, projection dim=16; không teacher features/EMA hoặc learned quantizer.
- Loss: CE chỉ masked patches, main categorical loss; không variance/covariance auxiliary được mô tả.
- Gradient/update: encoder, predictor, classifier và participating learned embeddings nhận gradient; R/codebook fixed. Token removal khác BEST-RQ original noise-in-place và one-stage BEST-RQ (ViT) comparator.
- Nghiệm/giới hạn: target diversity/predictability vẫn cần; split architecture không chứng minh tránh mọi collapse hoặc task transfer.
- Output: encoder final-block token features; predictor/head bỏ, downstream benchmark chọn pooling/probe.
- Data/budget/source: khoảng 1.9M filtered AudioSet 10 s clips; 200k updates, comparable compute reported, không strictly matched examples processed. BEST-RQ-2 §§2–4. Resource links trong paper đã ghi nhận, chưa audit code/load checkpoint.
2. Reconstruction và continuous contextual target#
AudioMAE#
- Student: log-mel → patches; encoder chỉ visible tokens, decoder nhận visible encoded tokens và learned mask tokens/positions.
- Target: original spectrogram patch values, có tùy chọn per-patch normalization; target là input space, không teacher embedding. Decoder local-window attention là architecture riêng của audio method.
- Loss: masked patch reconstruction MSE; averaging coordinate/vector/reduction phải đọc implementation trước khi so absolute values.
- Gradient/update: encoder/decoder/learned embedding nhận gradient; input values không là learned target branch; không EMA teacher.
- Nghiệm/giới hạn: fixed input targets, bottleneck/masking; decoder conditional mean có thể smooth trong MSE. Điều đó chưa chứng minh encoder mất artifact, xem counterexample bài 4.
- Output: encoder features, decoder bỏ; fine-tune/probe tùy downstream task.
- Data/budget/source: AudioSet-scale pretraining; main random mask 80%, native 10 s/16 kHz, 1,024 time ×128 mel, patch 16×16 →512 tokens theo setup/Appendix B. AudioMAE §§3–4. Không audit training budget/checkpoint history ở đây.
data2vec (bản gốc, speech)#
- Student: masked speech features qua contextual Transformer.
- Target: full-input EMA Transformer, selected top layers được normalize rồi average; speech dùng per-sample/per-feature normalization theo sequence (instance norm), lấy FFN output trước last residual theo method. Feature/positional encoders được shared; không sao chép mọi tham số frontend thành teacher EMA độc lập.
- Loss: generic formulation Smooth-L1; speech experimental setup dùng simple L2, K=8 upper layers. Target full input contextualized khác quantized local targets.
- Gradient/update: masked student có gradient; target Transformer under no-grad, EMA từ student. Shared frontend có gradient từ student, không direct gradient từ target regression branch.
- Nghiệm/giới hạn: target normalization/EMA/contextualization là training recipe, không variance penalty hoặc định lý anti-collapse. Modality-specific normalization không được tráo cho nhau.
- Output: contextual encoder features; predictor/head/task adaptation theo modality; không decoder reconstruct waveform.
- Data/budget/source: speech setup LibriSpeech 960 h; EMA .999→.9999 trong 30k updates đầu. Đây là bản paper gốc, không nhập data2vec 2.0 efficiency recipe. data2vec §§3–4.2.
3. Matching views và collapse regularization#
Các paper dưới đây có bằng chứng chính trên ảnh/ImageNet. Chúng dạy cơ chế chung; không phải experiments deepfake audio.
BYOL#
- Student/target: hai augmented views. Online encoder + projector + predictor match target encoder + projector của view kia; target normalized continuous vector. Target không có online predictor copy để regression theo cùng cách.
- Loss/positions: symmetric normalized squared distance giữa two-view pooled embeddings; không masked audio positions hoặc negatives.
- Gradient/update: gradient online encoder/projector/predictor; target detached; encoder/projector target EMA. Momentum tăng theo schedule, không một fixed tau cho cả run.
- Nghiệm/giới hạn: predictor, stop-gradient, normalization/teacher dynamics và augmentation trong recipe thực nghiệm; constant vectors vẫn là counterexample của objective alone.
- Output/data/source: online encoder bỏ projector/predictor khi evaluate; ImageNet setup, không chốt mọi ablation budget. BYOL §3.
SimSiam#
- Student/target: shared encoder/projector của hai augmented views; predictor online từng direction, target là view kia bị detach trong term đó.
- Loss/positions: symmetric negative cosine, pooled image features; không negatives hoặc masked targets.
- Gradient/update: encoder/projector/predictor cập nhật bằng gradient từ online paths của cả hai terms; không separate EMA target weights. Đừng nói “branch kia không bao giờ nhận gradient”: roles đảo trong symmetric term.
- Nghiệm/giới hạn: stop-gradient + predictor là asymmetry được ablate trong paper; không universal proof cho mọi loss/data.
- Output/data/source: encoder downstream, heads bỏ; ImageNet. SimSiam §3/Algorithm 1.
VICReg#
- Student/target: two augmented views qua shared encoder/expander; continuous paired vectors, không detached EMA teacher.
- Loss: alignment MSE giữa views + variance floor từng coordinate trên batch của mỗi view + covariance off-diagonal penalty mỗi view. Covariance dùng centered samples và .
- Gradient/update: cả view paths/encoder/expander nhận gradient từ các terms; không teacher update riêng.
- Nghiệm/giới hạn: constant output bị variance penalty, covariance term alone không đủ; exact constant point có thể có zero gradient với smoothed std. Penalty value và khả năng rời stationary point là hai câu hỏi khác nhau.
- Output/data/source: encoder bỏ expander khi downstream; ImageNet, coefficients/budget tùy setup, không tự gán cho audio. VICReg §4.1.
Barlow Twins#
- Student/target: paired augmented views, shared encoder/projector, batch-normalized coordinates.
- Loss: squared diagonal error của cross-correlation so với 1, cộng squared off-diagonals; mean cross-view alignment và redundancy khác VICReg covariance từng view.
- Gradient/update: cả paths và shared parameters học bằng gradient; không EMA/detached teacher riêng.
- Nghiệm/giới hạn: identity cross-correlation thúc đẩy non-redundant aligned features. At constant batch, ideal correlation denominator bằng 0; numerical epsilon không tạo task information.
- Output/data/source: encoder bỏ projector; ImageNet. Barlow Twins §2.1.
4. JEPA-family recipes#
I-JEPA#
- Student/target: image visible context → encoder/predictor với target positions; full-image EMA encoder → contextualized target output rồi gather. Không teacher chỉ nhìn target crop.
- Loss: paper squared L2 ở target blocks; official train code target feature LayerNorm rồi Smooth-L1 mean. Multi-block masks không là random ratio toàn grid.
- Gradient/update: context encoder/predictor gradient; target no-grad và ngoài optimizer; source EMA sau optimizer update. Paper EMA .996→1; exact schedule/config cần artifact cụ thể.
- Nghiệm/giới hạn: asymmetry, spatial task geometry và teacher dynamics; paper/code chênh một số loss/schedule values. Không universal anti-collapse proof từ EMA.
- Output/data/source: encoder features (evaluation convention cần ghi context/target); ImageNet-1K, paper main batch 2048 và ViT configs khác nhau. Paper, pinned code, crosswalk.
Audio-JEPA#
- Student/target: native input time256 ×mel128, patch16×16 →16 time ×8 mel =128 tokens; context encoder chỉ visible patches; predictor position-conditioned, width384/6 blocks; teacher full spectrogram, output target gather dưới no-grad. Mask ratio .4–.6, random permutation split tại pinned code.
- Loss: paper squared vector L2; code
norm_pix_loss=truechuẩn hóa target từng token theo coordinates,norm_mseL2-normalize prediction/target rồi mean . Không raw MSE identical. - Gradient/update: context encoder/predictor optimizer; teacher EMA callback .996→1 ở batch end. Batch-hook clock khác optimizer
global_stepkhi accumulation/limits; run setting chưa xác minh. - Nghiệm/giới hạn: asymmetry/masking/teacher dynamics, không explicit variance term ở recipe chính. Latest scheduler peak LR1e-3/WD1e-6 khác paper3e-4/.05; chưa truy recipe của public weights.
- Output/data/source: paper downstream frozen target encoder, predictor bỏ; manuscript SOICT dùng context encoder. AudioSet-2M/general audio, paper reported batch256/100k updates; không tái lập budget. Paper v2, pinned code, bài 6.
V-JEPA#
- Student/target: video visible context → predictor; full-video EMA target features selected ở masked spatiotemporal positions, positions điều kiện prediction.
- Loss/gradient/update: paper L1 feature regression; gradient context/predictor, target stop-gradient và EMA.
- Nghiệm/giới hạn: masking/asymmetry/teacher dynamics; feature prediction alone chưa chứng minh action planning hoặc thế giới vật lý được identify.
- Output/data/source: frozen visual/video encoder cho downstream; video datasets/pretraining mixtures và scale riêng của paper, không audit số clip/steps trong pass này. V-JEPA §3.
A-JEPA#
- Student/target: audio patch context → encoder/predictor; EMA target latent vectors. Curriculum chuyển từ block masking sang time-frequency-aware masks.
- Loss/gradient/update: masked latent L2; context/predictor gradient, target EMA/detach. Không dùng downstream attention-mask regularization để mô tả thay pretraining objective.
- Nghiệm/giới hạn: curriculum thay task geometry; không đồng nhất paper/model với Audio-JEPA Tuncay hoặc giả định mọi masking curriculum tăng ratio.
- Output/data/source: audio encoder downstream, AudioSet setting; không audit checkpoint/budget trong pass này. A-JEPA §3.
GMM-Anchored JEPA#
- Student/target: masked/augmented speech student; clean full-view EMA continuous targets và frozen GMM soft posteriors từ log-mel.
- Loss: masked continuous regression cộng auxiliary KL ; anchor weight linear decay1→.01, residual anchor vẫn hiện diện.
- Gradient/update: student encoder/predictor và cluster head gradient; teacher EMA/no direct gradient; GMM fit một lần, frozen trong training.
- Nghiệm/giới hạn: fixed acoustic anchor chống co-adaptation trong setting reported. Paper không ablate soft vs hard GMM nên “softness là nguyên nhân improvement” vẫn là hypothesis; “JEPA luôn collapse nếu không anchor” vượt evidence.
- Output/data/source: encoder/features, training auxiliaries bỏ theo downstream recipe; khoảng50k h Libri-Light large subset +English Granary reported, variants Conformer/Transformer. Method/setup/limitations. Author repo tồn tại, chưa audit run/provenance toàn pipeline.
S-JEPA#
- Student/target: waveform encoder + masked predictor/cluster head; target là posterior distribution, pha1 fixed MFCC GMM K100, pha2 GMM K500 từ EMA features, adaptive layer theo effective rank.
- Loss: single KL , không continuous latent regression. Stage-specific positions: pha1 masked+visible; pha2 ban đầu như pha1 rồi masked-only và augmentation off.
- Gradient/update: encoder/predictor/head gradient; GMM không gradient, pha2 cập nhật statistics online; EMA encoder supplies features, decay luân phiên .999/.9999. Pha1 không dùng EMA encoder làm target.
- Nghiệm/giới hạn: fixed starting anchors rồi slowly evolving clusters; không pretrained external teacher. Effective-rank layer rule là proxy, không causal proof forensic usefulness.
- Output/data/source: encoder giữ; predictor/head/GMM bỏ. Paper51.8M/6 layers, khoảng83k h Libri-Light +English Granary; repo default không được gán là checkpoint paper. S-JEPA §§3–4.
GLaS-JEPA#
- Student/target: cùng current encoder + shared token-wise linear projector128; full view tạo target, masked spans tạo prediction. Không separate temporal predictor, không separate EMA teacher.
- Loss: ; masked MSE dùng detached target; SIGReg trên full-view representations nhận gradient. Main population Full Marginal across batch/time; controlled ablation có fixed-time và shuffled marginal.
- Gradient/update: prediction gradient qua masked path; distribution penalty qua full path; shared encoder/projector học từ cả hai. Targets đổi theo current weights, không EMA update riêng.
- Nghiệm/giới hạn: explicit Gaussian-distribution regularization là anti-collapse incentive; finite correlated speech không tự thỏa IID/Gaussian theory. Không tuyên bố scaling mọi cỡ model hoặc forensic utility.
- Output/data/source: encoder giữ, projector bỏ; LibriSpeech960 h,57M encoder,220k AdamW updates reported. Paper comparisons khác architecture/budget. GLaS-JEPA §§3–5. Official code chưa xác minh trong lượt này.
5. Dùng hồ sơ để kiểm hiểu#
Che tên model rồi đọc student/target/loss/update. Nếu chỉ từ các ô đó nhận ra được nhóm và vẽ đúng gradient, bạn đang hiểu recipe. Nếu chỉ nhận ra nhờ tên hoặc “có masking”, quay lại bài 2–6. Giờ data, loss thấp và rank đẹp không tự là bằng chứng cue forensic còn được giữ.