ajaudio / studyAUDIO-JEPA · RESEARCH NOTES
8 phút đọc · Toàn văn
Mục lục bài · 7 mục

Bài 10 — Đọc biến thể bằng recipe và bằng chứng#

Đích học: nhận diện một model từ luồng target/loss/update; phân biệt phương pháp có thật, implementation có thật, checkpoint có thật và kết quả có thể tái lập; không suy forensic utility từ tên JEPA hoặc benchmark ASR. Cần bài 2, 6, 9.

1. “JEPA” chưa điền đủ một recipe#

Với một paper mới, điền bảy ô trước khi đọc bảng thắng/thua: student thấy gì; target chứa gì; target tính từ đâu; loss tính ở đâu; gradient đi đâu; tham số target cập nhật ra sao; feature nào giữ downstream. Hai paper cùng tên họ có thể khác cả target, loss và teacher.

Nguồn primaryStudent / targetLoss và cập nhậtBài học phân biệt
I-JEPA, Assran et al.Visible image context → predictor; full-image EMA encoder → target vectorsPaper squared L2; code target LayerNorm + Smooth-L1, gradient context/predictor, teacher EMAFull-view contextualized target khác patch-local target
V-JEPA, Bardes et al.Video context → predictor; full-video EMA featuresPaper L1 trên target features; teacher stop-gradient/EMANgay họ continuous JEPA cũng không có một loss MSE duy nhất
A-JEPA, Fei et al.Audio context → predictor; EMA target vectorsLatent L2; curriculum block mask → time-frequency maskPaper riêng với Audio-JEPA của Tuncay; curriculum là geometry chứ không chỉ ratio
Audio-JEPA, Tuncay et al.Visible log-mel patches → predictor; full-spectrogram EMA featuresPaper squared L2; released code target standardization + unit-normalized lossCần ghi paper/config/checkpoint riêng, xem bài 6
GMM-Anchored JEPAStudent masked/augmented speech; clean EMA latent target và frozen log-mel GMM posteriorContinuous regression + auxiliary KL anchor giảm trọng số 1→.01Hybrid hai loại target; GMM này không cập nhật online
S-JEPAMasked speech → predictor/head; GMM posteriorSingle KL, pha 1 frozen MFCC GMM; pha 2 online GMM nhận EMA featuresEMA cung cấp features cho clustering; không phải teacher vector để regression
GLaS-JEPAFull/masked views qua cùng encoder hiện hành + token-wise projectorMasked MSE với stop-gradient target + full-view SIGReg; không EMA teacher riêngLoss prediction và regularization có gradient paths khác nhau
BEST-RQ-2Visible-only ViT → predictor/classifier; fixed random-quantizer patch IDsMasked CE; gradient encoder/predictor/classifier; R/codebook fixedDecomposition context/predictor không bắt buộc continuous target hoặc teacher EMA

Đọc phương pháp ở các phiên bản được dẫn, không ghép tên giống nhau thành cùng model. Bảng là bản đồ; hồ sơ recipe ghi dữ liệu, head giữ lại, normalization và giới hạn nguồn.

2. Hai loại GMM target phải tách riêng#

GMM cho posterior qk(m)=πkN(m;μk,Σk)/∑lπlN(m;μl,Σl)q_k(m)=\pi_k\mathcal N(m;\mu_k,\Sigma_k)/\sum_l\pi_l\mathcal N(m;\mu_l,\Sigma_l). Một điểm gần ranh giới có thể có q=(.5,.5)q=(.5,.5) thay vì hard ID. Với logits student pp, KL q∥pq\|p hoặc soft CE có cùng gradient nếu qq cố định; chúng khác entropy hằng số của target, xem bài 2.

GMM-Anchored: một GMM fit trên log-mel rồi giữ cố định, posterior là anchor bổ sung cho latent JEPA loss. Trọng số anchor giảm trong training. Target acoustic không tự là phoneme hay forensic cue. Method §§3.1–3.4.

S-JEPA: pha 1 GMM MFCC 39 chiều, K=100 cố định; pha 2 K=500 trên EMA encoder features, cập nhật GMM từ minibatch sufficient statistics và chọn layer theo effective rank. Predictor/head match posterior bằng KL, không có latent-regression term thứ hai. EMA decay pha 2 luân phiên .999/.9999. Set positions không bất biến suốt training: pha 1 dùng masked + visible; pha 2 chuyển từ cấu hình đó sang masked-only và tắt augmentation. Phần architecture viết visible logits không supervised ở mô tả masked-only; khi ghi recipe phải kèm stage ở §3.3 và Appendix C. S-JEPA §§3.1–3.4.

Điểm dễ nhầm: “không pretrained teacher distillation” không đồng nghĩa không có EMA encoder. S-JEPA vẫn dùng EMA features để tạo vocabulary online; chỉ không hồi quy trực tiếp features đó hoặc kế thừa một teacher pretrained ngoài pipeline.

3. GLaS-JEPA: stop-gradient không chặn regularizer#

Full path tạo z=pα(fθ(h))z=p_\alpha(f_\theta(h)), masked path tạo z^=pα(fθ(m(h)))\hat z=p_\alpha(f_\theta(m(h))). Projector tuyến tính từng token có 128 chiều, không trộn các vị trí; encoder làm contextual prediction, không temporal predictor riêng. Objective paper:

L=(1−λ)Lpred+λLSIGReg,λ=.01.L=(1-\lambda)L_{pred}+\lambda L_{SIGReg},\quad\lambda=.01.

LpredL_{pred} ở masked positions dùng sg⁡(z)\operatorname{sg}(z) nên gradient qua masked path. SIGReg dùng full-view zz không detach để regularize distribution; gradient qua full path. Hai paths cùng tham số nên cùng encoder nhận hai contributions. “Không teacher EMA” không phải “không target”: target vẫn là full-view current representation. Cụm “without engineered targets” cần hiểu theo định nghĩa paper. GLaS-JEPA §§3.1–3.3.

Population SIGReg có thể là fixed-time across utterances hoặc marginal across times/utterances. Thay population đổi ý nghĩa penalty, xem phản ví dụ position-only ở bài 8. Các giả định của lý thuyết Gaussian/identifiability được paper dẫn không tự đúng với speech; không dùng một theorem của bài khác để khẳng định forensic signal được bảo toàn.

4. BEST-RQ-2: architecture split với categorical target#

Trace: patch input gốc → random projection/codebook cố định → ID; visible patches → ViT encoder; context features + mask tokens/vị trí → predictor → classifier logits; CE chỉ masked positions. Không teacher EMA, không learned quantizer. Predictor bị bỏ downstream. BEST-RQ gốc dùng in-place corruption trong sequence encoder; BEST-RQ-2 bỏ masked tokens khỏi context encoder rồi thêm mask tokens ở predictor. BEST-RQ-2 §§2.1–2.6.

Paper có BEST-RQ (ViT) để đối chiếu architecture/patchification với decomposition hai bước. Đối chứng này hỗ trợ một câu hỏi cơ chế trong protocol của paper, chưa đủ chuyển kết luận sang spoof detection. Không nhập target IDs ngẫu nhiên vào cùng nhóm “speech units từ clustering” chỉ vì đều dùng CE.

5. Chuỗi bằng chứng có nhiều nấc#

NấcĐiều đã xác nhậnChưa xác nhận
arXiv metadata và abstractTitle, authors, version, claim tác giả nêuCơ chế chi tiết, implementation đúng
Method equations/algorithmTarget, mask, loss, updates được mô tảCode/checkpoint thực hiện y hệt
Pinned code/configCách một revision wired functions/schedulesRevision đó tạo checkpoint hoặc reproduces table
Model card/file metadataResource tồn tại, file lịch sử ra saoExact training recipe, weights integrity hoặc score thực tế
Run manifest/log + artifact hashLiên kết một run với config/checkpoint cụ thểKết quả robust ngoài protocol đó
Independent evaluationKết quả trong một protocol được chạy lạiMọi domain/task khác

Với Audio-JEPA, source code và later card có năm 2026; checkpoint file có commit tháng 5/2025. Không dùng latest config để gán recipe cho checkpoint cũ. Với GLaS-JEPA chưa xác minh official repo trong lượt này. Với BEST-RQ-2 paper liên kết resource nhưng chưa audit implementation/checkpoint. Crosswalk và provenance.

6. “Hơn baseline” trả lời câu hỏi nào?#

Một checkpoint A pretrained trên general audio và B pretrained trên speech, khác architecture/frontend/compute, rồi cùng train downstream head: đo hai hệ thống checkpoint + adaptation. Nó không isolate ảnh hưởng objective A vs B. Giữ head giống nhau giảm một số confound, không làm các yếu tố pretraining tự bằng nhau.

Toy decomposition của một improvement quan sát Δ\Delta:

Δ=Δobjective+Δdata+Δarchitecture+Δbudget+Δfrontend+Δinteraction.\Delta=\Delta_{objective}+\Delta_{data}+\Delta_{architecture}+\Delta_{budget}+\Delta_{frontend}+\Delta_{interaction}.

Đây là ký hiệu để nhắc các yếu tố, không là mô hình additive đã được nhận dạng từ dữ liệu. Một tổng quan sát không cho từng term; interactions có thể không additive. Cần đối chứng tương ứng trước khi viết causal claim.

GLaS-JEPA tự ghi baselines khác architecture/budget, nên bảng contextual comparison không isolate SSL objective. S-JEPA đo ASR/emotion/slot-filling; Audio-JEPA/BEST-RQ-2 đo general-audio tasks. Những kết quả đó giúp biết loại information đọc được trong các protocol cụ thể, chưa là bằng chứng unseen-spoof robustness. GLaS-JEPA §5.1, S-JEPA §4.

7. Bài tập tổng hợp — nhận diện recipe không nhìn tên#

  1. Model có R/codebook fixed, visible-only encoder, predictor và CE masked. Nó có bắt buộc teacher EMA không? Thuộc nhóm target nào?
  2. Model có EMA encoder nhưng head dự đoán GMM posterior bằng KL. Có thể gọi loss continuous latent regression không?
  3. Một full-view path nhận zero gradient từ MSE nhưng nhận gradient từ SIGReg. Câu “full-view network không học” sai ở đâu?
  4. Repo có checkpoint tháng 5/2025 và config thêm tháng 7/2026. Có thể khẳng định checkpoint dùng WD trong config mới không?
  5. Encoder ASR tốt hơn trên LibriSpeech có đủ bằng chứng chọn nó cho unseen generator không?
  6. Hai checkpoints có cùng downstream recipe nhưng khác pretraining data. Claim “JEPA objective tốt hơn MAE objective” cần giới hạn thế nào?
Đáp án và lý do
  1. Không. Đây là categorical fixed-random target, architecture encoder–predictor split; tương ứng BEST-RQ-2 ở source đã đọc.
  2. Không. EMA features làm input cho clustering; target optimization là distribution, loss KL.
  3. Chặn một term không chặn term khác. Shared parameters vẫn nhận gradient full-view regularization.
  4. Không. Cần run/config provenance cùng checkpoint; mốc file sau không chứng minh lịch sử run trước.
  5. Không. ASR và forensic cue khác nhau, protocol generalization cũng khác.
  6. Có thể nói checkpoint systems khác nhau dưới downstream protocol này; chưa isolate causal effect của objective.

Bài đánh giá cuối lớp: chọn một recipe trong hồ sơ 14, che tên, tự vẽ forward/gradient/update và nêu một shortcut, một phản ví dụ cho claim downstream, một bằng chứng còn thiếu. Làm xong mới đánh dấu “đã hiểu”; việc tài liệu được viết xong chưa tính là người học đã học.

DỪNG LẠI & TỰ KIỂM TRA

Bạn đã giải thích được cơ chế trong bài?

↓ Bản Markdown nguyên gốcGiữ nguyên nội dung · Công thức, bảng và nguồn đầy đủ.Các chat bàn giao được mở trong Codex.

Gõ từ khóa để tìm bài học và đoạn liên quan.