ajaudio / studyAUDIO-JEPA · RESEARCH NOTES
7 phút đọc · Toàn văn
Mục lục bài · 5 mục

Sổ nguồn lớp 3 — phiên bản, phần đọc và phạm vi claim#

Ngày đối chiếu: 11/10/2026. Nguồn phương pháp là paper/book primary và repository tác giả. Không dùng blog/tóm tắt làm bằng chứng cơ chế. “Đã đọc method” không đồng nghĩa đọc mọi appendix, tái lập kết quả hoặc xác nhận checkpoint.

1. Nguồn lý thuyết và SSL kinh điển#

Nguồn / tác giảPhiên bản được dùngPhần đối chiếuDùng cho bài
Goodfellow, Bengio, Courville — Deep Learning, ch. 15Representation Learning, sách 2016Mở đầu, §15.3 về unsupervised representation và downstream task1: information/accessibility/task dependence
Tishby, Pereira, Bialek — The Information Bottleneck Methodphysics/0004057§3, Eq. 14–15, objective và joint distribution1: MI/bottleneck; toy discrete tự xây dựng
van den Oord, Li, Vinyals — Representation Learning with Contrastive Predictive Coding1807.03748v2§§2.1–2.3, audio §3.12–3: future prediction, InfoNCE
Baevski et al. — wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations2006.11477v3§§2–3, pretraining setup §4.22–3: differentiable quantizer, masked contrastive loss, code diversity
Hsu et al. — HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units2106.07447v1§II-A–E, iterative targets và masked/unmasked loss3: offline hard units khác learned quantizer
Chen et al. — WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing2110.13900v5§IV-A–C, Algorithm 1, masked speech denoising3: primary-speech target khi student bị noise/overlap
Chiu et al. — Self-supervised Learning with Random-projection Quantizer for Speech RecognitionICML/PMLR 162, 2022 PDF; agent metadata 2202.01855v2§§3.1–3.2, đối chiếu §4.1 cho dữ liệu; agent thêm §3.43: fixed random target, normalized nearest code
Huang et al. — Masked Autoencoders that Listen2207.06405v3§3, §§4.2/4.4, frontend Appendix B4: visible-only encoder, local decoder, masked spectrogram reconstruction
Baevski et al. — data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language2202.03555v3§§3.1–3.4, speech §4.2; agent thêm §64: full-input contextual targets, top-layer normalization/average; speech L2 khác generic Smooth-L1
Grill et al. — Bootstrap Your Own Latent: A New Approach to Self-Supervised Learning2006.07733v3§§3.1–3.2; agent thêm ablations §55: normalized loss, stop-gradient, online predictor, EMA
Chen, He — Exploring Simple Siamese Representation Learning2011.10566v1§3, Algorithm 1, Eq. 4; agent §§4.1–4.25: shared encoder, symmetric detached targets, không EMA
Bardes, Ponce, LeCun — VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning2105.04906v3§4.1, Eq. 1–65/8: sample covariance, variance penalty và collapse caveats
Zbontar et al. — Barlow Twins: Self-Supervised Learning via Redundancy Reduction2103.03230v3§2.1, Eq. 1–2; agent pseudocode/§2.25/8: cross-correlation identity và redundancy

Primary URLs của các bài arXiv ở lessons có thể mở PDF hoặc HTML. Root đọc trực tiếp phần phương pháp của từng họ cốt lõi nêu trên, gồm các đoạn loss/target/gradient dùng để viết bài. Hai agent bổ sung metadata và notes; ghi chú của agent không được nâng thành “root đã đọc toàn văn”. Root kiểm lại các điểm dễ sai: wav2vec2 negatives trong cùng utterance, data2vec speech dùng L2, target normalization, shared frontend, SimSiam không teacher EMA, VICReg covariance/sample axis.

Các ghi chú chi tiết và tác giả đầy đủ hơn: hồ sơ SSL của agent. Code fairseq hiện có thể mô tả data2vec 2.0; không dùng README đó để chứng minh recipe data2vec gốc. Root không audit implementation các SSL models kinh điển trong lượt này.

2. JEPA và các nguồn mới được kiểm chứng method#

Nguồn / tác giảPhiên bản / metadataPhần đã đọcGiới hạn chính
Assran et al. — Self-Supervised Learning from Images with a Joint-Embedding Predictive ArchitectureI-JEPA 2301.08243v3, 13/04/2023Root §§2–3 bản v1; agent §3 + Appendix A.1 bản v3; root kiểm official train codeImageNet; method/code khác loss và một số schedule values
Bardes et al. — Revisiting Feature Prediction for Learning Visual Representations from VideoV-JEPA 2404.08471v1, 15/02/2024Root/agent §3.1 objectiveVideo, L1 feature loss; chưa audit code hoặc planning capability
Fei, Fan, Huang — A-JEPA: Joint-Embedding Predictive Architecture Can Listen2311.15830v3, 11/01/2024§3 architecture/curriculum và downstream maskingKhông phải Audio-JEPA của Tuncay; chưa audit checkpoint
Tuncay, Labbé, Benetos, Pellegrini — Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning2507.02915v2, revised 26/09/2026Root/agent §§III–IV, root một số result/limitations; official code bên dướiKhông đọc revision diff hoặc tái lập mọi bảng; latest code/card chưa chứng minh recipe checkpoint
Ioannides et al. — Soft Clustering Anchors for Self-Supervised Speech Representation Learning in Joint Embedding Prediction ArchitecturesGMM-Anchored 2602.09040v1, HTML/header 30/01/2026Root §§3.1–3.4, §4 setup, limitations; agent §§3.1–3.5Không ablate soft vs hard GMM theo limitations; không forensic evidence
Ioannides et al. — S-JEPA: Soft Clustering Anchors for Self-Supervised Speech Representation Learning2606.19398v1, 17/06/2026Root §§3.1–3.4 và setup; agent transition/EMA appendicesSingle KL, loss position đổi theo stage; paper encoder 51.8M không đồng nhất repo default
Tuncay, Labbé, Pellegrini — BEST-RQ-2: Contextualize–Then–Predict, a Two-Step Approach for Self-Supervised Audio Representations2606.30700v1, submitted 29/06/2026Root §§2.2–2.6, §3 và §§4.1/4.3; agent §§2–3 và selected resultsPaper resource links chưa được audit/load; không detector evidence
Botté et al. — GLaS-JEPA: Gaussian-Regularized Speech SSL without Engineered Prediction Targets2609.37798v1, 29/09/2026Root/agent §§3.1–3.3, setup và limitationsChưa xác minh official code; baselines không matched architecture/budget

Timestamps above là metadata của artifact, không phải ngày lượt học. Với GMM-Anchored, ngày header không trùng tháng trong ID; giữ nguyên ngày source ghi, không tự “sửa” để đoán lịch nộp. Title/version có thật và method đã được đọc; điều đó chưa xác nhận mọi claim thực nghiệm. Notes: JEPA methods, BEST-RQ-2 follow-up.

3. Official code và provenance#

Artifact bất biếnRoot kiểm trực tiếpDùng để kết luận
I-JEPA @ 52c1ae95d05f743e000e8f10a1f3a79b10cff048src/train.py: forward_target, loss_fn, optimizer/EMA order; agent đọc mask/predictor/helper/configTeacher full-image → LayerNorm → target selection; Smooth-L1; target no-grad; EMA sau optimizer
Audio-JEPA @ ddd97ee9b88c00572f59bf2a00eb42120e6a4deajepa_module.py, criterion, ViT/predictor, train.py, YAML; mask sampler, EMA callback, LR/WD schedulerNative grid, target visibility, normalization, gradient groups, scheduler overrides và batch-hook EMA caveat
HF Audio-JEPA later config/card revisionAgent đọc API/card/config, root đối chiếu consistency với code; root web open JSON không thành côngDocumentation revision 16/07/2026; axis labels cần kiểm bằng source dimensions
HF checkpoint file commitAgent file history metadata; không tải weightsFile commit 22/05/2025; later config không đủ chứng minh training recipe

Chi tiết function/config và line anchors ở crosswalk. Root đã mở lại scheduler: LR dùng ref_lr chứ không giữ optimizer initial LR; WD scheduler gán ref_wd/final_wd vào non-excluded groups. Đây là source-level behavior; không có run log để xác nhận effective trainer setting trong training đã phát hành.

Không tải, hash hoặc load checkpoints; không chạy inference/train; không xác nhận bitwise reproducibility. Public repository tồn tại không đồng nghĩa paper configuration có checkpoint được truy nguyên.

4. Manuscript local là nguồn downstream riêng#

  • File: SOICT_2026_paper_4308.pdf, 14 trang.
  • SHA-256 lúc đọc: c48da893ad2d7d1e5bbb036d609ccbadfb088a1d9e43a5a526a247e3e481cfd3.
  • Lượt lớp 3 đọc/extract trực tiếp §§3.1–3.4 ở trang 3–6, kiểm ảnh render method trang 4. Text method là working note, không thay bản PDF.
  • Hỗ trợ claim: context encoder downstream; frontend/patch adaptation; 13-output layer combination; attentive statistics pooling/MLP; frozen/fine-tune; weighted CE và logit-difference score.
  • Không tuyên bố tái chạy bảng kết quả, xác nhận notebook outputs hoặc checkpoint match. Mức đọc full 14 trang trong global start là lịch sử bộ ôn tập trước đó, không được dùng để nói lượt enrich này đã đọc lại toàn bộ.

5. Quy ước bằng chứng trong học liệu#

Lý thuyết/đại số: điều kiện và phép suy ra được trình bày, như conditional mean, loss gradient, collision/invertible transform. Toy: dữ liệu tự tạo để chứng minh khả năng hoặc phản ví dụ; không phải số liệu checkpoint. Paper reported: phương pháp/kết quả tác giả báo cáo trong version dẫn. Code verified: behavior đọc từ pinned source. Chưa xác minh: runtime/provenance/task transfer chưa có bằng chứng tương ứng.

Script kiểm tra xác minh các toy và đường dẫn nội bộ; nó không xác minh một model thực tế tránh collapse, phát hiện spoof hoặc tái lập paper.

↓ Bản Markdown nguyên gốcGiữ nguyên nội dung · Công thức, bảng và nguồn đầy đủ.Các chat bàn giao được mở trong Codex.

Gõ từ khóa để tìm bài học và đoạn liên quan.