ajaudio / studyAUDIO-JEPA · RESEARCH NOTES
11 phút đọc · Toàn văn
Mục lục bài · 6 mục

JEPA và Audio-JEPA — hồ sơ nguồn phương pháp#

Ngày đối chiếu: 11/10/2026.
Phạm vi: đọc nguồn primary để mô tả I-JEPA, Audio-JEPA và các biến thể gần đây; tách formulation trong paper, cấu hình/code công khai và provenance của checkpoint. Đây là ghi chú học thuật nội bộ, không phải xác nhận tái lập. Không tải checkpoint/corpus, không cài môi trường, không train hoặc chạy inference.

1. Kết luận dùng cho chương học#

JEPA nên được mô tả như một cách tổ chức dự đoán trong không gian biểu diễn, không phải một công thức cố định. Với I-JEPA và Audio-JEPA, recipe teacher–student EMA là một hiện thực cụ thể: encoder context chỉ thấy token visible; predictor nhận thêm vị trí target; target encoder cập nhật chậm theo EMA; target branch xử lý full input nhưng chỉ output ở vị trí masked được dùng làm target và không nhận gradient trực tiếp.

Không được viết rằng mọi JEPA đều dùng EMA, MSE, full-input teacher hoặc regression target liên tục. GLaS-JEPA dùng cùng encoder hiện hành cho full-view target và masked-view prediction, stop-gradient cùng SIGReg; S-JEPA dự đoán soft GMM posterior bằng KL; tên “JEPA” tự nó không xác định loại target/loss/teacher. Cần nói từng recipe.

Audio-JEPA có khác biệt quan trọng giữa paper và implementation đang công khai: paper v2 mô tả bình phương khoảng cách L2 trung bình tại target patches; config/code hiện hành chọn target normalization cộng normalized MSE/cosine-like loss. Tương tự, paper ghi weight decay 0.05 và LR peak 3e-4; config optimizer ghi những giá trị đó nhưng scheduler hiện hành đặt LR ref 1e-3 và WD ref/final 1e-6. Đây là mismatch giữa artifacts, chưa đủ bằng chứng để kết luận checkpoint public đã được train với config hiện hành.

Sai số dự đoán latent không tự là likelihood, xác suất spoof, hay một score detection đã hiệu chỉnh. Paper Audio-JEPA đánh giá representation bằng encoder downstream trên X-ARES; không dùng latent prediction error làm detector. Với bài SOICT trong workspace, context encoder + weighted CE là downstream supervised classification; không đổi tên nó thành JEPA pretraining hoặc latent-error scoring.

2. Phiếu đọc nguồn primary#

NguồnMetadata / bản được dùngPhần đọc trong lượt nàyKết luận có thể hỗ trợGiới hạn
I-JEPAMahmoud Assran et al., Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture, arXiv v3, 13/04/2023Method §3 (Targets, Context, Prediction, Loss) và Appendix A.1 (Pretraining)Multi-block context/target, target output contextualized từ full image, vị trí cho predictor, squared-L2 formulation, EMA và recipe paperNguồn ảnh/ImageNet; kết quả không tự chuyển thành bằng chứng cho audio hoặc forensic task
Code I-JEPAfacebookresearch/ijepa, commit 52c1ae95d05f743e000e8f10a1f3a79b10cff048, 13/06/2023Train step, helper optimizer, ViT/predictor, multiblock mask, V-H/14 configChi tiết chạy trong implementation và chênh lệch config/paperMột config không phải provenance của mọi checkpoint; code không đánh giá lại kết quả paper
Audio-JEPALudovic Tuncay, Etienne Labbé, Emmanouil Benetos, Thomas Pellegrini, Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning, arXiv v2, revised 26/09/2026; ICME 2025PDF §§III–IV, pp. 3–4 (method, objective, implementation); metadata và các phần setup/results được tra cứu khi cầnAudioSet input, random time-frequency patches, target/full-input relation, paper loss, architecture và stated optimizer recipeĐọc sâu ở đây tập trung vào method/implementation, không xác nhận mọi bảng kết quả hoặc revision diff v1–v2
Code Audio-JEPAddd97ee9b88c00572f59bf2a00eb42120e6a4dea, commit 16/07/2026model, criterion, masks, training wiring, optimizer schedulers, EMA callback, frontendCách chạy cụ thể tại snapshot commit nàyCommit code có trước paper v2 khoảng hai tháng; không chứng minh checkpoint nào được train từ commit này
Audio-JEPA model cardHF revision c65d33bfdef48ccfada785493f3cc0db7409c06f, README/config thêm 16/07/2026; repo API liệt kê JEPA.ckpt được thêm 22/05/2025Card/API metadata và config.json; không tải/load .ckptResource public có checkpoint/config; config mô tả model/lossCard config được commit sau checkpoint hơn một năm; không phải run manifest hay bằng chứng config train checkpoint
A-JEPAZhengcong Fei, Mingyuan Fan, Junshi Huang, A-JEPA: Joint-Embedding Predictive Architecture Can Listen, arXiv v3, 11/01/2024§3, masking curriculum và architectureNhánh audio khác, có chuyển dần từ block mask sang time-frequency-aware mask; latent L2 và EMA targetKhông đồng nhất với Audio-JEPA của Tuncay et al.; không dùng kết quả của paper này như đánh giá Audio-JEPA
V-JEPAAdrien Bardes et al., Revisiting Feature Prediction for Learning Visual Representations from Video, arXiv v1, 15/02/2024§3 training objectiveVideo feature prediction dùng stop-gradient target EMA và paper chọn L1 regression vì ổn định hơn L2Bài toán video; không suy rộng thành năng lực planning/action hoặc recipe chung cho mọi JEPA
GMM-Anchored JEPAGeorgios Ioannides et al., Soft Clustering Anchors for Self-Supervised Speech Representation Learning in Joint Embedding Prediction Architectures, arXiv v1, 30/01/2026§§3.1–3.5, limitationsFrozen GMM trên log-Mel tạo soft posterior; KL anchor bổ trợ EMA-teacher continuous JEPA loss, trọng số 1.0→0.01Speech SSL evidence; không phải deepfake evidence. Repo có public code, nhưng không kiểm tra độc lập toàn bộ run/checkpoint
S-JEPAGeorgios Ioannides et al., S-JEPA: Soft Clustering Anchors for Self-Supervised Speech Representation Learning, arXiv v1, 17/06/2026§§3.1–3.4, appendices về chuyển pha và EMASingle KL loss dự đoán soft GMM posterior: phase 1 GMM MFCC cố định K=100; phase 2 online GMM trên EMA features K=500, chọn layer theo effective rankĐây không phải continuous latent-regression JEPA; kết quả theo SUPERB không phải forensic evidence. Bản paper ghi encoder 51.8M, README public ghi cấu hình mặc định khác; không đồng nhất hai artifacts
GLaS-JEPAGaspard Botté et al., GLaS-JEPA: Gaussian-Regularized Speech SSL without Engineered Prediction Targets, arXiv v1, 29/09/2026§§3.1–3.3 và experimental setup/limitationsSame current encoder tạo full-view target và masked-view prediction; stop-gradient ở target, prediction MSE cộng SIGReg λ=.01; không separate EMA teacher/predictorPaper ghi so sánh hạn chế bởi khác biệt data/architecture/budget; không có official code được xác minh trong lượt này

3. Hai recipe gốc cần dạy tách biệt#

I-JEPA: hình ảnh, multi-block, target output được contextualize#

  1. Ảnh được chia patch. Context block lấy scale 0.85–1.0, aspect ratio 1; thường lấy bốn target blocks scale 0.15–0.20, aspect ratio 0.75–1.5. Vùng target chồng lên context bị gỡ khỏi context. Ví dụ config ViT-H/14: lưới 16×16, 1 context mask, 4 target masks, patch size 14.
  2. Context encoder chỉ xử lý visible context patches. Target encoder nhận toàn ảnh trước, tạo chuỗi feature contextualized; sau đó mới gather target blocks từ output. Vì self-attention thấy toàn ảnh, target vector có ngữ cảnh ngoài patch bị che.
  3. Predictor hẹp hơn encoder. Nó nhận context outputs cộng positional embeddings và learned mask token cộng vị trí target; chạy riêng theo từng target block trong formulation/code.
  4. Paper §3 Loss ghi trung bình trên target blocks của tổng squared-L2 theo patch. Target parameters được EMA từ context encoder; gradient cập nhật context encoder và predictor.
  5. Appendix A.1 nêu AdamW, batch global 2048, LR 1e-4 warmup lên 1e-3 trong 15 epochs rồi cosine về 1e-6; WD 0.04→0.4; EMA 0.996→1.0.

Audio-JEPA: spectrogram, random patch masking, target encoder thấy đủ input#

  1. Frontend được paper mô tả là mono AudioSet, 10 giây/32 kHz, 128-band log-Mel và 256 time bins. Config data/code hiện hành dùng Kaldi fbank, trừ mean waveform, Hanning, low cutoff 20 Hz, high cutoff Nyquist, log features; output shape code là (1, time=256, mel=128).
  2. Patch 16×16 trên trục time×mel tạo 16 temporal × 8 mel = 128 tokens. Tỉ lệ mask 40–60%; khoảng chính xác lấy uniform một lần cho mỗi batch-call, rồi mỗi ví dụ lấy permutation riêng. Paper thử I-JEPA block strategy nhưng báo random masking tốt hơn trong preliminary experiment.
  3. Encoder context xử lý token visible. Target encoder chạy trên toàn spectrogram; target vectors là full-context ViT outputs tại masked positions, lấy bằng mask indexing dưới no-grad. Do đó target là biểu diễn contextualized của vùng masked, không phải feature của patch độc lập.
  4. Predictor nhận visible tokens/vị trí context; target positions nhận learned mask token cộng fixed 2D sin-cos position. 768→384, 6 transformer blocks, LayerNorm rồi project ngược 768. Encoder/target 12 layers, 768 dim, 12 heads, MLP ratio 4.
  5. Paper objective (v2, Eq. 2) mô tả trung bình squared Euclidean distance giữa predicted và target vectors tại masked patches. Code config lại chọn norm_mse cùng norm_pix_loss=true: criterion chuẩn hóa target theo mean/variance từng token qua feature dimension, rồi L2-normalize prediction và target, loss 2−2·dot, mean trên tokens/batch. Vì vậy câu “MSE” cần gắn nhãn paper hay code.
  6. Paper ghi AdamW β=(.9,.95), WD .05, initial LR 3e-4, warmup 1e-6→3e-4 trong 1,000 steps, rồi cosine về 0. Thực nghiệm báo tổng batch 256, 4 NVIDIA V100, 100,000 steps (~13 epochs, ~14 giờ). Đây là setting paper báo cáo, chưa tái lập trong lượt này. Snapshot code config/scheduler đặt LR ref 1e-3 và WD 1e-6 không đổi; xem crosswalk ở jepa-code-doi-chieu.md.
  7. Paper nêu downstream dùng frozen target encoder/embedding + task probe; predictor bỏ. Không có latent prediction-error detector trong paper. Không lấy error của objective pretraining làm spoof probability nếu chưa có nghiên cứu riêng về score, threshold và calibration.

4. EMA: ví dụ số và cách nói chính xác#

Quy tắc là θˉnew=τθˉold+(1−τ)θonline\bar{\theta}_{new}=\tau\bar{\theta}_{old}+(1-\tau)\theta_{online}. Nếu τ=0.996\tau=0.996, teacher cũ bằng 2 và online mới bằng 7, teacher mới bằng 0.996×2 + 0.004×7 = 2.02: online weights chỉ đóng góp 0.4% ở bước đó. Nếu tau giữ cố định .996, phần trọng số ban đầu còn lại sau 100 bước là 0.996100≈0.6700.996^{100}\approx0.670; half-life là log⁡(0.5)/log⁡(0.996)≈172.9\log(0.5)/\log(0.996)\approx172.9 updates.

Đây là minh họa đại số với tau cố định, không phải tuổi thọ thực tế của run I-JEPA/Audio-JEPA. I-JEPA paper/code schedule momentum tăng về 1; Audio-JEPA callback dùng cosine .996→1. EMA là phép làm chậm target, không phải bảo đảm chống collapse cho mọi objective.

Đặc biệt, Audio-JEPA callback cập nhật EMA ở hook on_train_batch_end, trong khi tau dựa global_step nhưng mẫu số là len(train_dataloader) × max_epochs. Hai thang chỉ khớp đơn giản khi mỗi batch có đúng một optimizer step, không có batch limits/early stop/max_steps hoặc gradient accumulation. Với accumulation, hook có thể chạy nhiều lần khi student weights chưa đổi còn global_step chưa tăng; source alone chưa chứng minh callback được chạy trong setting đó. Không diễn giải công thức schedule như cập nhật trên mỗi optimizer step nếu không kiểm tra trainer config/run log.

5. Nhánh JEPA 2026 và giới hạn suy rộng#

  • GMM-Anchored JEPA: GMM diagonal covariance fit một lần trên log-Mel; GMM frozen. Phase train kết hợp JEPA MSE với EMA teacher và KL prediction của soft posterior, λ\lambda giảm 1.0 xuống residual .01. Paper còn thêm denoising augmentations cho student. Gọi đây là hybrid continuous prediction + acoustic anchor.
  • S-JEPA: “JEPA-style” encoder–predictor nhưng loss là KL giữa soft output và posterior GMM. Pha 1 GMM K=100 trên MFCC 39 chiều; pha 2 K=500 trên EMA encoder features, GMM cập nhật online và chọn layer theo effective rank. EMA decay luân phiên nhanh/chậm ở phase 2; phase 1 không dùng EMA encoder làm target. Paper ghi 83k giờ LibriLight + Granary, encoder 51.8M theo setup paper. Không nói đây là cùng recipe với GMM-Anchored.
  • GLaS-JEPA: Full view và masked view cùng đi qua encoder hiện hành + shared token-wise projector; gradient prediction chỉ qua masked path do stop-gradient ở full-view target. SIGReg trên full-view representation bổ sung chống collapse; λ=.01\lambda=.01. Không separate EMA target, không temporal predictor riêng. Paper ghi LibriSpeech 960h, 57M encoder. Đây là evidence speech SSL, không deepfake.
  • A-JEPA là bài Fei et al. 2023/2024, không phải Audio-JEPA Tuncay et al. Nó dùng curriculum mask từ random blocks về time-frequency masks và EMA target.
  • V-JEPA là video feature-prediction paper Bardes et al. 2024. Bản paper dùng L1 regression cùng EMA target/stop-gradient; không gộp với loss của I-JEPA/Audio-JEPA, và không tự chứng minh action planning/world-model behavior.

Tên mới trong map đều có bài arXiv primary theo ID hiện nêu ở trên. Author-linked code repositories đã được tìm thấy cho GMM-Anchored JEPA (snapshot edb58af13297c0a6108e1b6eba459176138592e4, 02/07/2026) và S-JEPA (snapshot aab0a1da37ffbb7c577df36f0f10fb0984e677d8, 23/06/2026). README/config parameter counts không đồng nhất hoàn toàn với paper S-JEPA; không coi code-default là paper checkpoint. GLaS-JEPA chưa tìm được official repo trong lượt đối chiếu. Không có hàng nào ở đây chứng minh hữu ích cho audio deepfake detection.

6. Quy tắc dùng claim#

  1. Gắn mỗi chi tiết vào paper, pinned source code hay model card; đừng trộn chúng thành một “paper recipe”.
  2. Nêu rõ loại target (continuous vector hay posterior distribution), nguồn target, context/target visibility, gradient path và loss reduction/normalization.
  3. Phân biệt update EMA theo train batch với optimizer step; kiểm tra accumulation và trainer limits trước khi quy đổi lịch theo “steps”.
  4. Không suy provenance checkpoint từ config mới hơn. Với Audio-JEPA, checkpoint file trên HF có commit 22/05/2025, còn README/config revision thêm 16/07/2026. Không tải hoặc load checkpoint trong lượt này.
  5. Kết quả ASR/emotion/slot-filling/general audio chỉ là kết quả trên các task/protocol đó; không phải evidence cho spoof detection, unseen-generator robustness hoặc calibrated detection score.
↓ Bản Markdown nguyên gốcGiữ nguyên nội dung · Công thức, bảng và nguồn đầy đủ.Các chat bàn giao được mở trong Codex.

Gõ từ khóa để tìm bài học và đoạn liên quan.