ajaudio / studyAUDIO-JEPA · RESEARCH NOTES
4 phút đọc · Toàn văn
Mục lục bài · 7 mục

Bài 6 — JEPA–AudioMAE: công bằng cho câu hỏi nào?#

Trước · Tiếp.

Mục tiêu và tiền đề#

Truy vết native input; phân biệt checkpoint transfer với tác động nhân quả của objective. Cần continuous targets lớp 3, recipe/source, fair contrast lớp 5. Metadata/version/unknowns ở dossier pipeline.

1. Targets upstream khác; loss downstream chung#

Audio-JEPA §III dự đoán continuous features do EMA teacher tạo. Teacher nhìn full spectrogram rồi lấy outputs ở masked positions. AudioMAE §3 dùng decoder dự đoán patch values hoặc normalized values bằng masked MSE; decoder có local attention. Paper detector bỏ predictor/decoder, dùng encoder features và cùng back-end/weighted CE ở hai arms.

Khác target tạo incentives khác nhưng chưa tự xác định forensic information được giữ. Conditional mean có thể làm mượt output reconstruction; chưa kết luận mọi hidden feature của encoder mất detail. Latent targets cũng có thể bỏ detail. Phải đo đúng readout và decision.

2. Cùng token count chưa giữ geometry#

Setting/sourceInput time×melPatch time×melGrid time×melFull NNominal time span
JEPA native256×128; 10 s; 32 kHz16×1616×812816×39,0625=625 ms
AudioMAE native1024×128; 10 s; 16 kHz16×1664×851216×10=160 ms
Common downstream256×128; 2,56 s; 32 kHz sau toolkit 16 kHz8×3232×41288×10=80 ms

Native JEPA config đặt hop=10000/256=39,0625 ms và window=2,5×hop=97,65625 ms; source fbank rồi pad/truncate. Các shape cố định chưa chứng minh có đúng từng ấy valid frames. N trong bảng là full grid trước masking, không phải số visible MAE tokens.

JEPA giữ N=128 nhưng grid 16×8 → 32×4. MAE giảm N=512 → 128. Kernel area vẫn 256 nhưng physical span/semantics đổi. Native JEPA fmax 16 kHz khác downstream 8 kHz; resampling không tạo high-band evidence mới. HF card ghi time/frequency grid đảo so với arithmetic và code: input time 256/mel 128, patch 16×16 phải cho 16 time ×8 mel. Pinned sources, PDF §§3.1/3.4 trang 4/6.

3. Factors chung và factors còn khác#

AxisChung trong paper comparisonChưa isolate
DownstreamInput/head/subset/loss/epochs/LRs/seeds/scoringCheckpoint×recipe interaction; chưa optimum riêng
ArchitectureViT-Base family, 12 blocks, width 768Exact artifact/encoder branch conventions
DataCùng nguồn AudioSet-2 MClip lists/filtering/sampling/exposures
ObjectiveCố ý khác targetsMask/normalization/optimizer/budget/native input cùng khác
AdaptationFrozen/full ở các comparisons đã chạyFrozen vẫn có learned head/input adaptation

Upstream papers báo JEPA 100k updates/batch 256 và MAE 32 epochs/batch 512. Đây chưa là run manifests của exact released weights. Current JEPA config/loss có differences với paper; MAE release script ghi 33 epochs; HF JEPA card/config được cập nhật sau weight file. Giữ từng version riêng, chưa âm thầm chọn một làm recipe thật của weights. Dossier commits.

4. Worked example: estimand khác có thể đảo thứ hạng#

Toy dưới không phải paper result:

Checkpointh₀ EER %h₁ EER %
A812
B106

Cùng h₀, A tốt hơn 2 pp. Nếu mỗi checkpoint chọn recipe trên dev rồi chấm final evaluation mới, B có thể đạt 6 so với A 8. Fixed-recipe transfer và performance của search procedure là hai estimands. Chọn h₁ bằng test chỉ tạo exploratory/oracle result.

Phản ví dụ cho objective attribution: cùng objective nhưng khác data/filtering/compute cũng có thể tạo gap. Common downstream chưa chặn các upstream paths. Muốn kiểm objective phải định nghĩa factors giữ cố định; nếu tuning riêng, cần search budget và population.

5. Scope paper đã tự nêu#

PDF §3.4/§8 coi đây là hai released checkpoints dưới common recipe; objective attribution là hypothesis. Ba lower means trong bốn comparisons vẫn là evidence cụ thể. Footnote 4 trang 3: upstream Audio-JEPA đánh giá target encoder; detector dùng context encoder, chưa đo tác động của choice đó.

6. Tự kiểm#

  1. Giữ 128 tokens có giữ native geometry JEPA chưa?
  2. Cho mỗi encoder native input đã isolate objective chưa?
  3. Frozen contrast loại được explanation nào, còn confounds nào?
  4. Có thể gán current normalized loss cho exact released run không?
Đáp án và reasoning
  1. Chưa: grid 16×8→32×4; hop/fmax/duration/kernel aspect ratio cũng đổi.
  2. Không: contrast lúc đó là hai native system packages.
  3. Gap không chỉ xuất hiện do fine-tuning encoder; upstream/input/learned-head interactions vẫn còn.
  4. Chưa: code revision và file history chưa tạo run-to-checkpoint link.

Đào sâu: viết contract cho fixed recipe và matched search budget; nêu câu hỏi mỗi thiết kế chưa trả lời. Đây là học design, chưa lựa chọn experiment mới.

DỪNG LẠI & TỰ KIỂM TRA

Bạn đã giải thích được cơ chế trong bài?

↓ Bản Markdown nguyên gốcGiữ nguyên nội dung · Công thức, bảng và nguồn đầy đủ.Các chat bàn giao được mở trong Codex.

Gõ từ khóa để tìm bài học và đoạn liên quan.