Bài 6 — JEPA–AudioMAE: công bằng cho câu hỏi nào?#
Mục tiêu và tiền đề#
Truy vết native input; phân biệt checkpoint transfer với tác động nhân quả của objective. Cần continuous targets lớp 3, recipe/source, fair contrast lớp 5. Metadata/version/unknowns ở dossier pipeline.
1. Targets upstream khác; loss downstream chung#
Audio-JEPA §III dự đoán continuous features do EMA teacher tạo. Teacher nhìn full spectrogram rồi lấy outputs ở masked positions. AudioMAE §3 dùng decoder dự đoán patch values hoặc normalized values bằng masked MSE; decoder có local attention. Paper detector bỏ predictor/decoder, dùng encoder features và cùng back-end/weighted CE ở hai arms.
Khác target tạo incentives khác nhưng chưa tự xác định forensic information được giữ. Conditional mean có thể làm mượt output reconstruction; chưa kết luận mọi hidden feature của encoder mất detail. Latent targets cũng có thể bỏ detail. Phải đo đúng readout và decision.
2. Cùng token count chưa giữ geometry#
| Setting/source | Input time×mel | Patch time×mel | Grid time×mel | Full N | Nominal time span |
|---|---|---|---|---|---|
| JEPA native | 256×128; 10 s; 32 kHz | 16×16 | 16×8 | 128 | 16×39,0625=625 ms |
| AudioMAE native | 1024×128; 10 s; 16 kHz | 16×16 | 64×8 | 512 | 16×10=160 ms |
| Common downstream | 256×128; 2,56 s; 32 kHz sau toolkit 16 kHz | 8×32 | 32×4 | 128 | 8×10=80 ms |
Native JEPA config đặt hop=10000/256=39,0625 ms và window=2,5×hop=97,65625 ms; source fbank rồi pad/truncate. Các shape cố định chưa chứng minh có đúng từng ấy valid frames. N trong bảng là full grid trước masking, không phải số visible MAE tokens.
JEPA giữ N=128 nhưng grid 16×8 → 32×4. MAE giảm N=512 → 128. Kernel area vẫn 256 nhưng physical span/semantics đổi. Native JEPA fmax 16 kHz khác downstream 8 kHz; resampling không tạo high-band evidence mới. HF card ghi time/frequency grid đảo so với arithmetic và code: input time 256/mel 128, patch 16×16 phải cho 16 time ×8 mel. Pinned sources, PDF §§3.1/3.4 trang 4/6.
3. Factors chung và factors còn khác#
| Axis | Chung trong paper comparison | Chưa isolate |
|---|---|---|
| Downstream | Input/head/subset/loss/epochs/LRs/seeds/scoring | Checkpoint×recipe interaction; chưa optimum riêng |
| Architecture | ViT-Base family, 12 blocks, width 768 | Exact artifact/encoder branch conventions |
| Data | Cùng nguồn AudioSet-2 M | Clip lists/filtering/sampling/exposures |
| Objective | Cố ý khác targets | Mask/normalization/optimizer/budget/native input cùng khác |
| Adaptation | Frozen/full ở các comparisons đã chạy | Frozen vẫn có learned head/input adaptation |
Upstream papers báo JEPA 100k updates/batch 256 và MAE 32 epochs/batch 512. Đây chưa là run manifests của exact released weights. Current JEPA config/loss có differences với paper; MAE release script ghi 33 epochs; HF JEPA card/config được cập nhật sau weight file. Giữ từng version riêng, chưa âm thầm chọn một làm recipe thật của weights. Dossier commits.
4. Worked example: estimand khác có thể đảo thứ hạng#
Toy dưới không phải paper result:
| Checkpoint | h₀ EER % | h₁ EER % |
|---|---|---|
| A | 8 | 12 |
| B | 10 | 6 |
Cùng h₀, A tốt hơn 2 pp. Nếu mỗi checkpoint chọn recipe trên dev rồi chấm final evaluation mới, B có thể đạt 6 so với A 8. Fixed-recipe transfer và performance của search procedure là hai estimands. Chọn h₁ bằng test chỉ tạo exploratory/oracle result.
Phản ví dụ cho objective attribution: cùng objective nhưng khác data/filtering/compute cũng có thể tạo gap. Common downstream chưa chặn các upstream paths. Muốn kiểm objective phải định nghĩa factors giữ cố định; nếu tuning riêng, cần search budget và population.
5. Scope paper đã tự nêu#
PDF §3.4/§8 coi đây là hai released checkpoints dưới common recipe; objective attribution là hypothesis. Ba lower means trong bốn comparisons vẫn là evidence cụ thể. Footnote 4 trang 3: upstream Audio-JEPA đánh giá target encoder; detector dùng context encoder, chưa đo tác động của choice đó.
6. Tự kiểm#
- Giữ 128 tokens có giữ native geometry JEPA chưa?
- Cho mỗi encoder native input đã isolate objective chưa?
- Frozen contrast loại được explanation nào, còn confounds nào?
- Có thể gán current normalized loss cho exact released run không?
Đáp án và reasoning
- Chưa: grid 16×8→32×4; hop/fmax/duration/kernel aspect ratio cũng đổi.
- Không: contrast lúc đó là hai native system packages.
- Gap không chỉ xuất hiện do fine-tuning encoder; upstream/input/learned-head interactions vẫn còn.
- Chưa: code revision và file history chưa tạo run-to-checkpoint link.
Đào sâu: viết contract cho fixed recipe và matched search budget; nêu câu hỏi mỗi thiết kế chưa trả lời. Đây là học design, chưa lựa chọn experiment mới.