# Bài 6 — JEPA–AudioMAE: công bằng cho câu hỏi nào?

[Trước](05-DOC-BANG-KET-QUA.md) · [Tiếp](07-LICH-SU-THU-NGHIEM.md).

## Mục tiêu và tiền đề

Truy vết native input; phân biệt checkpoint transfer với tác động nhân quả của objective. Cần [continuous targets lớp 3](../lop-03-ssl-jepa/04-RECONSTRUCTION-CONTINUOUS-TARGET.md), [recipe/source](../lop-03-ssl-jepa/14-HO-SO-RECIPE.md), [fair contrast lớp 5](../lop-05-danh-gia-thuc-nghiem/08-SO-SANH-ABLATION-CO-CHE.md). Metadata/version/unknowns ở [dossier pipeline](research/pipeline-sources.md).

## 1. Targets upstream khác; loss downstream chung

[Audio-JEPA §III](https://arxiv.org/pdf/2507.02915v2) dự đoán continuous features do EMA teacher tạo. Teacher nhìn full spectrogram rồi lấy outputs ở masked positions. [AudioMAE §3](https://arxiv.org/html/2207.06405v2#S3) dùng decoder dự đoán patch values hoặc normalized values bằng masked MSE; decoder có local attention. Paper detector bỏ predictor/decoder, dùng encoder features và cùng back-end/weighted CE ở hai arms.

Khác target tạo incentives khác nhưng chưa tự xác định forensic information được giữ. Conditional mean có thể làm mượt output reconstruction; chưa kết luận mọi hidden feature của encoder mất detail. Latent targets cũng có thể bỏ detail. Phải đo đúng readout và decision.

## 2. Cùng token count chưa giữ geometry

| Setting/source | Input time×mel | Patch time×mel | Grid time×mel | Full N | Nominal time span |
|---|---|---|---|---:|---|
| JEPA native | 256×128; 10 s; 32 kHz | 16×16 | 16×8 | 128 | 16×39,0625=625 ms |
| AudioMAE native | 1024×128; 10 s; 16 kHz | 16×16 | 64×8 | 512 | 16×10=160 ms |
| Common downstream | 256×128; 2,56 s; 32 kHz sau toolkit 16 kHz | 8×32 | 32×4 | 128 | 8×10=80 ms |

Native JEPA config đặt hop=10000/256=39,0625 ms và window=2,5×hop=97,65625 ms; source fbank rồi pad/truncate. Các shape cố định chưa chứng minh có đúng từng ấy valid frames. N trong bảng là full grid **trước masking**, không phải số visible MAE tokens.

JEPA giữ N=128 nhưng grid 16×8 → 32×4. MAE giảm N=512 → 128. Kernel area vẫn 256 nhưng physical span/semantics đổi. Native JEPA fmax 16 kHz khác downstream 8 kHz; resampling không tạo high-band evidence mới. HF card ghi time/frequency grid đảo so với arithmetic và code: input time 256/mel 128, patch 16×16 phải cho **16 time ×8 mel**. [Pinned sources](research/pipeline-sources.md), PDF §§3.1/3.4 trang 4/6.

## 3. Factors chung và factors còn khác

| Axis | Chung trong paper comparison | Chưa isolate |
|---|---|---|
| Downstream | Input/head/subset/loss/epochs/LRs/seeds/scoring | Checkpoint×recipe interaction; chưa optimum riêng |
| Architecture | ViT-Base family, 12 blocks, width 768 | Exact artifact/encoder branch conventions |
| Data | Cùng nguồn AudioSet-2 M | Clip lists/filtering/sampling/exposures |
| Objective | Cố ý khác targets | Mask/normalization/optimizer/budget/native input cùng khác |
| Adaptation | Frozen/full ở các comparisons đã chạy | Frozen vẫn có learned head/input adaptation |

Upstream papers báo JEPA 100k updates/batch 256 và MAE 32 epochs/batch 512. Đây chưa là run manifests của exact released weights. Current JEPA config/loss có differences với paper; MAE release script ghi 33 epochs; HF JEPA card/config được cập nhật sau weight file. Giữ từng version riêng, chưa âm thầm chọn một làm recipe thật của weights. [Dossier commits](research/pipeline-sources.md).

## 4. Worked example: estimand khác có thể đảo thứ hạng

Toy dưới **không phải paper result**:

| Checkpoint | h₀ EER % | h₁ EER % |
|---|---:|---:|
| A | 8 | 12 |
| B | 10 | 6 |

Cùng h₀, A tốt hơn 2 pp. Nếu mỗi checkpoint chọn recipe trên dev rồi chấm final evaluation mới, B có thể đạt 6 so với A 8. Fixed-recipe transfer và performance của search procedure là hai estimands. Chọn h₁ bằng test chỉ tạo exploratory/oracle result.

**Phản ví dụ cho objective attribution:** cùng objective nhưng khác data/filtering/compute cũng có thể tạo gap. Common downstream chưa chặn các upstream paths. Muốn kiểm objective phải định nghĩa factors giữ cố định; nếu tuning riêng, cần search budget và population.

## 5. Scope paper đã tự nêu

PDF §3.4/§8 coi đây là hai released checkpoints dưới common recipe; objective attribution là hypothesis. Ba lower means trong bốn comparisons vẫn là evidence cụ thể. Footnote 4 trang 3: upstream Audio-JEPA đánh giá target encoder; detector dùng context encoder, chưa đo tác động của choice đó.

## 6. Tự kiểm

1. Giữ 128 tokens có giữ native geometry JEPA chưa?
2. Cho mỗi encoder native input đã isolate objective chưa?
3. Frozen contrast loại được explanation nào, còn confounds nào?
4. Có thể gán current normalized loss cho exact released run không?

<details>
<summary>Đáp án và reasoning</summary>

1. Chưa: grid 16×8→32×4; hop/fmax/duration/kernel aspect ratio cũng đổi.
2. Không: contrast lúc đó là hai native system packages.
3. Gap không chỉ xuất hiện do fine-tuning encoder; upstream/input/learned-head interactions vẫn còn.
4. Chưa: code revision và file history chưa tạo run-to-checkpoint link.

</details>

**Đào sâu:** viết contract cho fixed recipe và matched search budget; nêu câu hỏi mỗi thiết kế chưa trả lời. Đây là học design, chưa lựa chọn experiment mới.
