ajaudio / studyAUDIO-JEPA · RESEARCH NOTES
6 phút đọc · Toàn văn
Mục lục bài · 7 mục

Bài 8 — Fair comparison, ablation và bằng chứng cơ chế#

Trước · Tiếp.

Mục tiêu và tiền đề#

Chọn contrast theo estimand, viết competing explanations và phân biệt decodability/reliance/mechanism. Cần JEPA recipe, probes/diagnostics, selection. Các causal diagrams dưới là toy assumptions, không diagnosis pretrained checkpoint.

1. “Giữ công bằng” chưa định nghĩa một contrast duy nhất#

Cho M(c,h)M(c,h) là metric checkpoint c trong downstream recipe h. Fixed-recipe contrast M(cA,h0)−M(cB,h0)M(c_A,h_0)-M(c_B,h_0) đo checkpoint-system transfer ở recipe chung. Model-specific search H\mathcal H với budgetK đo M(cA,h^A(K))−M(cB,h^B(K))M(c_A,\widehat h_A(K))-M(c_B,\widehat h_B(K)) trên final eval mới: attainable performance tìm được dưới search procedure, không global optimum. Search cost include failed trials/data feedback, không chỉ run winner. Cawley–Talbot 2010 §§4–5, Bouthillier 2021 §3 cho selection/pipeline variance basis.

Worked toy:

CheckpointRecipe h1 EER%Recipe h2 EER%
A812
B106

Fixedh 1: A tốt hơn 2 pp. Search cảh 1/h2: B đạt 6 so A8, B tốt hơn 2 pp. Không contradiction. Dùng test chọn h rồi báo minima là oracle/exploratory; dùng dev chọn h và test mới mới đánh giá selected systems. “Common recipe” có thể disadvantage một checkpoint, vẫn trả lời valid transfer question; không claim model-specific optimum.

2. Objective effect cần intervention cụ thể#

Released JEPA/MAE còn khác masking/optimizer/input/budget/data realization, nên objective O đi cùng recipe R; observed contrast không isolate O. Controlled experiment có thể gán O và giữ architecture/data/updates/optimizer policy/downstream fixed. Nó đo effect của objective implementation trong điều kiện đó. Objective tương tác optimization/mask; nếu một objective cần riêng tuning thì fixed recipe contrast có scope khác tuned-procedure contrast. Không có “pure objective effect” vô điều kiện về mọi scale/recipe.

flowchart LR
 O[Objective O] --> Z[Representation Z]
 R[Pretraining recipe và data R] --> Z
 Z --> H[Head và decision]
 D[Downstream data và tuning D] --> H
 H --> M[Metric trong eval population]
 E[Eval mix và threshold policy E] --> M
Xem mã sơ đồ
flowchart LR
 O[Objective O] --> Z[Representation Z]
 R[Pretraining recipe và data R] --> Z
 Z --> H[Head và decision]
 D[Downstream data và tuning D] --> H
 H --> M[Metric trong eval population]
 E[Eval mix và threshold policy E] --> M

Trong observational checkpoint comparison, O và R không independently assigned. Giữ common H/D/E chặn một số paths, không chặn R→Z. Pair data order/mask draws chỉ giảm noise khi thật sự controlled và semantically comparable; số seed giống không đảm bảo paths giống. Sơ đồ bỏ nhiều yếu tố, không đủ suy identify causal effect nếu không assumptions/design.

3. Observation → decodability → reliance#

Complete toy mechanism case. Representation z=(a,c)z=(a,c) chứa artifact cue a∈{−1,+1} và source cue c∈{−1,+1}. B label tương ứng a=+1;S a=−1. Source balanced độc lập a. Hai heads:

h1(z)=a,h2(z)=a+2c,h_1(z)=a,\qquad h_2(z)=a+2c,

accept B nếu score≥0. Source probe g(z)=cg(z)=c accuracy 100% cho cả hai encoder/readouts. h1 detection accuracy 100%; h2 đúng khi c cùng sign a, sai khi opposite: accuracy 50%. Source decodable trong both, nhưng h1 không reliance vào c.

Can thiệp toy do(c←−c)do(c\leftarrow-c) giữ a: h1 score không đổi; h2 score đổi−4 c, mọi decisions đổi trong bốn combinations. Đây là reliance của h2 vào coordinate c trong toy. Trong audio thật, không có knob c sạch sẵn: re-recording/codec/denoising có thể đổi artifact a, content hoặc out-of-distribution quality đồng thời. Score đổi sau perturbation chưa đặc hiệu source reliance.

Để phân biệt competing explanations, thiết kế manipulation checks: intervention có giữ content/label/artifact đủ không; có matched perturbation control cùng severity nhưng không đổi proposed cue không; effect lặp qua relevant generators/sources hay chỉ một channel? Nếu source-related cue change cùng other nuisance, report effect của transformation, chưa effect riêng cue. Hewitt–Liang 2019 §§2–3 dùng control tasks giới hạn probe-capacity interpretation; independence/reliance toy ở đây là proof tự biên soạn, không empirical audio claim.

4. Inference ablation và retraining ablation#

Với learned system ftrainedf_{trained}, bỏ component/cue ở inference đo sensitivity/reliance của hệ đang có vào intervention. Retrain không component đo learning procedure có thể thích nghi và đạt gì khi resource ấy không còn. Hai estimands khác.

Toy redundant representation z=(a,a), head chỉ đọc coordinate 1. Bỏ coordinate 1 ở inference về 0 làm decisions mất discrimination. Retrain head đọc coordinate 2 khôi phục 100%. Inference necessity với head hiện tại không nghĩa information resource 1 cần cho mọi retrained system. Reverse case: removing a cue improves OOD ranking có thể vì regularization/distribution shift; không tự chứng minh correct causal mechanism.

Module ablation còn có parameter capacity/training time/optimization confounds. Matched-capacity dummy module hoặc simpler baseline giúp kiểm competing explanations “extra capacity”, “more compute”, “changed preprocessing”. Không yêu cầu mọi ablation giữ mọi đại lượng đồng thời; chọn quantity chính rồi ghi tradeoffs. Attention/t-SNE/CKA là descriptions/diagnostics, không đủ prove causal reliance. Randomization controls cho saliency có thể kiểm explanation dependence vào weights/data, vẫn không chứng minh intervention specificity. Không dạy saliency theorem khi chưa đọc method phù hợp.

5. Neo paper và viết claim#

Paper Table 4: released checkpoints common downstream recipe, JEPA lower mean ở3/4 comparisons. PDF §3.4/§8 nói pretraining recipes còn khác. Frozen head là learned layers+ASP+nonlinear MLP, không linear probe; frozen contrast vẫn có head optimization/capacity interactions. PDF discussion hypothesis artifact retention chưa đo trực tiếp predictability/retention. Chương 06 §§4–5.

Claim có scope: “Checkpoint A chuyển giao tốt hơn B trong recipe h1 trên corpus X theo EER, với run variation được báo.” Muốn “latent prediction giữ forensic artifacts hơn reconstruction” cần cue definition/manipulation/representation + decision evidence, matched pretraining contrasts và competing explanations. Không đi từ t-SNE đẹp trực tiếp tới sentence ấy.

6. Tự kiểm và nhánh sâu#

  1. Source probe 100% có chứng minh h1 dùng c không? Tính score change khi flip c.
  2. Vì sao h1-vs-h2 ranking trong recipe table có thể đảo sau tuning?
  3. Bỏ coordinate 1 làm inference fail có chứng minh retraining luôn fail không?
  4. Objective switch cùng mask switch support claim nào?
Đáp án reasoning
  1. Không; h1=a, flip c không đổi output, Δscore=0. Probe có readout khác head.
  2. Metric là checkpoint×recipe interaction; fixedh 1 và selected h trả lời khác estimands, selection phải dev/test separated.
  3. Không; coordinate 2 có redundant a, retrained head đọc được. Inference intervention giữ learned function cố định.
  4. Effect của objective+mask package trong scope ấy; không riêng objective nếu không disentangle hoặc gán rõ interaction design.

Đào sâu: tự lập 2×2 factorial objective×mask toy, định nghĩa main effects và interaction bằng contrasts. Chốt population/budget trước; một pooled main effect không đảm bảo từng setting cùng sign. Đây là học thiết kế, chưa chọn hướng nghiên cứu.

DỪNG LẠI & TỰ KIỂM TRA

Bạn đã giải thích được cơ chế trong bài?

↓ Bản Markdown nguyên gốcGiữ nguyên nội dung · Công thức, bảng và nguồn đầy đủ.Các chat bàn giao được mở trong Codex.

Gõ từ khóa để tìm bài học và đoạn liên quan.