ajaudio / studyAUDIO-JEPA · RESEARCH NOTES
5 phút đọc · Toàn văn
Mục lục bài · 7 mục

Bài 8 — Hypotheses: bằng chứng nào phân biệt cơ chế?#

Trước · Tiếp.

Mục tiêu và tiền đề#

Biến explanation thành predictions có thể phân biệt; đọc diagnostics đúng đơn vị. Cần collapse lớp 3, masking, error scoring, probe/reliance lớp 4, mechanism lớp 5. Evidence contracts dưới phục vụ học reasoning, chưa là tasks triển khai được chọn.

1. Source association có competing explanations#

P §6 trang 11: hai VCTK-source corpora có EER thấp hơn ba corpora có sources khác. Candidates gồm source/channel similarity, generator/language/content/quality, input/head interaction và exposure chưa audit. AASIST cũng train VCTK nhưng DFADD 41,87; shared source chưa đảm bảo error thấp. LibriSeVoc clean nhưng JEPA error cao; clean/noisy riêng chưa giải thích hết pattern.

Toy h=(a,c), artifact cue a=±1 và source c=±1, bốn combinations cân bằng. Source probe g=c đạt 100%. Head1 d=a không dùng c; head2 d=a+2c chỉ đúng 50%. Flip c giữ a: Δd₁=0, Δd₂=−4c. Cùng decodability, khác reliance. Với audio thật, source manipulation còn có thể đổi channel/content/artifact; cần matched controls và manipulation checks. Probe/t-SNE chưa đo causal reliance riêng.

2. Random frozen gần pretrained đo gì?#

H v38: DFADD random 8,66±2,25 vs JEPA 8,47±0,35, gap 0,19 pp chưa là equivalence test. Random vẫn nhận mel input, có patch projection/positions/nonlinear Transformer và learned back-end≈0,496 M. Inductive bias có thể giữ cues cho head. Pretraining cũng có thể giữ thông tin head/task chưa khai thác hoặc ảnh hưởng optimization stability.

Random full FT 13,68 n=1 so JEPA 7,39 n=3 chưa đủ kết luận pretraining chỉ hữu ích khi fine-tuning. Similar downstream EER chưa chứng minh representation collapse. Cần matched arms, head capacity/budget/selection và readout/decision evidence. History.

3. Giải đúng effective rank 263,1 ở v41#

C: collapse_metrics.py, v41 G4/G4b và replacement. H[B,N,D] là final encoder token outputs; B=64,N=128,D=768. Code reshape BN=8.192 rows, center raw matrix rồi SVD, chưa mean-pool thành 64 utterance rows.

pj=σj/∑kσk,reff=exp⁡(−∑jpjlog⁡pj).p_j=\sigma_j/\sum_k\sigma_k,\qquad r_{eff}=\exp(-\sum_jp_j\log p_j).

Centered token rank≤min(8192−1,768)=768; 263,1 không vi phạm bound 63. Bound≤63 chỉ áp dụng khi center 64 pooled utterance rows. rank_raw dùng normalized token rows không center; khác cả normalization, chưa quy toàn difference cho centering.

Statistic CXử lýChưa tự kết luận
rank_effRaw token rows, center, entropy trên σUseful utterance-level rank
rank_rawL2-normalized token rows, không centerCausal contribution riêng của mean
std_tokenCoordinate std qua normalized tokensContent hay position tạo variation
cos_crossToken pairs thuộc khác utterancesDetector reliance vào cue
std_uttMean raw tokens; normalize từng utterance; std qua batchAccessibility qua nonlinear head

Toy σ=[2,1]: entropy trên σ cho erank≈1,8899; trên σ² cho≈1,6494. Hai definitions khác nhau. Entropy rank bỏ scale: tiny centered noise vẫn có thể high rank.

H notes: native/detector rank 263,1/247,9; cos 0,6793/0,7192; std_utt 0,0061/0,0060. Chưa đo mới ở đây. G4b lọc clips≥10 s, G4≥2,56 s và mở streams riêng; chưa khóa paired IDs. Source cast 16 kHz rồi resample 32 kHz, nên fmax 16 kHz không khôi phục content trên 8 kHz. Chưa gọi comparison này isolated causal input test.

Phản ví dụ: h+=(100,1), h−=(100,−1) có cosine 9999/10001≈0,9998 nhưng coordinate 2 phân loại hoàn hảo. Tokens có thể đa dạng theo positions trong khi pooled utterances giống nhau. Anisotropy/low variance trên speech batch/readout chưa chứng minh global checkpoint collapse hoặc nguyên nhân EER.

4. Mask adjacency khác learned shortcut#

H notes v41 báo≈89,9% targets còn visible four-neighbor trong sampled masks. C code đo geometry, chưa đọc learner strategy. Toy cố định 64 masked trong 128 tokens, grid 32×4: conditional probability all k neighbors cũng masked là C(63,k)/C(127,k). Có 4 corners k=2, 64 noncorner edges k=3, 60 interiors k=4. Weighted mean cho ≈90,58% có ít nhất một visible neighbor. Đây là combinatorial toy, chưa replicate random ratio 0,4–0,6 của notes.

Nearby context có thể hỗ trợ interpolation; chưa chứng minh model chỉ dùng local shortcut. Adjacent noise có thể không predict target; cue có thể cần long context. Structured mask cũng có thể bỏ context hữu ích. Cần learner/perturbation controls để nói strategy.

5. Output averaging và predictor compatibility#

Toy target t=±1 equally likely, context không có info: prediction 0 có expected MSE 1; sample±1 độc lập có MSE 2. Đây là optimal output prediction dưới fixed target distribution; chưa nói mọi hidden features mất detail. PDF §6 trang9–10 dùng lý thuyết này làm motivation, chưa đo cue retention.

Predictor g học với coordinates/positions của source checkpoint. Interpolation, positions mới hoặc context fine-tuning có thể đổi interface. Toy rotate h→Qh: classifier có thể reparameterize để giữ decisions nhưng fixed predictor có thể đổi error. Genuine OOD/noisy cũng có thể khó predict hơn clean fake. Prediction error chưa tự là spoof probability/LLR. Error scoring lớp 3.

6. Tự kiểm#

  1. 263,1 với 64 clips có phải sửa xuống≤63 không? Nêu rows.
  2. Source probe 100% đã chứng minh reliance chưa?
  3. Adjacency≈90% đã chứng minh model chỉ nội suy local chưa?
  4. Random gần JEPA đã chứng minh pretraining vô ích chưa?
  5. Nêu evidence cho hypothesis “JEPA giữ cue X tốt hơn MAE”.
Đáp án và reasoning
  1. Không:64×128 token rows, bound 768. Giữ statistic kèm đơn vị/provenance.
  2. Chưa: head1 không đọc c dù probe đọc được c. Cần decision interventions/checks.
  3. Chưa: available geometry khác predictability và learned strategy.
  4. Chưa: uncertainty/equivalence, inductive bias, learned head và recipe còn ảnh hưởng.
  5. Define cue; đo retention tại readout; đo decision reliance/transfer; matched pretraining contrast và kiểm label/các cues khác.

Đào sâu: chọn hypothesis; viết observation có thể bác bỏ nó và explanation vẫn còn phù hợp. Đây là học evidence reasoning, chưa ưu tiên project mới.

DỪNG LẠI & TỰ KIỂM TRA

Bạn đã giải thích được cơ chế trong bài?

↓ Bản Markdown nguyên gốcGiữ nguyên nội dung · Công thức, bảng và nguồn đầy đủ.Các chat bàn giao được mở trong Codex.

Gõ từ khóa để tìm bài học và đoạn liên quan.