Bài 8 — Hypotheses: bằng chứng nào phân biệt cơ chế?#
Mục tiêu và tiền đề#
Biến explanation thành predictions có thể phân biệt; đọc diagnostics đúng đơn vị. Cần collapse lớp 3, masking, error scoring, probe/reliance lớp 4, mechanism lớp 5. Evidence contracts dưới phục vụ học reasoning, chưa là tasks triển khai được chọn.
1. Source association có competing explanations#
P §6 trang 11: hai VCTK-source corpora có EER thấp hơn ba corpora có sources khác. Candidates gồm source/channel similarity, generator/language/content/quality, input/head interaction và exposure chưa audit. AASIST cũng train VCTK nhưng DFADD 41,87; shared source chưa đảm bảo error thấp. LibriSeVoc clean nhưng JEPA error cao; clean/noisy riêng chưa giải thích hết pattern.
Toy h=(a,c), artifact cue a=±1 và source c=±1, bốn combinations cân bằng. Source probe g=c đạt 100%. Head1 d=a không dùng c; head2 d=a+2c chỉ đúng 50%. Flip c giữ a: Δd₁=0, Δd₂=−4c. Cùng decodability, khác reliance. Với audio thật, source manipulation còn có thể đổi channel/content/artifact; cần matched controls và manipulation checks. Probe/t-SNE chưa đo causal reliance riêng.
2. Random frozen gần pretrained đo gì?#
H v38: DFADD random 8,66±2,25 vs JEPA 8,47±0,35, gap 0,19 pp chưa là equivalence test. Random vẫn nhận mel input, có patch projection/positions/nonlinear Transformer và learned back-end≈0,496 M. Inductive bias có thể giữ cues cho head. Pretraining cũng có thể giữ thông tin head/task chưa khai thác hoặc ảnh hưởng optimization stability.
Random full FT 13,68 n=1 so JEPA 7,39 n=3 chưa đủ kết luận pretraining chỉ hữu ích khi fine-tuning. Similar downstream EER chưa chứng minh representation collapse. Cần matched arms, head capacity/budget/selection và readout/decision evidence. History.
3. Giải đúng effective rank 263,1 ở v41#
C: collapse_metrics.py, v41 G4/G4b và replacement. H[B,N,D] là final encoder token outputs; B=64,N=128,D=768. Code reshape BN=8.192 rows, center raw matrix rồi SVD, chưa mean-pool thành 64 utterance rows.
Centered token rank≤min(8192−1,768)=768; 263,1 không vi phạm bound 63. Bound≤63 chỉ áp dụng khi center 64 pooled utterance rows. rank_raw dùng normalized token rows không center; khác cả normalization, chưa quy toàn difference cho centering.
| Statistic C | Xử lý | Chưa tự kết luận |
|---|---|---|
| rank_eff | Raw token rows, center, entropy trên σ | Useful utterance-level rank |
| rank_raw | L2-normalized token rows, không center | Causal contribution riêng của mean |
| std_token | Coordinate std qua normalized tokens | Content hay position tạo variation |
| cos_cross | Token pairs thuộc khác utterances | Detector reliance vào cue |
| std_utt | Mean raw tokens; normalize từng utterance; std qua batch | Accessibility qua nonlinear head |
Toy σ=[2,1]: entropy trên σ cho erank≈1,8899; trên σ² cho≈1,6494. Hai definitions khác nhau. Entropy rank bỏ scale: tiny centered noise vẫn có thể high rank.
H notes: native/detector rank 263,1/247,9; cos 0,6793/0,7192; std_utt 0,0061/0,0060. Chưa đo mới ở đây. G4b lọc clips≥10 s, G4≥2,56 s và mở streams riêng; chưa khóa paired IDs. Source cast 16 kHz rồi resample 32 kHz, nên fmax 16 kHz không khôi phục content trên 8 kHz. Chưa gọi comparison này isolated causal input test.
Phản ví dụ: h+=(100,1), h−=(100,−1) có cosine 9999/10001≈0,9998 nhưng coordinate 2 phân loại hoàn hảo. Tokens có thể đa dạng theo positions trong khi pooled utterances giống nhau. Anisotropy/low variance trên speech batch/readout chưa chứng minh global checkpoint collapse hoặc nguyên nhân EER.
4. Mask adjacency khác learned shortcut#
H notes v41 báo≈89,9% targets còn visible four-neighbor trong sampled masks. C code đo geometry, chưa đọc learner strategy. Toy cố định 64 masked trong 128 tokens, grid 32×4: conditional probability all k neighbors cũng masked là C(63,k)/C(127,k). Có 4 corners k=2, 64 noncorner edges k=3, 60 interiors k=4. Weighted mean cho ≈90,58% có ít nhất một visible neighbor. Đây là combinatorial toy, chưa replicate random ratio 0,4–0,6 của notes.
Nearby context có thể hỗ trợ interpolation; chưa chứng minh model chỉ dùng local shortcut. Adjacent noise có thể không predict target; cue có thể cần long context. Structured mask cũng có thể bỏ context hữu ích. Cần learner/perturbation controls để nói strategy.
5. Output averaging và predictor compatibility#
Toy target t=±1 equally likely, context không có info: prediction 0 có expected MSE 1; sample±1 độc lập có MSE 2. Đây là optimal output prediction dưới fixed target distribution; chưa nói mọi hidden features mất detail. PDF §6 trang9–10 dùng lý thuyết này làm motivation, chưa đo cue retention.
Predictor g học với coordinates/positions của source checkpoint. Interpolation, positions mới hoặc context fine-tuning có thể đổi interface. Toy rotate h→Qh: classifier có thể reparameterize để giữ decisions nhưng fixed predictor có thể đổi error. Genuine OOD/noisy cũng có thể khó predict hơn clean fake. Prediction error chưa tự là spoof probability/LLR. Error scoring lớp 3.
6. Tự kiểm#
- 263,1 với 64 clips có phải sửa xuống≤63 không? Nêu rows.
- Source probe 100% đã chứng minh reliance chưa?
- Adjacency≈90% đã chứng minh model chỉ nội suy local chưa?
- Random gần JEPA đã chứng minh pretraining vô ích chưa?
- Nêu evidence cho hypothesis “JEPA giữ cue X tốt hơn MAE”.
Đáp án và reasoning
- Không:64×128 token rows, bound 768. Giữ statistic kèm đơn vị/provenance.
- Chưa: head1 không đọc c dù probe đọc được c. Cần decision interventions/checks.
- Chưa: available geometry khác predictability và learned strategy.
- Chưa: uncertainty/equivalence, inductive bias, learned head và recipe còn ảnh hưởng.
- Define cue; đo retention tại readout; đo decision reliance/transfer; matched pretraining contrast và kiểm label/các cues khác.
Đào sâu: chọn hypothesis; viết observation có thể bác bỏ nó và explanation vẫn còn phù hợp. Đây là học evidence reasoning, chưa ưu tiên project mới.