ajaudio / studyAUDIO-JEPA · RESEARCH NOTES
7 phút đọc · Toàn văn
Mục lục bài · 9 mục

Bài 4 — Reconstruction và contextualized latent targets#

Bắt đầu · Trước: speech units · Tiếp: self-distillation

Mục tiêu và tiền đề#

Bạn sẽ phân biệt decoder output với encoder information; tự tạo một contextualized target; đọc data2vec theo modality và version. Cần bài 2 về MSE/normalization và lớp 1 về token/attention. Trực giác: dự đoán “giá trị của patch” và “cách một teacher hiểu patch trong toàn clip” là hai đáp án khác, dù cùng có mask.

Notation: X∈RN×pX\in\mathbb R^{N\times p} là NN input patches, mỗi patch có pp scalar; encoder HC∈R∣C∣×dH_C\in\mathbb R^{|C|\times d}; decoder/predictor outputs ở MM là X^M∈R∣M∣×p\hat X_M\in\mathbb R^{|M|\times p} hoặc H^M∈R∣M∣×d\hat H_M\in\mathbb R^{|M|\times d}. Target-space dimension cần ghi; loss “trên spectrogram” không phải waveform reconstruction.

1. AudioMAE: visible-only encoder, input-space target#

AudioMAE chuyển log-mel thành patches, encoder xử lý visible patches; decoder chèn mask tokens/restores order để reconstruct masked spectrogram patches. Main pretraining dùng random 80% masking, MSE trên masked patches; target có lựa chọn patch normalization. Decoder được bỏ sau pretraining. Masked Autoencoders that Listen, §3 và §4.2/4.4, Appendix B.

Toy trace. Có bốn patches, hai scalar/patch:

X=[11243348],C={1,3},M={2,4}.X=\begin{bmatrix}1&1\\2&4\\3&3\\4&8\end{bmatrix},\quad C=\{1,3\},\quad M=\{2,4\}.

Encoder sees first/third rows only, có shape 2×d2\times d. Decoder predicts X^2=[2,2],X^4=[4,4]\hat X_2=[2,2],\hat X_4=[4,4]. Squared errors 4 và 16; mean trên hai patches và hai scalar là (4+16)/4=5(4+16)/4=5. Nếu loss mean trên patch vector norms thì 10. Hai conventions không thay nhau.

Một patch-normalized target cũng là target khác: [2,4][2,4] center mean 3, population std 1 → [−1,1][-1,1], còn [4,8][4,8] mean 6, std 2 → [−1,1][-1,1]. Target normalization đã bỏ offset/gain ở loss, nhưng encoder vẫn nhận visible patches theo front-end recipe. Không suy mọi latent invariant với gain từ target này.

Mask token ở decoder không có giá trị patch bị che; nó cho biết chỗ cần reconstruct. Encoder visible-only làm sequence ngắn hơn. Đây khác student full-length transformer nhận mask embeddings trong data2vec v1/speech.

2. Conditional mean không là định lý encoder quên artifact#

Giả sử một hidden target detail U∈{−1,+1}U\in\{-1,+1\} independent context, balanced. Squared-loss optimum prediction là 0 và expected error 1, như proof bài 2. Nhưng encoder có thể giữ một visible cue VV trong một direction mà decoder không dùng, vẫn cùng optimum. Hoặc encoder giữ VV vì nó giúp predict một target khác. Những solutions đó không bị loại bởi proof về UU.

Counterexample đủ cụ thể. Context là scalar V∈{−1,+1}V\in\{-1,+1\}, target UU independent. Encoder f1(V)=0f_1(V)=0, f2(V)=Vf_2(V)=V; decoder ở cả hai trả 0. Hai hệ có cùng expected reconstruction loss 1. Downstream label Y=VY=V: encoder thứ nhất vô ích, encoder thứ hai hoàn hảo. Do đó reconstruction optimum alone không identify encoder information.

Ngược lại, latent prediction không tự tránh averaging: nếu target teacher còn chứa một detail unpredictable từ context, MSE predictor cũng trả conditional mean của detail ấy khi teacher fixed trong step. Teacher learned có thể thay target qua training; phải xét dynamics, không lấy “latent” thay cho proof.

3. Contextualization: target ở vị trí j có thể chứa context khác#

Teacher self-attention nhận toàn input. Output ở một patch không đơn thuần là projection của riêng patch đó. Một toy tuyến tính giúp nhìn:

t2=0,75x1+0,25x2,x1=2, x2=10⇒t2=4.t_2=0,75x_1+0,25x_2,\qquad x_1=2,\ x_2=10\Rightarrow t_2=4.

Student thấy x1=2x_1=2, target position 2, không thấy x2x_2. Nếu luôn trả 0,75x1=1,50,75x_1=1,5, nó đã dự đoán được phần do context đóng góp; residual 2,5 là phần thiếu trong case này. Đây là toy attention mixture, không weights/behavior của data2vec hoặc JEPA thực.

Teacher full-input trong training tạo supervision từ data đã có; không phải sử dụng test labels. Tuy nhiên một recipe có thể để target chứa nhiều visible context hoặc position structure khiến prediction dễ theo một cách khác mong muốn. Câu đúng cần kiểm là “target chứa gì và context cần dùng gì?”, không mặc định kết luận leakage. Train/test contamination là câu hỏi dataset/protocol riêng.

4. data2vec: teacher features thay discrete labels#

Ở data2vec bản gốc, student nhận chuỗi đã mask; teacher nhận đầy đủ sample chưa mask. Target tại masked timesteps là average top K normalized teacher block features; teacher Transformer dùng EMA. Paper nói feature/positional encoders được shared; no-gradient teacher forward không ngăn shared weights đổi qua student update. Target thường lấy FFN output trước residual cuối của block. data2vec §§3.1–3.4.

Viết target:

tj=1K∑ℓ=L−K+1LNℓ(aj(ℓ)).t_j=\frac1K\sum_{\ell=L-K+1}^L \mathcal N_\ell(a_j^{(\ell)}).

aj(ℓ)∈Rda_j^{(\ell)}\in\mathbb R^d, KK là số block average, LL tổng blocks; N\mathcal N cần đọc axes. Speech paper dùng instance normalization theo sequence/sample cho từng feature dimension; vision/NLP dùng parameter-less layer normalization. Main formulation trình bày Smooth L1, speech setup §4.2 dùng simple L2, K=8K=8, EMA từ 0,999 tới 0,9999 trong 30.000 updates. Không gọi mọi data2vec loss là Smooth L1 hoặc mọi target normalization là LayerNorm. Speech setup §4.2.

Toy normalize-then-average. Hai teacher layers, mỗi layer có 3 steps và 1 feature: a(1)=[0,2,4]a^{(1)}=[0,2,4], a(2)=[10,10,14]a^{(2)}=[10,10,14]. Dùng population variance, epsilon 0 cho minh họa:

N(a(1))=[−3/2,0,+3/2],\mathcal N(a^{(1)})=[-\sqrt{3/2},0,+\sqrt{3/2}],
N(a(2))=[−1/2,−1/2,2].\mathcal N(a^{(2)})=[-1/\sqrt2,-1/\sqrt2,\sqrt2].

Target step 3 là (1,224745+1,414214)/2≈1,319479(1,224745+1,414214)/2\approx1,319479. Average raw layer tại step 3 là 9; layer lớn hơn chi phối. Normalize từng layer rồi average không bằng average rồi normalize. Toy dùng một feature chỉ để tính sequence normalization; không một tensor runtime data2vec.

Teacher output chứa continuous context, không unit ID. Matching head không cần phân biệt positive với negatives. Target normalization/teacher dynamics là recipe hỗ trợ training, không phải theorem EMA tự đủ. Cụ thể KK/norm/loss/data phải gắn checkpoint/version. data2vec 2.0 có thay efficiency/architecture; không chép recipe bản gốc vào bản 2.0 chỉ vì tên.

5. Một phép so sánh giữ đúng axes#

Câu hỏiAudioMAE core recipedata2vec bản gốc, speechAudio-JEPA recipe sẽ đọc ở bài 6
TargetMasked spectrogram valuesContextual teacher featuresContextual target patch features
Target có học cùng training không?Input values fixedEMA/shared componentsEMA target encoder
Student Transformer sequenceVisible patchesFull-length, masked embeddingsVisible patches
Prediction outputPatch scalarsFeature vector per masked stepFeature vector per target patch
Loss positionsMasked patchesMasked time stepsMasked patches
GradientEncoder + decoderStudent path; shared weights qua studentContext encoder + predictor

Đây là comparison recipe cốt lõi đã đọc, không mọi hybrid/variant. Bảng trên không xếp hạng forensic utility. Học input reconstruction tốt hoặc ASR tốt đều còn thiếu cầu nối tới nhãn detection của lớp 2.

6. Neo vào paper của mình#

Paper SOICT so hai released encoders dưới common downstream input/head/recipe. AudioMAE native sequence 1.024 time frames và Audio-JEPA 256 khác nhau; chuyển về 256 ở downstream đổi token/time geometry không đối xứng. Cùng AudioSet source chưa kiểm cùng data realization, steps, mask hay optimizer. Chương 06.

Để nói “objective latent giữ artifact tốt hơn reconstruction”, cần evidence isolate objective và đo cue/task, không chỉ EER gap giữa hai checkpoints. Toy mục 2 cũng cho biết low reconstruction error không identify duy nhất encoder information. Bài 10 sẽ tự review một causal claim.

Bài tập tăng dần#

  1. Với bốn-patch toy, encoder có bao nhiêu input rows? Decoder reconstruction target width là p hay d?
  2. Nếu X^2=[2,4],X^4=[4,7]\hat X_2=[2,4],\hat X_4=[4,7], mean scalar MSE là bao nhiêu?
  3. Hai encoder ở counterexample V/UV/U cùng optimum; downstream Y=VY=V khác ra sao? Proof conditional mean đã nói gì và chưa nói gì?
  4. Teacher target ở patch 2 có information từ patch 1: có được gọi “pure local target” không? Có tự là test-label leakage không?
  5. data2vec v1 speech và vision có bắt buộc cùng normalization/loss/settings không?
Đáp án
  1. Hai rows visible; target output width p=2, không latent width d.
  2. Chỉ một scalar error −1, tổng 1; mean trên 4 scalar là 0,25.
  3. f1f_1 không phân biệt V; f2f_2 giữ nhãn. Proof chỉ xác định optimal output estimate của U với context/distribution fixed, không xác định toàn hidden representation.
  4. Không pure local. Teacher contextualization là recipe supervision; leakage cần contamination/test information khác, không tự suy từ teacher full-input trong train.
  5. Không. Phải khóa modality/version; source §3 định nghĩa chung và §4 ghi settings cụ thể, có speech L2 khác formulation Smooth L1.

Đào sâu tự chọn#

Khai triển target covariance của average layers: covariance chứa cross-layer terms, nên average không tự tăng information hoặc “lọc noise” nếu layer noises/cues correlated. Không suy từ phép average thành semantics guarantee. Thử dựng hai tầng có cues đối dấu; chúng có thể triệt nhau ngay cả khi mỗi tầng riêng giữ cue.

DỪNG LẠI & TỰ KIỂM TRA

Bạn đã giải thích được cơ chế trong bài?

↓ Bản Markdown nguyên gốcGiữ nguyên nội dung · Công thức, bảng và nguồn đầy đủ.Các chat bàn giao được mở trong Codex.

Gõ từ khóa để tìm bài học và đoạn liên quan.