# Bài 4 — Reconstruction và contextualized latent targets

[Bắt đầu](00-BAT-DAU-LOP-03.md) · Trước: [speech units](03-SPEECH-UNITS-CONTRASTIVE.md) · Tiếp: [self-distillation](05-SELF-DISTILLATION-ASYMMETRY.md)

## Mục tiêu và tiền đề

Bạn sẽ phân biệt decoder output với encoder information; tự tạo một contextualized target; đọc data2vec theo modality và version. Cần bài 2 về MSE/normalization và lớp 1 về token/attention. Trực giác: dự đoán “giá trị của patch” và “cách một teacher hiểu patch trong toàn clip” là hai đáp án khác, dù cùng có mask.

Notation: $X\in\mathbb R^{N\times p}$ là $N$ input patches, mỗi patch có $p$ scalar; encoder $H_C\in\mathbb R^{|C|\times d}$; decoder/predictor outputs ở $M$ là $\hat X_M\in\mathbb R^{|M|\times p}$ hoặc $\hat H_M\in\mathbb R^{|M|\times d}$. Target-space dimension cần ghi; loss “trên spectrogram” không phải waveform reconstruction.

## 1. AudioMAE: visible-only encoder, input-space target

AudioMAE chuyển log-mel thành patches, encoder xử lý visible patches; decoder chèn mask tokens/restores order để reconstruct masked spectrogram patches. Main pretraining dùng random 80% masking, MSE trên masked patches; target có lựa chọn patch normalization. Decoder được bỏ sau pretraining. [*Masked Autoencoders that Listen*, §3 và §4.2/4.4, Appendix B](https://arxiv.org/html/2207.06405v3).

**Toy trace.** Có bốn patches, hai scalar/patch:

$$X=\begin{bmatrix}1&1\\2&4\\3&3\\4&8\end{bmatrix},\quad C=\{1,3\},\quad M=\{2,4\}.$$

Encoder sees first/third rows only, có shape $2\times d$. Decoder predicts $\hat X_2=[2,2],\hat X_4=[4,4]$. Squared errors 4 và 16; mean trên hai patches và hai scalar là $(4+16)/4=5$. Nếu loss mean trên patch **vector norms** thì 10. Hai conventions không thay nhau.

Một patch-normalized target cũng là target khác: $[2,4]$ center mean 3, population std 1 → $[-1,1]$, còn $[4,8]$ mean 6, std 2 → $[-1,1]$. Target normalization đã bỏ offset/gain ở loss, nhưng encoder vẫn nhận visible patches theo front-end recipe. Không suy mọi latent invariant với gain từ target này.

Mask token ở decoder không có giá trị patch bị che; nó cho biết chỗ cần reconstruct. Encoder visible-only làm sequence ngắn hơn. Đây khác student full-length transformer nhận mask embeddings trong data2vec v1/speech.

## 2. Conditional mean không là định lý encoder quên artifact

Giả sử một hidden target detail $U\in\{-1,+1\}$ independent context, balanced. Squared-loss optimum prediction là 0 và expected error 1, như proof bài 2. Nhưng encoder có thể giữ một visible cue $V$ trong một direction mà decoder không dùng, vẫn cùng optimum. Hoặc encoder giữ $V$ vì nó giúp predict một target khác. Những solutions đó không bị loại bởi proof về $U$.

**Counterexample đủ cụ thể.** Context là scalar $V\in\{-1,+1\}$, target $U$ independent. Encoder $f_1(V)=0$, $f_2(V)=V$; decoder ở cả hai trả 0. Hai hệ có cùng expected reconstruction loss 1. Downstream label $Y=V$: encoder thứ nhất vô ích, encoder thứ hai hoàn hảo. Do đó reconstruction optimum alone không identify encoder information.

Ngược lại, latent prediction không tự tránh averaging: nếu target teacher còn chứa một detail unpredictable từ context, MSE predictor cũng trả conditional mean của detail ấy khi teacher fixed trong step. Teacher learned có thể thay target qua training; phải xét dynamics, không lấy “latent” thay cho proof.

## 3. Contextualization: target ở vị trí j có thể chứa context khác

Teacher self-attention nhận toàn input. Output ở một patch không đơn thuần là projection của riêng patch đó. Một toy tuyến tính giúp nhìn:

$$t_2=0,75x_1+0,25x_2,\qquad x_1=2,\ x_2=10\Rightarrow t_2=4.$$

Student thấy $x_1=2$, target position 2, không thấy $x_2$. Nếu luôn trả $0,75x_1=1,5$, nó đã dự đoán được phần do context đóng góp; residual 2,5 là phần thiếu trong case này. Đây là **toy attention mixture**, không weights/behavior của data2vec hoặc JEPA thực.

Teacher full-input trong training tạo supervision từ data đã có; không phải sử dụng test labels. Tuy nhiên một recipe có thể để target chứa nhiều visible context hoặc position structure khiến prediction dễ theo một cách khác mong muốn. Câu đúng cần kiểm là “target chứa gì và context cần dùng gì?”, không mặc định kết luận leakage. Train/test contamination là câu hỏi dataset/protocol riêng.

## 4. data2vec: teacher features thay discrete labels

Ở data2vec bản gốc, student nhận chuỗi đã mask; teacher nhận đầy đủ sample chưa mask. Target tại masked timesteps là average **top K normalized teacher block features**; teacher Transformer dùng EMA. Paper nói feature/positional encoders được shared; no-gradient teacher forward không ngăn shared weights đổi qua student update. Target thường lấy FFN output trước residual cuối của block. [data2vec §§3.1–3.4](https://arxiv.org/pdf/2202.03555).

Viết target:

$$t_j=\frac1K\sum_{\ell=L-K+1}^L \mathcal N_\ell(a_j^{(\ell)}).$$

$a_j^{(\ell)}\in\mathbb R^d$, $K$ là số block average, $L$ tổng blocks; $\mathcal N$ cần đọc **axes**. Speech paper dùng instance normalization theo sequence/sample cho từng feature dimension; vision/NLP dùng parameter-less layer normalization. Main formulation trình bày Smooth L1, **speech setup §4.2 dùng simple L2**, $K=8$, EMA từ 0,999 tới 0,9999 trong 30.000 updates. Không gọi mọi data2vec loss là Smooth L1 hoặc mọi target normalization là LayerNorm. [Speech setup §4.2](https://arxiv.org/pdf/2202.03555).

**Toy normalize-then-average.** Hai teacher layers, mỗi layer có 3 steps và 1 feature: $a^{(1)}=[0,2,4]$, $a^{(2)}=[10,10,14]$. Dùng population variance, epsilon 0 cho minh họa:

$$\mathcal N(a^{(1)})=[-\sqrt{3/2},0,+\sqrt{3/2}],$$
$$\mathcal N(a^{(2)})=[-1/\sqrt2,-1/\sqrt2,\sqrt2].$$

Target step 3 là $(1,224745+1,414214)/2\approx1,319479$. Average raw layer tại step 3 là 9; layer lớn hơn chi phối. **Normalize từng layer rồi average không bằng average rồi normalize**. Toy dùng một feature chỉ để tính sequence normalization; không một tensor runtime data2vec.

Teacher output chứa continuous context, không unit ID. Matching head không cần phân biệt positive với negatives. Target normalization/teacher dynamics là recipe hỗ trợ training, không phải theorem EMA tự đủ. Cụ thể $K$/norm/loss/data phải gắn checkpoint/version. data2vec 2.0 có thay efficiency/architecture; không chép recipe bản gốc vào bản 2.0 chỉ vì tên.

## 5. Một phép so sánh giữ đúng axes

| Câu hỏi | AudioMAE core recipe | data2vec bản gốc, speech | Audio-JEPA recipe sẽ đọc ở bài 6 |
|---|---|---|---|
| Target | Masked spectrogram values | Contextual teacher features | Contextual target patch features |
| Target có học cùng training không? | Input values fixed | EMA/shared components | EMA target encoder |
| Student Transformer sequence | Visible patches | Full-length, masked embeddings | Visible patches |
| Prediction output | Patch scalars | Feature vector per masked step | Feature vector per target patch |
| Loss positions | Masked patches | Masked time steps | Masked patches |
| Gradient | Encoder + decoder | Student path; shared weights qua student | Context encoder + predictor |

Đây là comparison recipe cốt lõi đã đọc, không mọi hybrid/variant. Bảng trên không xếp hạng forensic utility. Học input reconstruction tốt hoặc ASR tốt đều còn thiếu cầu nối tới nhãn detection của lớp 2.

## 6. Neo vào paper của mình

Paper SOICT so hai **released encoders** dưới common downstream input/head/recipe. AudioMAE native sequence 1.024 time frames và Audio-JEPA 256 khác nhau; chuyển về 256 ở downstream đổi token/time geometry không đối xứng. Cùng AudioSet source chưa kiểm cùng data realization, steps, mask hay optimizer. [Chương 06](../06-GIAI-PHAU-PAPER.md).

Để nói “objective latent giữ artifact tốt hơn reconstruction”, cần evidence isolate objective và đo cue/task, không chỉ EER gap giữa hai checkpoints. Toy mục 2 cũng cho biết low reconstruction error không identify duy nhất encoder information. Bài 10 sẽ tự review một causal claim.

## Bài tập tăng dần

1. Với bốn-patch toy, encoder có bao nhiêu input rows? Decoder reconstruction target width là p hay d?
2. Nếu $\hat X_2=[2,4],\hat X_4=[4,7]$, mean scalar MSE là bao nhiêu?
3. Hai encoder ở counterexample $V/U$ cùng optimum; downstream $Y=V$ khác ra sao? Proof conditional mean đã nói gì và chưa nói gì?
4. Teacher target ở patch 2 có information từ patch 1: có được gọi “pure local target” không? Có tự là test-label leakage không?
5. data2vec v1 speech và vision có bắt buộc cùng normalization/loss/settings không?

<details>
<summary>Đáp án</summary>

1. Hai rows visible; target output width p=2, không latent width d.
2. Chỉ một scalar error −1, tổng 1; mean trên 4 scalar là 0,25.
3. $f_1$ không phân biệt V; $f_2$ giữ nhãn. Proof chỉ xác định optimal output estimate của U với context/distribution fixed, không xác định toàn hidden representation.
4. Không pure local. Teacher contextualization là recipe supervision; leakage cần contamination/test information khác, không tự suy từ teacher full-input trong train.
5. Không. Phải khóa modality/version; source §3 định nghĩa chung và §4 ghi settings cụ thể, có speech L2 khác formulation Smooth L1.

</details>

## Đào sâu tự chọn

Khai triển target covariance của average layers: covariance chứa cross-layer terms, nên average không tự tăng information hoặc “lọc noise” nếu layer noises/cues correlated. Không suy từ phép average thành semantics guarantee. Thử dựng hai tầng có cues đối dấu; chúng có thể triệt nhau ngay cả khi mỗi tầng riêng giữ cue.
