# Bài 4 — Layer fusion và pooling: gộp thông tin nào, mất thông tin nào?

[Điểm vào](00-BAT-DAU-LOP-04.md) · Trước: [adaptation](03-ADAPTATION-GRADIENT-PEFT.md) · Tiếp: [loss/sampling](05-LOSS-SAMPLING-OPTIMIZATION.md).

## Mục tiêu và tiền đề

Tính readout có padding/outlier, phân biệt token/time/channel axes, và giải thích vì sao attention weights chưa là causal importance. Cần [mean/variance và geometry](../lop-01-nen-tang/07-HINH-HOC-BIEU-DIEN.md), [attention](../lop-01-nen-tang/08-MANG-NORON-ATTENTION-TOI-UU.md), [token geometry](../lop-01-nen-tang/05-MEL-CEPSTRUM-CHUAN-HOA-TOKEN.md).

## 1. Layer sum: cùng shape chưa đủ cùng ý nghĩa

H_l ∈ R^(b×N×D) là readout layer l, L layers. Weighted sum:

$$\alpha_l=\frac{e^{a_l}}{\sum_{k=1}^L e^{a_k}},\qquad H=\sum_{l=1}^L\alpha_l H_l.$$

Softmax trên **layer axis**, dùng chung scalars cho samples/tokens/channels trong kiểu này. Init a_l=0 cho α_l=1/L. Concatenate đổi width thành LD, learned sum giữ D. Sum yêu cầu token correspondence; khác grid/mask positions phải align trước. Các layers cùng mạng có correspondence vị trí nhưng khác scale và coordinate semantics; một sum là trainable readout, không định lý representations đã aligned.

Toy tự biên soạn: h₂=10h₁, α₁=0,9, α₂=0,1. Tổng h=1,9h₁; dù weight layer 2 nhỏ, term của nó là h₁, lớn hơn term layer 1 là 0,9h₁. Nếu h₂=h₁, mọi α cho cùng h. Weight không xác định causal contribution. LayerNorm làm đổi scale và có thể đổi information accessible; block 12 và final-LN không là hai nguồn độc lập. [ELMo §3.2 về task-specific layer weighting](https://aclanthology.org/N18-1202.pdf), [paper SOICT §3.2](<C:/Users/LENOVO/Downloads/SOICT_2026_paper_4308.pdf>).

## 2. Mean/max/statistics: cùng dữ liệu, summary khác nhau

Một sample có valid tokens hᵢ ∈ R^D, mask mᵢ ∈ {0,1}; M=Σmᵢ > 0. Mean μ=Σmᵢhᵢ/M; max theo từng coordinate trong valid positions. Population variance v=Σmᵢ(hᵢ−μ)²/M, std σ=√v, statistics vector [μ;σ] width 2D. Bình phương/sqrt ở đây theo coordinate; nếu h có đơn vị u thì v có u² và σ có u. Embeddings thường không có đơn vị vật lý.

Toy D=1, valid [0,0,6], padding [0]. Có mask: μ=2, v=(4+4+16)/3=8, σ≈2,828427; max=6. Bỏ mask: μ=1,5, v=6,75, σ≈2,598076. Max vẫn 6 nhưng nếu valid [−2,−1] và pad 0 thì unmasked max thành 0. Padding bằng zero không vô hại cho mọi pooling.

Rare cue 6 giữa nhiều zeros bị mean pha loãng; max giữ 6 nhưng một noise spike 60 thắng cue. Statistics cho thấy dispersion nhưng không biết đó là artifact hay noise. Lặp **toàn bộ** [0,0,6] hai lần giữ mean/population std; chỉ lặp silence làm dilution đổi. Tiling không thêm independent evidence.

## 3. Attentive statistics: weights và denominator

Scalar attention ví dụ:

$$e_i=v^\top\tanh(Wh_i+c)+c_0,\qquad
\beta_i=\frac{m_i e^{e_i}}{\sum_j m_j e^{e_j}},$$

$$\mu=\sum_i\beta_i h_i,\qquad
v=\sum_i\beta_i(h_i-\mu)^2
=\sum_i\beta_i h_i^2-\mu^2,\qquad \sigma=\sqrt{v}.$$

W shape K×D, v shape K; β normalized trên valid token axis, cùng scalar weight cho D channels. Channel-dependent variant có e_ic và β_ic normalized theo i **cho mỗi c**; không normalize lẫn channel axis. Multi-head tạo vài summaries; không tự đảm bảo heads dùng cues khác nhau. [Okabe, Koshinaka, Shinoda, §3 equations 3–6](https://www.isca-archive.org/interspeech_2018/okabe18_interspeech.pdf); nguồn gốc là speaker verification, không universal deepfake gain. [ECAPA-TDNN §3.1, channel-dependent attention](https://arxiv.org/pdf/2005.07143).

Toy valid h=[0,2,4], weights β=[0,1;0,2;0,7]. μ=0+0,4+2,8=3,2. Second moment=0+0,8+11,2=12; v=12−10,24=1,76; σ≈1,326650. Centered computation cho 0,1×10,24+0,2×1,44+0,7×0,64=1,76. Đây là **weighted population std**, không unbiased sample estimator. Unbiased correction phụ thuộc statistical sampling/weights assumptions; không tự chia T−1 cho learned-attention representation.

Finite precision có thể làm second-moment subtraction hơi âm. Hai cách smoothing khác nhau:

- √max(v, ε): floor σ≥√ε, derivative theo v bằng 0 bên dưới floor.
- √(max(v,0)+ε): mọi v≥0 tăng thêm ε; derivative khác.

Với v=0, ε=10⁻⁸ cả hai trả 10⁻⁴; với v=ε, cách thứ hai trả √(2ε). Chọn ε có đơn vị variance trong toy. Code notebook mẫu v38 dùng centered weighted variance rồi clamp(min=1e−8). Đây là source fact, không xác minh numerical stability của run; [root review](research/root-source-review.md).

All-padding: denominator 0/softmax toàn −∞ không được xử lý như sample thật. Toy helper của lớp **raise error** trước normalize; production cần policy rõ (reject/empty output/dummy với valid flag). Softmax ổn định subtract max **trên valid scores**, không để padding thắng max. Mask tại pooling chưa tự xóa ảnh hưởng padding đã đi qua encoder attention, positional embeddings hoặc normalization.

![Toy pooling và thứ tự](assets/pooling-counterexamples.png)

Hình tự biên soạn từ các số ở bài này, không model outputs. Nó cho thấy summary đổi khi padding bị tính vào và mean/std không phân biệt thứ tự.

## 4. Pooling mất order, nhưng encoder có thể đã mã hóa order

Sequences [0,0,6] và [6,0,0] có cùng mean/std/max. Pooling trên multiset raw features không phân biệt onset vs offset. Nếu token hᵢ đã mang positions/context từ encoder, đổi waveform order có thể đổi chính hᵢ; vì vậy không nói cả ViT detector bất biến mọi temporal permutation. Định lý permutation invariance của summary áp dụng cho **permute cùng tập token vectors**, không tái tính encoder trên audio đã permute.

Temporal TDNN/RNN/Transformer đọc transitions với thứ tự; graph đọc quan hệ theo node/edge construction; MIL đọc bag label có positive instance. MIL max cho giả định “clip positive nếu có ít nhất một fake segment”, nhưng utterance CE không tự train segment localization; nuisance outlier có thể thắng. [Attention-based MIL §2](https://proceedings.mlr.press/v80/ilse18a/ilse18a.pdf).

## 5. Mapping 2D thành temporal sequence

Paper có H_grid shape b×32×4×768. Có ít nhất hai mappings khác nhau:

1. Mean over frequency → b×32×768: rẻ, bỏ khác biệt giữa bốn vùng mel.
2. Concatenate four frequency vectors → b×32×3.072, rồi projection về width D': giữ slot identity nhưng thêm parameters và normalization assumptions.

Toy grid time rows [[1,10],[2,20]]. Flatten time-major [1,10,2,20] **không là bốn successive times**: token 1→token 2 cùng time, token 2→token 3 vừa đổi frequency vừa đổi time. Mean-frequency sequence [5,5;11] có hai steps; concat [[1,10],[2,20]] giữ hai frequency slots. Frequency-major flatten còn tạo adjacency khác. Phải kiểm actual layout code, không suy từ shape duy nhất.

ViT full attention đã trộn time/frequency context. Gắn graph rồi gọi nodes “pure temporal/spectral cues” cần construction và evidence; không tương đương AASIST chỉ vì dùng graph library.

## 6. Paper và bài tập

Paper dùng scalar layer weighting L=13 và scalar token attention K=128, concat mean/std 1.536 dims; không frame classifier. Học weights tốt cho label task có thể tập trung silence/channel, không bắt buộc artifact. Inference remove token/layer so với retrain readout trả lời hai câu hỏi khác nhau, học ở bài 9.

1. Valid [2,4], pad 0: masked mean/std và unmasked mean là gì?
2. β=[0,5;0,5], h=[1,3]. Population variance và unbiased sample variance bằng nhau không?
3. All-padding sample được softmax weight uniform để “cho chạy” có vấn đề gì?
4. Một layer có weight 0,8. Kết luận “80% evidence đến từ layer đó” được không?
5. Reshape b×128×768 sang b×128×768 rồi đưa RNN có tự giữ đúng time sequence không?

<details>
<summary>Đáp án giải thích</summary>

1. Mean 3, population std 1; unmasked mean 2. Pad làm thay observed multiset.
2. Population variance 1; sample estimator với hai independent equal-weight observations là 2. Learned summary không mặc định là sample estimator.
3. Summary của padding được coi là evidence, và mask trước encoder cũng chưa rõ. Cần explicit empty-input policy.
4. Không. Scale/correlation/LN/backend interaction đổi contribution; weighting không intervention.
5. Không. 128 index gồm 32×4 grid; temporal adjacency phải map rõ rồi mới chọn head.

</details>
