Bài 4 — Layer fusion và pooling: gộp thông tin nào, mất thông tin nào?#
Điểm vào · Trước: adaptation · Tiếp: loss/sampling.
Mục tiêu và tiền đề#
Tính readout có padding/outlier, phân biệt token/time/channel axes, và giải thích vì sao attention weights chưa là causal importance. Cần mean/variance và geometry, attention, token geometry.
1. Layer sum: cùng shape chưa đủ cùng ý nghĩa#
H_l ∈ R^(b×N×D) là readout layer l, L layers. Weighted sum:
Softmax trên layer axis, dùng chung scalars cho samples/tokens/channels trong kiểu này. Init a_l=0 cho α_l=1/L. Concatenate đổi width thành LD, learned sum giữ D. Sum yêu cầu token correspondence; khác grid/mask positions phải align trước. Các layers cùng mạng có correspondence vị trí nhưng khác scale và coordinate semantics; một sum là trainable readout, không định lý representations đã aligned.
Toy tự biên soạn: h₂=10h₁, α₁=0,9, α₂=0,1. Tổng h=1,9h₁; dù weight layer 2 nhỏ, term của nó là h₁, lớn hơn term layer 1 là 0,9h₁. Nếu h₂=h₁, mọi α cho cùng h. Weight không xác định causal contribution. LayerNorm làm đổi scale và có thể đổi information accessible; block 12 và final-LN không là hai nguồn độc lập. ELMo §3.2 về task-specific layer weighting, paper SOICT §3.2.
2. Mean/max/statistics: cùng dữ liệu, summary khác nhau#
Một sample có valid tokens hᵢ ∈ R^D, mask mᵢ ∈ {0,1}; M=Σmᵢ > 0. Mean μ=Σmᵢhᵢ/M; max theo từng coordinate trong valid positions. Population variance v=Σmᵢ(hᵢ−μ)²/M, std σ=√v, statistics vector [μ;σ] width 2D. Bình phương/sqrt ở đây theo coordinate; nếu h có đơn vị u thì v có u² và σ có u. Embeddings thường không có đơn vị vật lý.
Toy D=1, valid [0,0,6], padding [0]. Có mask: μ=2, v=(4+4+16)/3=8, σ≈2,828427; max=6. Bỏ mask: μ=1,5, v=6,75, σ≈2,598076. Max vẫn 6 nhưng nếu valid [−2,−1] và pad 0 thì unmasked max thành 0. Padding bằng zero không vô hại cho mọi pooling.
Rare cue 6 giữa nhiều zeros bị mean pha loãng; max giữ 6 nhưng một noise spike 60 thắng cue. Statistics cho thấy dispersion nhưng không biết đó là artifact hay noise. Lặp toàn bộ [0,0,6] hai lần giữ mean/population std; chỉ lặp silence làm dilution đổi. Tiling không thêm independent evidence.
3. Attentive statistics: weights và denominator#
Scalar attention ví dụ:
W shape K×D, v shape K; β normalized trên valid token axis, cùng scalar weight cho D channels. Channel-dependent variant có e_ic và β_ic normalized theo i cho mỗi c; không normalize lẫn channel axis. Multi-head tạo vài summaries; không tự đảm bảo heads dùng cues khác nhau. Okabe, Koshinaka, Shinoda, §3 equations 3–6; nguồn gốc là speaker verification, không universal deepfake gain. ECAPA-TDNN §3.1, channel-dependent attention.
Toy valid h=[0,2,4], weights β=[0,1;0,2;0,7]. μ=0+0,4+2,8=3,2. Second moment=0+0,8+11,2=12; v=12−10,24=1,76; σ≈1,326650. Centered computation cho 0,1×10,24+0,2×1,44+0,7×0,64=1,76. Đây là weighted population std, không unbiased sample estimator. Unbiased correction phụ thuộc statistical sampling/weights assumptions; không tự chia T−1 cho learned-attention representation.
Finite precision có thể làm second-moment subtraction hơi âm. Hai cách smoothing khác nhau:
- √max(v, ε): floor σ≥√ε, derivative theo v bằng 0 bên dưới floor.
- √(max(v,0)+ε): mọi v≥0 tăng thêm ε; derivative khác.
Với v=0, ε=10⁻⁸ cả hai trả 10⁻⁴; với v=ε, cách thứ hai trả √(2ε). Chọn ε có đơn vị variance trong toy. Code notebook mẫu v38 dùng centered weighted variance rồi clamp(min=1e−8). Đây là source fact, không xác minh numerical stability của run; root review.
All-padding: denominator 0/softmax toàn −∞ không được xử lý như sample thật. Toy helper của lớp raise error trước normalize; production cần policy rõ (reject/empty output/dummy với valid flag). Softmax ổn định subtract max trên valid scores, không để padding thắng max. Mask tại pooling chưa tự xóa ảnh hưởng padding đã đi qua encoder attention, positional embeddings hoặc normalization.

Hình tự biên soạn từ các số ở bài này, không model outputs. Nó cho thấy summary đổi khi padding bị tính vào và mean/std không phân biệt thứ tự.
4. Pooling mất order, nhưng encoder có thể đã mã hóa order#
Sequences [0,0,6] và [6,0,0] có cùng mean/std/max. Pooling trên multiset raw features không phân biệt onset vs offset. Nếu token hᵢ đã mang positions/context từ encoder, đổi waveform order có thể đổi chính hᵢ; vì vậy không nói cả ViT detector bất biến mọi temporal permutation. Định lý permutation invariance của summary áp dụng cho permute cùng tập token vectors, không tái tính encoder trên audio đã permute.
Temporal TDNN/RNN/Transformer đọc transitions với thứ tự; graph đọc quan hệ theo node/edge construction; MIL đọc bag label có positive instance. MIL max cho giả định “clip positive nếu có ít nhất một fake segment”, nhưng utterance CE không tự train segment localization; nuisance outlier có thể thắng. Attention-based MIL §2.
5. Mapping 2D thành temporal sequence#
Paper có H_grid shape b×32×4×768. Có ít nhất hai mappings khác nhau:
- Mean over frequency → b×32×768: rẻ, bỏ khác biệt giữa bốn vùng mel.
- Concatenate four frequency vectors → b×32×3.072, rồi projection về width D': giữ slot identity nhưng thêm parameters và normalization assumptions.
Toy grid time rows [[1,10],[2,20]]. Flatten time-major [1,10,2,20] không là bốn successive times: token 1→token 2 cùng time, token 2→token 3 vừa đổi frequency vừa đổi time. Mean-frequency sequence [5,5;11] có hai steps; concat [[1,10],[2,20]] giữ hai frequency slots. Frequency-major flatten còn tạo adjacency khác. Phải kiểm actual layout code, không suy từ shape duy nhất.
ViT full attention đã trộn time/frequency context. Gắn graph rồi gọi nodes “pure temporal/spectral cues” cần construction và evidence; không tương đương AASIST chỉ vì dùng graph library.
6. Paper và bài tập#
Paper dùng scalar layer weighting L=13 và scalar token attention K=128, concat mean/std 1.536 dims; không frame classifier. Học weights tốt cho label task có thể tập trung silence/channel, không bắt buộc artifact. Inference remove token/layer so với retrain readout trả lời hai câu hỏi khác nhau, học ở bài 9.
- Valid [2,4], pad 0: masked mean/std và unmasked mean là gì?
- β=[0,5;0,5], h=[1,3]. Population variance và unbiased sample variance bằng nhau không?
- All-padding sample được softmax weight uniform để “cho chạy” có vấn đề gì?
- Một layer có weight 0,8. Kết luận “80% evidence đến từ layer đó” được không?
- Reshape b×128×768 sang b×128×768 rồi đưa RNN có tự giữ đúng time sequence không?
Đáp án giải thích
- Mean 3, population std 1; unmasked mean 2. Pad làm thay observed multiset.
- Population variance 1; sample estimator với hai independent equal-weight observations là 2. Learned summary không mặc định là sample estimator.
- Summary của padding được coi là evidence, và mask trước encoder cũng chưa rõ. Cần explicit empty-input policy.
- Không. Scale/correlation/LN/backend interaction đổi contribution; weighting không intervention.
- Không. 128 index gồm 32×4 grid; temporal adjacency phải map rõ rồi mới chọn head.