ajaudio / studyAUDIO-JEPA · RESEARCH NOTES
7 phút đọc · Toàn văn
Mục lục bài · 7 mục

Bài 2 — Front-end và baseline: trace cơ chế thay vì nhớ tên#

Điểm vào · Trước: hệ thống · Tiếp: adaptation.

Mục tiêu và tiền đề#

Giải thích vài baseline đủ sâu để dùng chúng chất vấn một detector lớn; phân biệt waveform input với toàn front-end học được. Cần STFT/power/phase, cepstrum, pretraining taxonomy. Baseline là một recipe có phiên bản, không tên architecture đứng riêng.

1. LFCC/CQCC-GMM: hai density models trên frame features#

LFCC: waveform → short-time spectrum → linear-spaced filter energies → log → DCT → cepstral coefficients, có thể thêm delta/delta-delta. CQCC bắt đầu từ constant-Q representation với resolution phụ thuộc frequency, rồi log/resampling/DCT theo recipe. Chúng đổi cách chia information budget trên phổ; không mặc định giữ phase chi tiết hoặc có cùng feature geometry. Phần DSP đã học ở lớp 1.

Với v_t ∈ R^D, GMM cho lớp c ∈ {B,S}:

pc(v)=∑k=1Kπck N(v;μck,Σck),πck≥0,∑kπck=1.p_c(v)=\sum_{k=1}^{K}\pi_{ck}\,\mathcal N(v;\mu_{ck},\Sigma_{ck}),\qquad \pi_{ck}\ge0,\quad\sum_k\pi_{ck}=1.

Gaussian centers/covariances và mixture weights được fit theo từng lớp. Trách nhiệm component γ_tk proportional π_ck N(v_t;μ_ck,Σ_ck), normalized over k; EM luân phiên responsibilities và parameter updates. Mixture component không tự là generator identity. Diagonal/full covariance, initialization/regularization và feature scaling đều đổi density. Compute log density bằng log-sum-exp, không cộng log của từng component như thể cùng sinh một frame.

Frame-independent approximation cho utterance score trung bình:

s(x)=1T∑t=1T[ln⁡pB(vt)−ln⁡pS(vt)].s(x)=\frac1T\sum_{t=1}^{T}\bigl[\ln p_B(v_t)-\ln p_S(v_t)\bigr].

Higher s nghiêng B. Density không phải class posterior; chỉ under assumptions và cộng prior/cost mới có decision interpretation. Mean thay sum normalize duration, vẫn không làm frames độc lập hoặc xóa source cues.

Toy đầy đủ, tự biên soạn: K=1,D=1, p_B=N(0,1),p_S=N(2,1). Gaussian log density difference = −v²/2+(v−2)²/2=2−2v. Frames [0,1,2] cho [2,0,−2], mean score 0. Permute [2,1,0] vẫn 0; [1,1,1] cũng 0 dù trajectory/variance khác. GMM likelihood tổng quát có thể phụ thuộc distribution/variance, nhưng frame averaging vẫn không encode order nếu chỉ permute cùng features. Delta features có local order trước aggregation, nên không suy toàn hệ bất biến với permuting raw waveform.

2. Official configuration là evidence khác chạy baseline#

ASVspoof 2021 LA source ở commit 9b33f5… dùng train protocol LA 2019, eval LA 2021, mặc định load pretrained GMM. LFCC script cấu hình window 30 ms, FFT 1024, 70 filters, no_coeff=19 gồm c0 theo comment source, range 0–4 kHz; nhánh tự train K=512, tối đa 10 iterations. CQCC script có 12 bins/octave, fmin=62,5 Hz, fmax=4 kHz, cf=19 không c0, option ZsdD; cũng K=512. Không nhập những defaults ấy thành config mọi LFCC/CQCC paper. LFCC source, CQCC source.

Caveat source đã thấy ở cả hai scripts: vòng spoof references bonafideIdx(i) khi dựng path. Đây là indexing discrepancy ở nhánh tự train, không kết luận pretrained weights cũng bị lỗi; không chạy nhánh ấy trong lớp này. Helpers/config/checkpoint provenance cần audit trước reproduction. Học baseline còn là học phân biệt artifact có sẵn với experiment verified.

3. RawNet2 anti-spoofing: fixed filterbank + learned hierarchy#

RawNet2 adaptation cho anti-spoofing gốc dùng fixed Sinc filters trên waveform, residual blocks + filter-wise feature map scaling, GRU tổng hợp time, FC phân loại B/S. Paper §3 giữ cutoff/bandwidth filters fixed vì recipe-data sparsity; không gọi mọi weight của waveform frontend trainable. FMS điều chỉnh feature-map channels; không giống attentive statistics over tokens. Tak et al., RawNet2 v3 §3, Table 1.

SettingPaper anti-spoofing v3Official ASVspoof 2021 LA config/source
Fixed input64.000 samples64.600 samples
Sinc filters/kernel128 / 12920 / first_conv=1024 được tăng thành 1025 trong code
Later residual width512 theo Table 1128 ở config được đọc
Temporal backendGRU 1024GRU 1024, 3 layers

Đây là hai recipe khác nhau của cùng họ, không chép config 2021 thành Table 1 paper. YAML pinned, model.py pinned. Waveform network vẫn bị crop/length/sampling decisions; GRU toàn đoạn không tự chứng minh online detector đáp ứng latency.

4. AASIST: topology xuất phát từ hai axes của feature map#

AASIST dùng waveform Sinc front-end/residual encoder tạo E ∈ R^(b×C×S_f×T_f), C là channels, S_f spectral positions, T_f temporal positions. Chữ S_f khác label S. Official code lấy max |E| theo T_f để tạo spectral nodes, theo S_f để tạo temporal nodes; mỗi node vector width C. Graph attention trong từng group, pooling, rồi heterogeneous attention với type-specific relations và master node. Hai pathways/readouts được hợp trước classifier. Jung et al., §§2–3/Figure 1, pinned Model.forward.

Shape toy tự biên soạn: E có b=1,C=2,S_f=3,T_f=4. Max-time cho 3 spectral nodes × 2 features; max-frequency cho 4 temporal nodes × 2 features. Graph combined có hai types với tổng 7 nodes trước pooling, không 12 flattened nodes tương đương. Nếu chỉ có utterance vector 768 chiều, frequency/time axes đã mất; tùy ý chia 768 channels thành nodes không khôi phục cùng topology.

Node relations không bắt buộc là một spatial adjacency grid; AASIST learns pairwise attention theo graph design. Trong source pinned, Sinc filter bank được dựng cố định, còn residual/graph blocks học. Gắn một graph module lên ViT là thiết kế cần mapping/readout/capacity validation riêng; không gọi là AASIST equivalent. Chính max reduction có thể bỏ order/locality ở một branch, thêm graph không tự phục hồi mọi information bị reduce.

Root kiểm main.py thấy scorer lấy raw class-B logit batch_out[:,1]; Model trả out_layer không log-softmax. Đây là convention của source này, khác B−S trong paper mình; không sửa source hoặc suy hậu quả benchmark nếu chưa chạy. Scorer pinned.

5. Spectrogram/pretrained encoders và multi-view#

Spectrogram CNN giữ locality theo time/frequency; ViT patch projection chuyển patch values thành tokens rồi contextualizes. AST có image-pretraining adaptation trong recipe gốc; “spectrogram Transformer” không tự nói nguồn supervision. AST §2. Speech SSL và general-audio SSL khác corpora/tasks/resolution; encoder giữ speech semantics hữu ích không bảo đảm giữ forensic residual. Prerequisite lớp 3 speech SSL, AudioMAE/contextual targets.

Whisper học speech/text bằng weak supervision, không pure SSL như HuBERT. Dùng encoder Whisper cho detector là supervised/weakly-supervised pretrained features + downstream readout; cần nói layer và adaptation. Whisper §2.

Multi-view có thể dùng waveform + magnitude + phase/prosody. Complementarity là giả thuyết: hai views có thể trùng information, hoặc thêm model capacity/data giải thích gain. Toy errors trên 100 samples: model A sai 10, B sai 10; nếu cùng 10 thì voting không có diverse error coverage, nếu error sets disjoint có tiềm năng complementarity nhưng fusion rule chưa biết ai đúng. Đo view-only, same-capacity duplicate-view control, aligned inputs/splits và fusion fitted dev. Không infer causal cue chỉ từ fusion gain; RawNet2 §4.4 là evidence recipe-specific.

6. Liên hệ paper và bài tập#

Audio-JEPA paper mình dùng log-mel ViT context encoder pretrained AudioSet; downstream geometry khác native. LFCC/RawNet2/AASIST giúp đặt câu hỏi về input information, topology, learning scale và recipe fairness. Chúng không là cam kết ghép module vào đề tài.

  1. Vì sao toy Gaussian [0,1,2] và [1,1,1] cùng score? Điều gì của full GMM vẫn có thể khác toy?
  2. Raw waveform input có nghĩa mọi filters được học không?
  3. Map 3×4 thành 12 nodes rồi dense graph có equivalent AASIST toy không?
  4. Fusion hai views tốt hơn một view có chứng minh phase cue gây gain chưa?
  5. Source baseline có bug nhánh train: được kết luận checkpoint phát hành lỗi không?
Đáp án giải thích
  1. Với hai unit-variance Gaussians toy, log-ratio affine theo v nên cùng mean cho cùng score. General mixtures có nonlinear density ratio; order vẫn bị frame averaging bỏ nếu feature set giữ nguyên.
  2. Không; fixed Sinc/filter/preprocess có thể đứng trước learned layers.
  3. Không: 12 patch-like nodes khác 3 spectral + 4 temporal nodes tạo bằng axis reductions và type-specific interactions.
  4. Chưa: capacity/optimization/data/error complementarity là competing explanations; cần controls/interventions.
  5. Không. Phải nối source revision, executed branch và checkpoint provenance; ở đây mới source discrepancy.
DỪNG LẠI & TỰ KIỂM TRA

Bạn đã giải thích được cơ chế trong bài?

↓ Bản Markdown nguyên gốcGiữ nguyên nội dung · Công thức, bảng và nguồn đầy đủ.Các chat bàn giao được mở trong Codex.

Gõ từ khóa để tìm bài học và đoạn liên quan.