Bài 8 — Inference/scoring: coverage, ranking và ý nghĩa quyết định#
Điểm vào · Trước: domain robustness · Tiếp: diagnostics và recipe.
Mục tiêu và tiền đề#
Trace đúng score direction, hiểu fusion đổi ranking, window aggregation đổi label assumptions, và phân biệt chunked với streaming causal. Cần logits/Bayes, encoder vs predictor error, task output. Calibration/operating costs chỉ nối khái niệm; evaluation chi tiết học lớp 5.
1. Score không tự có semantics từ tên encoder#
Hai logits z_B,z_S cho:
Đây là algebra identity của cùng head, không bằng chứng q_B calibrated với deployment. Weighted CE/sampling/shift còn đổi meaning; bài 5 đã xét.
Toy A:(z_B,z_S)=(7,6),B:(4,0). Raw z_B xếp A>B; d xếp 1<4; q_B≈0,731059<0,982014. Cộng c(x) vào cả logits giữ d/q, có thể đổi raw-logit ranking. Một-logit BCE model nhận d riêng là hợp lệ; “single logit” lỗi trong toy là chọn một component của two-logit head, không mọi single-logit classifier.
PDF paper ghi B−S, high=B. Official AASIST pinned scorer lấy raw class1 logit; nguồn không được diễn giải thành same posterior ordering nếu không kiểm. Paper §3.2/§4.1, AASIST scorer. Đây là source observation, không conclusion về performance difference.
JEPA predictor error, attention weight, embedding norm và head score là các đại lượng khác. Attention weight normalized để pooling; error đo predictability theo target/mask; không scalar nào tự là spoof probability. Bài 9 lớp 3 đã có counterexamples nên không lặp proofs.
2. Monotonic transform, calibration và fusion#
Strictly increasing global transform f(d) giữ ranking và attainable decision sets nếu threshold được map tương ứng; không tự cải thiện discrimination. Positive affine ad+b,a>0 là một ví dụ; sigmoid cũng vậy trong real arithmetic. Negative slope đảo polarity; clipping/rounding/finite-precision saturation có thể tạo ties. Transform sample-dependent không có guarantee toàn ordering.
Calibration fit trên dev có thể sửa ý nghĩa probability/threshold trong distribution được kiểm. Temperature scaling dùng σ(d/T), T>0; thêm intercept b thành σ(d/T+b) là mở rộng affine/Platt-style, không công thức temperature-only. Calibration không tự chữa generator shift hoặc tăng rank separation. Guo et al., §4 temperature scaling. Dev phải đại diện và không lấy test labels để tune; test-set normalization/access có thể đổi protocol.
Score fusion: s_f=w₁d₁+w₂d₂+b. Không phải transform đơn điệu của riêng d₁. Toy:
| Mẫu | d₁ | d₂ | Fusion 0,5d₁+0,5d₂ |
|---|---|---|---|
| A | 4 | −8 | −2 |
| B | 1 | 2 | 1,5 |
Hệ 1 xếp A cao hơn B; fusion đảo thứ tự. Đảo ấy tốt hay xấu tùy labels; fusion chỉ có thể giúp khi signals/errors/cost phù hợp. Same seeds có errors tương quan nên không bảo đảm ensemble hơn best seed. Scales/polarities phải align; fit weights/calibration/aggregation trên dev rồi fixed evaluation. RawNet2 §4.4, recipe-specific SVM fusion.
Embedding fusion concat [u₁;u₂] trước learned head đổi input space/capacity, alignment và training cost; learned interaction khác score fusion. Same-capacity duplicate-view control giúp phân biệt extra capacity với complementary cues.
3. Mean logits khác mean probabilities#
Window logits d=[10,−2]: mean d=4 → σ(4)≈0,982014. Mean probabilities=(σ(10)+σ(−2))/2≈0,559579. Vì sigmoid nonlinear, hai aggregation rules không commute. Một rất confident positive window có thể chi phối mean logits; mean q giới hạn contribution trong [0,1]. Cả hai chưa tự là clip posterior: label semantics, coverage và correlation cần model rõ.
4. Fixed/sliding crops và “any fake” label#
Đổi sang spoof-oriented r_j=−d_j, r cao nghiêng S. Windows r=[0,1;0,1;0,9;0,1] có mean 0,3, max 0,9, top-2 mean 0,5. Max giữ rare positive window, mean pha loãng, top-k tradeoff tùy k và number windows. Với d high=B, any-fake max-r tương đương min-d, không max-d. Không đổi polarity rồi vẫn nói max nhạy fake.
Toy clip 4 s, window 2,56 s, hop 0,64 s: naive full-window starts 0; 0,64; 1,28 s, window cuối kết thúc 3,84 s. Tail [3,84;4] s chưa được nhìn nếu không thêm shifted final window start 1,44 s. Padding/last-window policy là phần coverage, không implementation detail vô hại. Với fake interval [3,9;3,95] s, prefix không thấy, naive grid cũng không thấy; thêm tail window mới có observation. Vẫn cần model học local cue và duration sensitivity.
Max có false alarms khi nhiều chances: nếu mỗi bona fide window có false-positive probability p độc lập, clip-max false positive=1−(1−p)^K. p=0,01,K=100 cho≈0,633968. Overlap windows correlated nên independence formula chỉ toy, không prediction cho actual detector. Mean dưới equal variance/equicorrelation ρ có Var(mean)=σ²[1+(K−1)ρ]/K;ρ=1 không giảm variance dù K lớn. Lặp clip không tạo K independent pieces of evidence.
MIL/partial-fake literature đặt instance/segment assumptions rõ; utterance-only head và pooling attention không tự trở thành frame classifier. MIL §2, PartialSpoof v2 §§2–3.
5. Streaming: score của quá khứ có nhìn tương lai không?#
Streaming causal tại thời điểm t chỉ dùng samples ≤ t, cộng permitted look-ahead nếu công bố. Chunked batch inference trên complete windows có thể nhìn toàn window; không đồng nghĩa zero-lookahead incremental output.
Các chỗ có future access: centered STFT, symmetric convolution/padding, bidirectional/full ViT attention, utterance normalization/pooling, last-window aggregation. Chỉ thay RNN thành causal chưa làm preprocessing causal. Feature masking, state reset, overlap recompute và accumulation cũng cần định nghĩa. PyTorch STFT center, ViT method full self-attention.
Ví dụ window 2,56 s xuất score sau khi cửa sổ hoàn tất: initial acquisition wait ít nhất 2,56 s rồi compute/buffering. Hop 0,64 s cho cadence tiềm năng 0,64 s, không latency 0,64 s cho first evidence. Causal sliding window dùng history[t−2,56;t] có thể là deployment design, nhưng output/threshold/error under streaming cần protocol riêng. Full-clip model không tự bảo đảm stateful efficient streaming.
6. Liên hệ paper và bài tập#
Paper reports fixed toolkit input/prefix crop, score B−S. Không có streaming/partial localization hoặc tuned sliding-window evaluation đã xác minh trong lớp 4. Fusion historical notes là dữ liệu lịch sử, không re-run. Thêm windows sẽ đổi inference system/compute/evidence coverage; phải ghi khi đọc benchmark comparison.
- Với d high=B, clip fake nếu có window fake: dùng max-d có cùng semantics max-r không?
- Calibrating một positive affine transform có thể đổi ideal ranking không? Fusion có thể không?
- Mean q của [10,−2] giống q của mean d không?
- Hop 0,64 s và window 2,56 s đủ chứng minh latency 0,64 s không?
- Prefix crop miss fake cuối clip: đó có chứng minh JEPA representation không giữ artifact không?
Đáp án giải thích
- Không:max-d chọn most-B; max-r chọn most-S, tương đương min-d.
- Affine slope positive giữ rank trong lý tưởng. Fusion dùng nhiều coordinates có thể đổi rank như toy; tốt/xấu cần labels/protocol.
- Không:≈0,559579 vs≈0,982014 vì nonlinear sigmoid.
- Không:initial observation wait, lookahead, processing và cadence là các clocks khác.
- Không:input observation chưa chứa cue; phải phân biệt coverage với information loss trong encoder.