# Bài 8 — Inference/scoring: coverage, ranking và ý nghĩa quyết định

[Điểm vào](00-BAT-DAU-LOP-04.md) · Trước: [domain robustness](07-DATA-DOMAIN-ROBUSTNESS.md) · Tiếp: [diagnostics và recipe](09-DIAGNOSTICS-RECIPE-PAPER.md).

## Mục tiêu và tiền đề

Trace đúng score direction, hiểu fusion đổi ranking, window aggregation đổi label assumptions, và phân biệt chunked với streaming causal. Cần [logits/Bayes](../lop-01-nen-tang/06-XAC-SUAT-LOSS-DETECTION.md), [encoder vs predictor error](../lop-03-ssl-jepa/09-ENCODER-ERROR-SCORING.md), [task output](../lop-02-deepfake-bai-toan/06-THREAT-MODEL-DAU-RA-SUPERVISION.md). Calibration/operating costs chỉ nối khái niệm; evaluation chi tiết học lớp 5.

## 1. Score không tự có semantics từ tên encoder

Hai logits z_B,z_S cho:

$$q_B=\frac{e^{z_B}}{e^{z_B}+e^{z_S}}=\sigma(d),\qquad d=z_B-z_S.$$

Đây là algebra identity của **cùng head**, không bằng chứng q_B calibrated với deployment. Weighted CE/sampling/shift còn đổi meaning; bài 5 đã xét.

Toy A:(z_B,z_S)=(7,6),B:(4,0). Raw z_B xếp A>B; d xếp 1<4; q_B≈0,731059<0,982014. Cộng c(x) vào cả logits giữ d/q, có thể đổi raw-logit ranking. Một-logit BCE model nhận d riêng là hợp lệ; “single logit” lỗi trong toy là chọn một component của two-logit head, không mọi single-logit classifier.

PDF paper ghi B−S, high=B. Official AASIST pinned scorer lấy raw class1 logit; nguồn không được diễn giải thành same posterior ordering nếu không kiểm. [Paper §3.2/§4.1](<C:/Users/LENOVO/Downloads/SOICT_2026_paper_4308.pdf>), [AASIST scorer](https://github.com/clovaai/aasist/blob/a04c9863f63d44471dde8a6abcb3b082b07cd1/main.py#L264). Đây là source observation, không conclusion về performance difference.

JEPA predictor error, attention weight, embedding norm và head score là các đại lượng khác. Attention weight normalized để pooling; error đo predictability theo target/mask; không scalar nào tự là spoof probability. [Bài 9 lớp 3](../lop-03-ssl-jepa/09-ENCODER-ERROR-SCORING.md) đã có counterexamples nên không lặp proofs.

## 2. Monotonic transform, calibration và fusion

Strictly increasing global transform f(d) giữ ranking và attainable decision sets nếu threshold được map tương ứng; không tự cải thiện discrimination. Positive affine ad+b,a>0 là một ví dụ; sigmoid cũng vậy trong real arithmetic. Negative slope đảo polarity; clipping/rounding/finite-precision saturation có thể tạo ties. Transform sample-dependent không có guarantee toàn ordering.

Calibration fit trên dev có thể sửa ý nghĩa probability/threshold trong distribution được kiểm. Temperature scaling dùng σ(d/T), T>0; thêm intercept b thành σ(d/T+b) là mở rộng affine/Platt-style, không công thức temperature-only. Calibration không tự chữa generator shift hoặc tăng rank separation. [Guo et al., §4 temperature scaling](https://proceedings.mlr.press/v70/guo17a/guo17a.pdf). Dev phải đại diện và không lấy test labels để tune; test-set normalization/access có thể đổi protocol.

Score fusion: s_f=w₁d₁+w₂d₂+b. Không phải transform đơn điệu của riêng d₁. Toy:

| Mẫu | d₁ | d₂ | Fusion 0,5d₁+0,5d₂ |
|---|---:|---:|---:|
| A | 4 | −8 | −2 |
| B | 1 | 2 | 1,5 |

Hệ 1 xếp A cao hơn B; fusion đảo thứ tự. Đảo ấy tốt hay xấu tùy labels; fusion chỉ có thể giúp khi signals/errors/cost phù hợp. Same seeds có errors tương quan nên không bảo đảm ensemble hơn best seed. Scales/polarities phải align; fit weights/calibration/aggregation trên dev rồi fixed evaluation. [RawNet2 §4.4, recipe-specific SVM fusion](https://arxiv.org/html/2011.01108v3#S4.SS4).

Embedding fusion concat [u₁;u₂] trước learned head đổi input space/capacity, alignment và training cost; learned interaction khác score fusion. Same-capacity duplicate-view control giúp phân biệt extra capacity với complementary cues.

## 3. Mean logits khác mean probabilities

Window logits d=[10,−2]: mean d=4 → σ(4)≈0,982014. Mean probabilities=(σ(10)+σ(−2))/2≈0,559579. Vì sigmoid nonlinear, hai aggregation rules không commute. Một rất confident positive window có thể chi phối mean logits; mean q giới hạn contribution trong [0,1]. Cả hai chưa tự là clip posterior: label semantics, coverage và correlation cần model rõ.

## 4. Fixed/sliding crops và “any fake” label

Đổi sang spoof-oriented r_j=−d_j, r cao nghiêng S. Windows r=[0,1;0,1;0,9;0,1] có mean 0,3, max 0,9, top-2 mean 0,5. Max giữ rare positive window, mean pha loãng, top-k tradeoff tùy k và number windows. Với d high=B, any-fake max-r tương đương **min-d**, không max-d. Không đổi polarity rồi vẫn nói max nhạy fake.

Toy clip 4 s, window 2,56 s, hop 0,64 s: naive full-window starts 0; 0,64; 1,28 s, window cuối kết thúc 3,84 s. Tail [3,84;4] s chưa được nhìn nếu không thêm shifted final window start 1,44 s. Padding/last-window policy là phần coverage, không implementation detail vô hại. Với fake interval [3,9;3,95] s, prefix không thấy, naive grid cũng không thấy; thêm tail window mới có observation. Vẫn cần model học local cue và duration sensitivity.

Max có false alarms khi nhiều chances: nếu mỗi bona fide window có false-positive probability p độc lập, clip-max false positive=1−(1−p)^K. p=0,01,K=100 cho≈0,633968. Overlap windows correlated nên independence formula **chỉ toy**, không prediction cho actual detector. Mean dưới equal variance/equicorrelation ρ có Var(mean)=σ²[1+(K−1)ρ]/K;ρ=1 không giảm variance dù K lớn. Lặp clip không tạo K independent pieces of evidence.

MIL/partial-fake literature đặt instance/segment assumptions rõ; utterance-only head và pooling attention không tự trở thành frame classifier. [MIL §2](https://proceedings.mlr.press/v80/ilse18a/ilse18a.pdf), [PartialSpoof v2 §§2–3](https://arxiv.org/pdf/2104.02518v2).

## 5. Streaming: score của quá khứ có nhìn tương lai không?

Streaming causal tại thời điểm t chỉ dùng samples ≤ t, cộng permitted look-ahead nếu công bố. Chunked batch inference trên complete windows có thể nhìn toàn window; không đồng nghĩa zero-lookahead incremental output.

Các chỗ có future access: centered STFT, symmetric convolution/padding, bidirectional/full ViT attention, utterance normalization/pooling, last-window aggregation. Chỉ thay RNN thành causal chưa làm preprocessing causal. Feature masking, state reset, overlap recompute và accumulation cũng cần định nghĩa. [PyTorch STFT `center`](https://docs.pytorch.org/docs/2.14/generated/torch.stft.html), [ViT method full self-attention](https://arxiv.org/pdf/2010.11929).

Ví dụ window 2,56 s xuất score sau khi cửa sổ hoàn tất: initial acquisition wait ít nhất 2,56 s rồi compute/buffering. Hop 0,64 s cho cadence tiềm năng 0,64 s, không latency 0,64 s cho first evidence. Causal sliding window dùng history[t−2,56;t] có thể là deployment design, nhưng output/threshold/error under streaming cần protocol riêng. Full-clip model không tự bảo đảm stateful efficient streaming.

## 6. Liên hệ paper và bài tập

Paper reports fixed toolkit input/prefix crop, score B−S. Không có streaming/partial localization hoặc tuned sliding-window evaluation đã xác minh trong lớp 4. Fusion historical notes là dữ liệu lịch sử, không re-run. Thêm windows sẽ đổi inference system/compute/evidence coverage; phải ghi khi đọc benchmark comparison.

1. Với d high=B, clip fake nếu có window fake: dùng max-d có cùng semantics max-r không?
2. Calibrating một positive affine transform có thể đổi ideal ranking không? Fusion có thể không?
3. Mean q của [10,−2] giống q của mean d không?
4. Hop 0,64 s và window 2,56 s đủ chứng minh latency 0,64 s không?
5. Prefix crop miss fake cuối clip: đó có chứng minh JEPA representation không giữ artifact không?

<details>
<summary>Đáp án giải thích</summary>

1. Không:max-d chọn most-B; max-r chọn most-S, tương đương min-d.
2. Affine slope positive giữ rank trong lý tưởng. Fusion dùng nhiều coordinates có thể đổi rank như toy; tốt/xấu cần labels/protocol.
3. Không:≈0,559579 vs≈0,982014 vì nonlinear sigmoid.
4. Không:initial observation wait, lookahead, processing và cadence là các clocks khác.
5. Không:input observation chưa chứa cue; phải phân biệt coverage với information loss trong encoder.

</details>
