# Bài 3 — DCF, LLR, Bayes threshold và calibration

[Trước](02-RANKING-ROC-PR-EER.md) · [Tiếp](04-TANDEM-SASV-LOCALIZATION.md).

## Mục tiêu và tiền đề

Suy threshold từ costs/priors, tính raw/normalized cost và phân biệt minDCF/actDCF/Cllr. Cần [Bayes/weighted CE](../lop-01-nen-tang/06-XAC-SUAT-LOSS-DETECTION.md) và [ranking bài 2](02-RANKING-ROC-PR-EER.md). Log dùng natural ln; LLR dưới đây **B/S**, high=B.

## 1. Bayes rule: threshold từ hai hành động

Đặt $\ell(x)=\log[p(x|B)/p(x|S)]$. Accept có conditional risk $C_{fa}P(S|x)$; reject có risk $C_{miss}P(B|x)$, coi correct decisions cost 0. Accept nếu:

$$\frac{P(B|x)}{P(S|x)}\ge\frac{C_{fa}}{C_{miss}}
\quad\Longleftrightarrow\quad
\ell(x)\ge\log\frac{C_{fa}\pi_S}{C_{miss}\pi_B}.$$

Dùng Bayes posterior odds=$e^\ell\pi_B/\pi_S$. Equal conditional risks dùng accept như convention≥. Priors positive, costs positive; density/LLR phải đúng với population được quyết định. Đây là rule tối ưu **theo mô hình xác suất/cost đó**, không guarantee optimal cho arbitrary neural score ngoài miền.

Raw DCF=$C_{miss}\pi_BP_{miss}+C_{fa}\pi_SP_{fa}$. Normalizer thường dùng $DCF_{def}=\min(C_{miss}\pi_B,C_{fa}\pi_S)$: cost tốt hơn giữa always-reject và always-accept. Normalized DCF=raw/default; có thể >1. minDCF là minimum sweep trên eval score/labels; actDCF cần rule/threshold source đã chốt. Một số research gọi actual cost tại transferred dev threshold; official ASVspoof 5 actDCF dùng **LLR Bayes threshold**, phải phân biệt.

[ASVspoof 5 v0.6 Appendix 11.1](https://www.asvspoof.org/file/ASVspoof5___Evaluation_Plan_Phase2.pdf): $C_{miss}=1,C_{fa}=10,\pi_S=0,05$, nên weights 0,95 và 0,5; default 0,5; normalized DCF=$1,9P_{miss}+P_{fa}$, $\tau_{Bayes}=\log(0,5/0,95)=-\log1,9≈-0,641854$. Không áp những constants này mặc định cho mọi benchmark.

## 2. Worked example: cùng ranking khác actual risk

**Toy scores**, chỉ giả định diễn giải LLR để kiểm rule; không đo calibration thật. A: B=[2;1], S=[−1;−2]. Hệ B dùng $s'=s+10$: B=[12;11], S=[9;8]. Cả hai AUC 1, EER 0 và minDCF 0 vì vẫn tách hoàn toàn.

Tại official numerical threshold−0,641854: A acceptB/rejectS, raw/act normalized 0. B accept tất cả: miss 0,FA 1, raw 0,5 và normalized 1. Nếu chuyển decision threshold thành 9,358146 thì decisions của B khớp A. Nhưng đó là map rule, chưa khẳng định shifted score đã là LLR đúng để dùng **cùng** Bayes numerical threshold.

Với toy bài 2 B=[0,9;0,6;0,4],S=[0,8;0,4;0,1], normalized DCF tại 0,6 là $1,9/3+1/3=29/30≈0,966667$. Brute-force sweep cho minimum 2/3 tại threshold 0,4 (miss 0,FA 2/3); accept-all cho 1. Ngưỡng 0,6 đạt EER nhưng không tối ưu cost khi miss có weight 1,9. Bayes numerical threshold âm trên những probability-like scores lại accept-all. Không gắn LLR meaning cho score[0,1] chỉ vì high=B.

## 3. Two logits, posterior odds và LLR khác nhau

$$q_B=\frac{e^{z_B}}{e^{z_B}+e^{z_S}},\qquad\log\frac{q_B}{1-q_B}=z_B-z_S.$$

Đây là identity của head. Trong population optimum của weighted CE với effective training prior $\rho_B$, weights $w_B,w_S$, nếu class-conditional training likelihoods đúng và optimization/model đủ lý tưởng:

$$d^*(x)=\ell_{train}(x)+\log\frac{\rho_B}{\rho_S}+\log\frac{w_B}{w_S}.$$

Muốn dùng như LLR cần tháo prior/weight offset trong giả định ấy. Sampling, finite data, regularization, neural misspecification và conditional domain shift làm correction đơn giản không đủ. $d=z_B-z_S$ của paper weighted CE chưa tự calibrated LLR. Xem [loss reduction/caveats lớp 4](../lop-04-ky-thuat-detector/05-LOSS-SAMPLING-OPTIMIZATION.md), không cộng offset lần nữa nếu sampling đã đổi prior mà chưa kiểm recipe.

Fit affine $\ell_{cal}=as+b$ trên development labels theo calibration objective/prior rồi freeze. $a>0$ giữ ranking; temperature-only có $b=0$, có thể không sửa prior offset. Negative slope đổi rank/polarity, không được chọn bằng test labels. Dev/test conditional shift vẫn có thể phá calibration. [Guo et al.2017 §4](https://proceedings.mlr.press/v70/guo17a/guo17a.pdf) định nghĩa temperature scaling cho posterior calibration, không tự là proof LLR deployment.

## 4. Cllr và empirical oracle diagnostic

Với LLR high=B:

$$C_{llr}=\frac{\overline{softplus(-\ell)}_B+\overline{softplus(\ell)}_S}{2\ln2},\quad
softplus(u)=\log(1+e^u).$$

Tính ổn định bằng $\max(u,0)+\log(1+e^{-|u|})$ hoặc logaddexp(0,u); không naive exp 1000. All-zero LLR cho Cllr 1. Toy A ở trên cho≈0,3175; shifted B cho≈6,1316 dù ranking không đổi. Một S bị rất confidently positive bị phạt lớn. Cllr kết hợp discrimination và meaning/scale của LLR, không calibration-only score.

Đào sâu có điều kiện: PAV tìm monotone nondecreasing calibration trên **finite labelled sample**, gộp adjacent blocks vi phạm monotonicity; empirical minimum gọi $C_{llr}^{min}$, difference observed−minimum là empirical calibration loss theo convention này. PAV có thể tạo ties; không gọi nó strictly increasing transform. Optimal blocks pure-class có thể cho posterior 0/1 và LLR±∞. [Brümmer–du Preez 2013, algorithm/theorem](https://arxiv.org/pdf/1304.2331) chứng minh optimum cho binary regular proper scoring rules với constraint/order/weights được định nghĩa; không theorem về new-domain calibration.

Toy đã separated cho oracle minimum 0 với±∞ class scores; finite original cost≈0,318. Fit eval labels để có minimum là **diagnostic oracle**, không deployed calibrator. Data in-sample minimum không ước lượng free-of-bias attainable risk ngoài sample. Không fit PAV trên test rồi báo deployment calibration đã cải thiện.

![Toy: cùng ranking khác Cllr và bootstrap nhóm](assets/cost-bootstrap-toy.png)

Panel trái minh họa score scale; panel phải được giải ở bài 7. Đây là số mô phỏng.

## 5. Paper, counterexample và tự kiểm

Paper báo EER, không Cllr/actDCF. Không suy calibration quality từ seed EER. [Chương 06](../06-GIAI-PHAU-PAPER.md). AUC cao vẫn có thể poor low-FPR rank và domain calibration; sửa scale chỉ giải một phần.

1. Với equal costs/priors, LLR threshold bao nhiêu? Nếu score là posterior B thì threshold bao nhiêu?
2. Shift+10 ở toy có đổi minDCF không? Fixed threshold thì sao?
3. Cllr 1 đủ chứng minh mọi score calibrated chưa?
4. Fit PAV eval hợp lệ cho oracle diagnostic nhưng vì sao không deployment claim?

<details>
<summary>Đáp án reasoning</summary>

1. LLR 0; posterior 0,5. Hai scales khác nhau nhưng Bayes rule cùng decisions nếu đúng mappings.
2. min không đổi khi decision sets giữ. Fixed numerical threshold có thể accept-all; cần map threshold hoặc calibrate dev.
3. Chưa. All-zero đúng LLR cho hoàn toàn không phân biệt dưới equal class evidence; một hệ miscalibrated/discriminative cũng có thể có same scalar. Cần scope và discrimination diagnostic.
4. Đã dùng labels của chính population được đo để fit; không kiểm transfer tới unseen trials. Oracle diagnostic cần ghi in-sample status.

</details>

**Đào sâu:** tự derive optimum posterior từng PAV block bằng derivative class-balanced log loss; kiểm weights và prior correction trước khi map posterior block thành LLR. Không cần implement PAV để tiếp tục lớp.
