Bài 3 — DCF, LLR, Bayes threshold và calibration#
Mục tiêu và tiền đề#
Suy threshold từ costs/priors, tính raw/normalized cost và phân biệt minDCF/actDCF/Cllr. Cần Bayes/weighted CE và ranking bài 2. Log dùng natural ln; LLR dưới đây B/S, high=B.
1. Bayes rule: threshold từ hai hành động#
Đặt . Accept có conditional risk ; reject có risk , coi correct decisions cost 0. Accept nếu:
Dùng Bayes posterior odds=. Equal conditional risks dùng accept như convention≥. Priors positive, costs positive; density/LLR phải đúng với population được quyết định. Đây là rule tối ưu theo mô hình xác suất/cost đó, không guarantee optimal cho arbitrary neural score ngoài miền.
Raw DCF=. Normalizer thường dùng : cost tốt hơn giữa always-reject và always-accept. Normalized DCF=raw/default; có thể >1. minDCF là minimum sweep trên eval score/labels; actDCF cần rule/threshold source đã chốt. Một số research gọi actual cost tại transferred dev threshold; official ASVspoof 5 actDCF dùng LLR Bayes threshold, phải phân biệt.
ASVspoof 5 v0.6 Appendix 11.1: , nên weights 0,95 và 0,5; default 0,5; normalized DCF=, . Không áp những constants này mặc định cho mọi benchmark.
2. Worked example: cùng ranking khác actual risk#
Toy scores, chỉ giả định diễn giải LLR để kiểm rule; không đo calibration thật. A: B=[2;1], S=[−1;−2]. Hệ B dùng : B=[12;11], S=[9;8]. Cả hai AUC 1, EER 0 và minDCF 0 vì vẫn tách hoàn toàn.
Tại official numerical threshold−0,641854: A acceptB/rejectS, raw/act normalized 0. B accept tất cả: miss 0,FA 1, raw 0,5 và normalized 1. Nếu chuyển decision threshold thành 9,358146 thì decisions của B khớp A. Nhưng đó là map rule, chưa khẳng định shifted score đã là LLR đúng để dùng cùng Bayes numerical threshold.
Với toy bài 2 B=[0,9;0,6;0,4],S=[0,8;0,4;0,1], normalized DCF tại 0,6 là . Brute-force sweep cho minimum 2/3 tại threshold 0,4 (miss 0,FA 2/3); accept-all cho 1. Ngưỡng 0,6 đạt EER nhưng không tối ưu cost khi miss có weight 1,9. Bayes numerical threshold âm trên những probability-like scores lại accept-all. Không gắn LLR meaning cho score[0,1] chỉ vì high=B.
3. Two logits, posterior odds và LLR khác nhau#
Đây là identity của head. Trong population optimum của weighted CE với effective training prior , weights , nếu class-conditional training likelihoods đúng và optimization/model đủ lý tưởng:
Muốn dùng như LLR cần tháo prior/weight offset trong giả định ấy. Sampling, finite data, regularization, neural misspecification và conditional domain shift làm correction đơn giản không đủ. của paper weighted CE chưa tự calibrated LLR. Xem loss reduction/caveats lớp 4, không cộng offset lần nữa nếu sampling đã đổi prior mà chưa kiểm recipe.
Fit affine trên development labels theo calibration objective/prior rồi freeze. giữ ranking; temperature-only có , có thể không sửa prior offset. Negative slope đổi rank/polarity, không được chọn bằng test labels. Dev/test conditional shift vẫn có thể phá calibration. Guo et al.2017 §4 định nghĩa temperature scaling cho posterior calibration, không tự là proof LLR deployment.
4. Cllr và empirical oracle diagnostic#
Với LLR high=B:
Tính ổn định bằng hoặc logaddexp(0,u); không naive exp 1000. All-zero LLR cho Cllr 1. Toy A ở trên cho≈0,3175; shifted B cho≈6,1316 dù ranking không đổi. Một S bị rất confidently positive bị phạt lớn. Cllr kết hợp discrimination và meaning/scale của LLR, không calibration-only score.
Đào sâu có điều kiện: PAV tìm monotone nondecreasing calibration trên finite labelled sample, gộp adjacent blocks vi phạm monotonicity; empirical minimum gọi , difference observed−minimum là empirical calibration loss theo convention này. PAV có thể tạo ties; không gọi nó strictly increasing transform. Optimal blocks pure-class có thể cho posterior 0/1 và LLR±∞. Brümmer–du Preez 2013, algorithm/theorem chứng minh optimum cho binary regular proper scoring rules với constraint/order/weights được định nghĩa; không theorem về new-domain calibration.
Toy đã separated cho oracle minimum 0 với±∞ class scores; finite original cost≈0,318. Fit eval labels để có minimum là diagnostic oracle, không deployed calibrator. Data in-sample minimum không ước lượng free-of-bias attainable risk ngoài sample. Không fit PAV trên test rồi báo deployment calibration đã cải thiện.

Panel trái minh họa score scale; panel phải được giải ở bài 7. Đây là số mô phỏng.
5. Paper, counterexample và tự kiểm#
Paper báo EER, không Cllr/actDCF. Không suy calibration quality từ seed EER. Chương 06. AUC cao vẫn có thể poor low-FPR rank và domain calibration; sửa scale chỉ giải một phần.
- Với equal costs/priors, LLR threshold bao nhiêu? Nếu score là posterior B thì threshold bao nhiêu?
- Shift+10 ở toy có đổi minDCF không? Fixed threshold thì sao?
- Cllr 1 đủ chứng minh mọi score calibrated chưa?
- Fit PAV eval hợp lệ cho oracle diagnostic nhưng vì sao không deployment claim?
Đáp án reasoning
- LLR 0; posterior 0,5. Hai scales khác nhau nhưng Bayes rule cùng decisions nếu đúng mappings.
- min không đổi khi decision sets giữ. Fixed numerical threshold có thể accept-all; cần map threshold hoặc calibrate dev.
- Chưa. All-zero đúng LLR cho hoàn toàn không phân biệt dưới equal class evidence; một hệ miscalibrated/discriminative cũng có thể có same scalar. Cần scope và discrimination diagnostic.
- Đã dùng labels của chính population được đo để fit; không kiểm transfer tới unseen trials. Oracle diagnostic cần ghi in-sample status.
Đào sâu: tự derive optimum posterior từng PAV block bằng derivative class-balanced log loss; kiểm weights và prior correction trước khi map posterior block thành LLR. Không cần implement PAV để tiếp tục lớp.