Bài 5 — Pooled, macro, per-group và worst-group#
Mục tiêu và tiền đề#
Tính score-offset counterexample và phân biệt shared threshold với per-group oracle. Cần EER/ROC, risk, domain shift.
1. Ba đại lượng tưởng giống nhau#
Pooled EER: gom scores, một scale và sweep threshold chung. Macro EER=, thường : từng EER có oracle threshold riêng. Shared-threshold group rates: fit threshold từ dev, freeze rồi tính miss/FA từng group trên test. Cái thứ ba mô tả policy chuyển giao; hai cái đầu là ranking/sweep summaries.
Với fixed threshold:
Weights theo lớp có thể khác nhau. EER nonlinear nên không lấy hai expressions ấy tại group-specific thresholds rồi gọi pooled EER. ASVspoof 5 Appendix 11.1 dùng pooled scores cho overall ranking; aggregate choice là protocol, không verdict mọi macro sai.
2. Worked example: hai group perfect, pooled 50%#
Toy:
| Group | B scores | S scores | Local separating threshold |
|---|---|---|---|
| A | [3;4] | [1;2] | 2,5 |
| B | [103;104] | [101;102] | 102,5 |
Mỗi group AUC 1/EER 0, equal-weight macroEER 0. Pooled B=[3;4;103;104], S=[1;2;101;102]. Threshold 50: miss 2/4,FA 2/4 → EER 50%. Pooled AUC: low B thắng 2 S mỗi mẫu (4 wins), high B thắng 4 mỗi mẫu (8 wins), tổng 12/16=0,75. Threshold 2,5 dùng toàn bộ thì groupB FA 100%; threshold 102,5 thì groupA miss 100%. Không có perfect common threshold.
Trừ 100 chỉ groupB giúp pooledEER 0, giữ local ranking. Đây là group-conditioned offset correction, không global monotone transform. Fit offset bằng eval labels rồi gọi transferred calibration là oracle; use metadata/target access/calibration policy phải hợp protocol. Toy không chứng minh actual detector có offsets 100.
3. Weights đổi population và đôi khi đảo thứ hạng#
Toy fixed-rule group error summaries: A errors(g1,g2)=(0;0,4), B=(0,1;0,2). Equal macroA 0,2>B0,15 nên B tốt hơn. Population 90%g1: A0,04<B0,11 nên A tốt hơn. Đây là change of estimand/population weights, không arithmetic contradiction. Không gọi “trung bình” khi không nêu equal groups hay sample/population weights. Với DCF, còn lớp/cost weights.
Per-attack thường dùng S của attack g và cùng toàn bộ genuine reference. Khi genuine scores đổi, nhiều attack EER đổi cùng nhau; attack rows không independent. Shared reference hợp lệ nếu disclosed, nhưng không coi G rows là G independent experiments. Nếu group chỉ có S thì own EER undefined; cần specify reference B. Không quietly bỏ missing-class groups rồi vẫn gọi macro toàn taxonomy. Có thể báo S acceptance ở shared threshold và B miss toàn reference như hai quantities rõ hơn.
4. Worst-group, selection và uncertainty#
Worst-group score= với lower-is-better metric; G và definitions phải chốt. Nhóm 2 clips có 1 error cho 50%; nhóm 1.000 clips có 400 cho 40%. Không thể biết chắc population nhóm nhỏ khó hơn chỉ từ ordering sample; intervals cần đúng sampling units. Tìm subgroup hậu nghiệm trong nhiều partitions rồi báo worst/best là một selection process. Với bootstrap statistic “worst-group”, cần recompute max trên cùng predefined groups mỗi replicate; CI của group được chọn duy nhất trả lời khác và bỏ selection uncertainty. Family/domain sample ít vẫn giới hạn inference dù within-group utterance intervals hẹp.
Báo cả aggregate theo primary protocol và phân rã có denominators/reference/counts/uncertainty. Không lựa mỗi subgroup hệ mình thắng để kể generalization. Bouthillier et al.2021 §§2–3 cung cấp sources-of-variation framing; maximum/weights toys là suy luận tự biên soạn từ estimand.
5. Paper và tự kiểm#
Table 2 paper là per-corpus EER, không pooled across five corpora và không một transferred operating threshold. Không tự lấy mean năm cells khi frozenADD chưa chạy và In-the-WildFT n=1 rồi coi bằng evidence 3 seed ở mọi corpus. Chương 06 Table 2.
- Macro EER 0 trong offset toy có chứng minh deploy một threshold zero-error không?
- Một attack có 100 S và dùng 755 B shared; một attack khác 200 S dùng cùng 755 B. Có independent EER rows không?
- Vì sao pooled miss/FA dùng hai sets weights khác nhau?
- Nếu chỉ báo worst-group quan sát, cần nói thêm gì?
Đáp án reasoning
- Không: mỗi group oracle riêng; threshold 50 trong pooled cho 50%/50%.
- Không: shared B và có thể shared speakers/source tạo dependence. Group rows hợp lệ descriptive, không independent replicates.
- Conditional error rates chia riêng . Group prevalence mỗi lớp có thể khác; overall sample count weights không đúng cả hai.
- Definitions/counts/reference/selection status, sampling units và uncertainty. Worst observed chưa chắc worst population hoặc mọi future family.
Đào sâu: nhân số samples groupA mười lần mà giữ local distributions, tính pooled curves lại. Phân biệt thay population weights với thêm independent evidence; lặp exact clips không tạo 10×sample information.