# Bài 5 — Pooled, macro, per-group và worst-group

[Trước](04-TANDEM-SASV-LOCALIZATION.md) · [Tiếp](06-SPLIT-LEAKAGE-SELECTION.md).

## Mục tiêu và tiền đề

Tính score-offset counterexample và phân biệt shared threshold với per-group oracle. Cần [EER/ROC](02-RANKING-ROC-PR-EER.md), [risk](03-DCF-LLR-CALIBRATION.md), [domain shift](../lop-02-deepfake-bai-toan/08-CUE-CONFOUND-DOMAIN-SHIFT.md).

## 1. Ba đại lượng tưởng giống nhau

Pooled EER: gom scores, một scale và sweep threshold chung. Macro EER=$\sum_gw_gEER_g$, thường $w_g=1/G$: từng EER có oracle threshold riêng. Shared-threshold group rates: fit threshold từ dev, freeze rồi tính miss/FA từng group trên test. Cái thứ ba mô tả policy chuyển giao; hai cái đầu là ranking/sweep summaries.

Với fixed threshold:

$$P_{miss}^{pool}=\sum_g\frac{N_{B,g}}{N_B}P_{miss,g},\quad
P_{fa}^{pool}=\sum_g\frac{N_{S,g}}{N_S}P_{fa,g}.$$

Weights **theo lớp** có thể khác nhau. EER nonlinear nên không lấy hai expressions ấy tại group-specific thresholds rồi gọi pooled EER. [ASVspoof 5 Appendix 11.1](https://www.asvspoof.org/file/ASVspoof5___Evaluation_Plan_Phase2.pdf) dùng pooled scores cho overall ranking; aggregate choice là protocol, không verdict mọi macro sai.

## 2. Worked example: hai group perfect, pooled 50%

**Toy**:

| Group | B scores | S scores | Local separating threshold |
|---|---|---|---:|
| A | [3;4] | [1;2] | 2,5 |
| B | [103;104] | [101;102] | 102,5 |

Mỗi group AUC 1/EER 0, equal-weight macroEER 0. Pooled B=[3;4;103;104], S=[1;2;101;102]. Threshold 50: miss 2/4,FA 2/4 → EER 50%. Pooled AUC: low B thắng 2 S mỗi mẫu (4 wins), high B thắng 4 mỗi mẫu (8 wins), tổng 12/16=0,75. Threshold 2,5 dùng toàn bộ thì groupB FA 100%; threshold 102,5 thì groupA miss 100%. Không có perfect common threshold.

Trừ 100 chỉ groupB giúp pooledEER 0, giữ local ranking. Đây là group-conditioned offset correction, không global monotone transform. Fit offset bằng eval labels rồi gọi transferred calibration là oracle; use metadata/target access/calibration policy phải hợp protocol. Toy không chứng minh actual detector có offsets 100.

## 3. Weights đổi population và đôi khi đảo thứ hạng

**Toy fixed-rule group error summaries**: A errors(g1,g2)=(0;0,4), B=(0,1;0,2). Equal macroA 0,2>B0,15 nên B tốt hơn. Population 90%g1: A0,04<B0,11 nên A tốt hơn. Đây là change of estimand/population weights, không arithmetic contradiction. Không gọi “trung bình” khi không nêu equal groups hay sample/population weights. Với DCF, còn lớp/cost weights.

Per-attack thường dùng S của attack g và **cùng toàn bộ genuine reference**. Khi genuine scores đổi, nhiều attack EER đổi cùng nhau; attack rows không independent. Shared reference hợp lệ nếu disclosed, nhưng không coi G rows là G independent experiments. Nếu group chỉ có S thì own EER undefined; cần specify reference B. Không quietly bỏ missing-class groups rồi vẫn gọi macro toàn taxonomy. Có thể báo S acceptance ở shared threshold và B miss toàn reference như hai quantities rõ hơn.

## 4. Worst-group, selection và uncertainty

Worst-group score=$\max_gM_g$ với lower-is-better metric; G và definitions phải chốt. Nhóm 2 clips có 1 error cho 50%; nhóm 1.000 clips có 400 cho 40%. Không thể biết chắc population nhóm nhỏ khó hơn chỉ từ ordering sample; intervals cần đúng sampling units. Tìm subgroup hậu nghiệm trong nhiều partitions rồi báo worst/best là một selection process. Với bootstrap statistic “worst-group”, cần recompute max trên **cùng predefined groups** mỗi replicate; CI của group được chọn duy nhất trả lời khác và bỏ selection uncertainty. Family/domain sample ít vẫn giới hạn inference dù within-group utterance intervals hẹp.

Báo cả aggregate theo primary protocol và phân rã có denominators/reference/counts/uncertainty. Không lựa mỗi subgroup hệ mình thắng để kể generalization. [Bouthillier et al.2021 §§2–3](https://arxiv.org/html/2103.03098v1) cung cấp sources-of-variation framing; maximum/weights toys là suy luận tự biên soạn từ estimand.

## 5. Paper và tự kiểm

Table 2 paper là **per-corpus EER**, không pooled across five corpora và không một transferred operating threshold. Không tự lấy mean năm cells khi frozenADD chưa chạy và In-the-WildFT n=1 rồi coi bằng evidence 3 seed ở mọi corpus. [Chương 06 Table 2](../06-GIAI-PHAU-PAPER.md).

1. Macro EER 0 trong offset toy có chứng minh deploy một threshold zero-error không?
2. Một attack có 100 S và dùng 755 B shared; một attack khác 200 S dùng cùng 755 B. Có independent EER rows không?
3. Vì sao pooled miss/FA dùng hai sets weights khác nhau?
4. Nếu chỉ báo worst-group quan sát, cần nói thêm gì?

<details>
<summary>Đáp án reasoning</summary>

1. Không: mỗi group oracle riêng; threshold 50 trong pooled cho 50%/50%.
2. Không: shared B và có thể shared speakers/source tạo dependence. Group rows hợp lệ descriptive, không independent replicates.
3. Conditional error rates chia riêng $N_B,N_S$. Group prevalence mỗi lớp có thể khác; overall sample count weights không đúng cả hai.
4. Definitions/counts/reference/selection status, sampling units và uncertainty. Worst observed chưa chắc worst population hoặc mọi future family.

</details>

**Đào sâu:** nhân số samples groupA mười lần mà giữ local distributions, tính pooled curves lại. Phân biệt thay population weights với thêm independent evidence; lặp exact clips không tạo 10×sample information.
