ajaudio / studyAUDIO-JEPA · RESEARCH NOTES
4 phút đọc · Toàn văn
Mục lục bài · 6 mục

Bài 5 — Pooled, macro, per-group và worst-group#

Trước · Tiếp.

Mục tiêu và tiền đề#

Tính score-offset counterexample và phân biệt shared threshold với per-group oracle. Cần EER/ROC, risk, domain shift.

1. Ba đại lượng tưởng giống nhau#

Pooled EER: gom scores, một scale và sweep threshold chung. Macro EER=∑gwgEERg\sum_gw_gEER_g, thường wg=1/Gw_g=1/G: từng EER có oracle threshold riêng. Shared-threshold group rates: fit threshold từ dev, freeze rồi tính miss/FA từng group trên test. Cái thứ ba mô tả policy chuyển giao; hai cái đầu là ranking/sweep summaries.

Với fixed threshold:

Pmisspool=∑gNB,gNBPmiss,g,Pfapool=∑gNS,gNSPfa,g.P_{miss}^{pool}=\sum_g\frac{N_{B,g}}{N_B}P_{miss,g},\quad P_{fa}^{pool}=\sum_g\frac{N_{S,g}}{N_S}P_{fa,g}.

Weights theo lớp có thể khác nhau. EER nonlinear nên không lấy hai expressions ấy tại group-specific thresholds rồi gọi pooled EER. ASVspoof 5 Appendix 11.1 dùng pooled scores cho overall ranking; aggregate choice là protocol, không verdict mọi macro sai.

2. Worked example: hai group perfect, pooled 50%#

Toy:

GroupB scoresS scoresLocal separating threshold
A[3;4][1;2]2,5
B[103;104][101;102]102,5

Mỗi group AUC 1/EER 0, equal-weight macroEER 0. Pooled B=[3;4;103;104], S=[1;2;101;102]. Threshold 50: miss 2/4,FA 2/4 → EER 50%. Pooled AUC: low B thắng 2 S mỗi mẫu (4 wins), high B thắng 4 mỗi mẫu (8 wins), tổng 12/16=0,75. Threshold 2,5 dùng toàn bộ thì groupB FA 100%; threshold 102,5 thì groupA miss 100%. Không có perfect common threshold.

Trừ 100 chỉ groupB giúp pooledEER 0, giữ local ranking. Đây là group-conditioned offset correction, không global monotone transform. Fit offset bằng eval labels rồi gọi transferred calibration là oracle; use metadata/target access/calibration policy phải hợp protocol. Toy không chứng minh actual detector có offsets 100.

3. Weights đổi population và đôi khi đảo thứ hạng#

Toy fixed-rule group error summaries: A errors(g1,g2)=(0;0,4), B=(0,1;0,2). Equal macroA 0,2>B0,15 nên B tốt hơn. Population 90%g1: A0,04<B0,11 nên A tốt hơn. Đây là change of estimand/population weights, không arithmetic contradiction. Không gọi “trung bình” khi không nêu equal groups hay sample/population weights. Với DCF, còn lớp/cost weights.

Per-attack thường dùng S của attack g và cùng toàn bộ genuine reference. Khi genuine scores đổi, nhiều attack EER đổi cùng nhau; attack rows không independent. Shared reference hợp lệ nếu disclosed, nhưng không coi G rows là G independent experiments. Nếu group chỉ có S thì own EER undefined; cần specify reference B. Không quietly bỏ missing-class groups rồi vẫn gọi macro toàn taxonomy. Có thể báo S acceptance ở shared threshold và B miss toàn reference như hai quantities rõ hơn.

4. Worst-group, selection và uncertainty#

Worst-group score=max⁡gMg\max_gM_g với lower-is-better metric; G và definitions phải chốt. Nhóm 2 clips có 1 error cho 50%; nhóm 1.000 clips có 400 cho 40%. Không thể biết chắc population nhóm nhỏ khó hơn chỉ từ ordering sample; intervals cần đúng sampling units. Tìm subgroup hậu nghiệm trong nhiều partitions rồi báo worst/best là một selection process. Với bootstrap statistic “worst-group”, cần recompute max trên cùng predefined groups mỗi replicate; CI của group được chọn duy nhất trả lời khác và bỏ selection uncertainty. Family/domain sample ít vẫn giới hạn inference dù within-group utterance intervals hẹp.

Báo cả aggregate theo primary protocol và phân rã có denominators/reference/counts/uncertainty. Không lựa mỗi subgroup hệ mình thắng để kể generalization. Bouthillier et al.2021 §§2–3 cung cấp sources-of-variation framing; maximum/weights toys là suy luận tự biên soạn từ estimand.

5. Paper và tự kiểm#

Table 2 paper là per-corpus EER, không pooled across five corpora và không một transferred operating threshold. Không tự lấy mean năm cells khi frozenADD chưa chạy và In-the-WildFT n=1 rồi coi bằng evidence 3 seed ở mọi corpus. Chương 06 Table 2.

  1. Macro EER 0 trong offset toy có chứng minh deploy một threshold zero-error không?
  2. Một attack có 100 S và dùng 755 B shared; một attack khác 200 S dùng cùng 755 B. Có independent EER rows không?
  3. Vì sao pooled miss/FA dùng hai sets weights khác nhau?
  4. Nếu chỉ báo worst-group quan sát, cần nói thêm gì?
Đáp án reasoning
  1. Không: mỗi group oracle riêng; threshold 50 trong pooled cho 50%/50%.
  2. Không: shared B và có thể shared speakers/source tạo dependence. Group rows hợp lệ descriptive, không independent replicates.
  3. Conditional error rates chia riêng NB,NSN_B,N_S. Group prevalence mỗi lớp có thể khác; overall sample count weights không đúng cả hai.
  4. Definitions/counts/reference/selection status, sampling units và uncertainty. Worst observed chưa chắc worst population hoặc mọi future family.

Đào sâu: nhân số samples groupA mười lần mà giữ local distributions, tính pooled curves lại. Phân biệt thay population weights với thêm independent evidence; lặp exact clips không tạo 10×sample information.

DỪNG LẠI & TỰ KIỂM TRA

Bạn đã giải thích được cơ chế trong bài?

↓ Bản Markdown nguyên gốcGiữ nguyên nội dung · Công thức, bảng và nguồn đầy đủ.Các chat bàn giao được mở trong Codex.

Gõ từ khóa để tìm bài học và đoạn liên quan.