Bài 7 — Uncertainty, paired resampling và phép so đúng câu hỏi#
Mục tiêu và tiền đề#
Tự thực hiện group draw paired, nói CI conditional trên cái gì, chọn test/null/margin phù hợp. Cần metrics, group aggregation, selection. Rates dùng fractions; run EER example ghi % và difference pp.
1. Không chỉ một nguồn variance#
| Random object | Estimand/variation đang đo | Không tạo bằng |
|---|---|---|
| Training run R | Mean metric qua run của recipe đã chốt | Resample scores của 1 checkpoint |
| Evaluation units D | Metric của checkpoint trên population samples | SD qua 3 seeds |
| Recording/speaker cluster G | New cluster samples, within-group dependence | IID crops cùng recording |
| Generator/family/domain | Transfer qua population groups được định nghĩa | Rất nhiều clips của 2 families |
| Search procedure H | Attainable/selected pipeline qua searches/data | Thêm seeds config thắng |
Bouthillier 2021 §§2–3/AppendixC.2 mô hình hóa các nguồn này. Same seed number ở architectures khác không tạo identical random path; pairing cần actual shared experimental units/randomization, không chỉ tên run.
Seed toy [6;8;10]% có mean 8%, sample SD=pp. Mean±SD=[6;10]% mô tả spread observed runs, không 95%CI. Nếu iid normal run-level outcomes và estimand mean recipe metric hợp lý, t-CI=%. Với 3 runs normality/tails khó kiểm; không lấy t-interval làm tự động valid. Nó không prediction interval run mới hay eval-sample interval. Bootstrap 3 points không invent information ở tails. NIST t-CI mean.
2. Worked paired group bootstrap có thể liệt kê hết#
Toy hai independent source groups, mỗi group gồm một real và một derivative spoof; giữ relation cùng group. Scores high=B:
| Group | Label | A | B model |
|---|---|---|---|
| g1 | B | 2 | 2 |
| g1 | S | 1 | 3 |
| g2 | B | 1 | 2 |
| g2 | S | 2 | 1 |
Local EER g1: A0,B1 (perfectly inverted); g2: A1,B0. Equal macro EER mỗi model 0,5; . Resample two whole groups with replacement; same draws cho A/B, recompute group EER và equal-weight mean theo draw multiplicities:
| Ordered draw | A macro | B macro | Δ |
|---|---|---|---|
| g1,g1 | 0 | 1 | −1 |
| g1,g2 | 0,5 | 0,5 | 0 |
| g2,g1 | 0,5 | 0,5 | 0 |
| g2,g2 | 1 | 0 | +1 |
Exact empirical bootstrap distribution có weights(1/4;1/2;1/4). Không cần 100.000 replicates để có thêm unique group information. Percentile 2,5%/97,5% theo discrete inverse-CDF cho[−1;+1]; đây là illustration của empirical resampling, không chứng minh nominal 95%coverage với 2 groups. Trong draw g1,g1, group g1 có multiplicity2; phải giữ weight đó. Deduplicate IDs sau draw sẽ thay estimator.
Nếu chọn pooled EER estimand, recompute pooled curve toàn draw thay macro. Trong toy này original pooled EER của cả A và B cũng bằng 0,5: hai estimators tình cờ trùng số, không phải identity. Counterexample score offset ở bài 5 cho per-group EER 0 nhưng pooled EER 0,5. Nếu quantity là actual error tại dev-fixed threshold, giữ threshold fixed mỗi replicate; bootstrap error bits khi đó đúng câu hỏi fixed-rule, không EER. Bootstrap EER/minDCF phải sweep lại trên mỗi replicate, để cả threshold-summary estimation biến thiên.
3. Resampling assumptions, stratification và degeneracy#
Nonparametric bootstrap lấy empirical sample làm approximation population; cần independent/exchangeable sampling units và đủ support, estimator có regularity thích hợp. Cluster bootstrap giữ dependence trong cluster, không tự sửa dependence giữa clusters. Unequal cluster size: pooling observations với multiplicities nhắm size-weighted quantity; equal-group macro nhắm group-weighted quantity. Nếu thật sự sampled speakers rồi recordings thì nested design cần rules từng level; đừng tự resample mọi level khi sampling mechanism không có randomness đó. Field–Welsh 2007 §§2–3,5; Shalizi bootstrap procedure.
Draw 4 IID trials từ 2 B+2 S có probability thiếu class=. EER/AUC undefined trên draw ấy. Không quietly discard rồi gọi unchanged unconditional estimand. Prespecify class-stratified design nếu mục tiêu conditional class distributions và group structure cho phép; hoặc use whole groups chứa cả classes như toy; hoặc báo invalid fraction/limitations và sửa design. Không stratify files riêng nếu điều đó phá real–fake parent relation. Nếu finite dataset/IDs là toàn population mục tiêu, descriptive metric không cần sampling CI cho imagined iid population.
4. Seed × sample thường crossed, không nested tự động#
Mỗi training run được chấm trên cùng eval groups: matrix scores là crossed factors R×G. Ví dụ estimand , delta giữa hai procedures. Toy algorithm cho empirical approximation: draw run indices theo run design; draw evaluation groups độc lập theo sampling design; dùng cùng group draw cho models và cho mọi run của replicate; tính metric từng run trên draw, rồi average run metrics. Nếu runs là paired experimental blocks, draw run pairs; nếu độc lập thì draw mỗi arm riêng. Không pool scores across seeds trừ khi estimand là ensemble/mixture mới.
Owen–Eckles 2012 v3 §§2–4 chứng minh product reweighting variance cho means trong crossed random-effects models. Không chứng minh percentile bootstrap EER valid cho mọi seed×cluster design. Ví dụ procedure trên là design proposal phải kiểm assumptions/coverage, nhất là ít seeds/families và nonsmooth threshold metrics. Báo separate seed spread và conditional sample interval thường rõ hơn false precision từ combined CI chưa được thẩm định.
5. Test chọn theo quantity và null#
McNemar: cùng independent test cases, binary correct/incorrect ở fixed dev-selected rules. Chỉ discordant pairs hữu ích: =A đúng B sai; =A sai B đúng. Under null equal marginal correctness, conditional trên , . Toy b7,c1: exact two-sided . A có 7 wins/1 loss vẫn chưa reject ở0,05 theo test này. Same-speaker dependence phá independent-pairs assumption; không dùng McNemar cho EER. Exact conservative/mid-p/asymptotic có tradeoffs; không chọn method sau nhìn p. Fagerland et al.2013 Methods.
Paired t-test/CI: independent run/block differences, mean difference dưới assumptions (normal hoặc đủ asymptotic support); không t-test mỗi utterance EER vì utterance không có EER. Paired swap/randomization test: null exchangeability of A/B outcomes trong independent experimental unit; cần swapping unit đúng group và statistic recomputed. Equality of mean alone không tự exchangeability. Label permutation hỏi X–y association khác paired system comparison; không đổi null để có p nhỏ. DeLong là method cho correlated ROC-AUC, không EER/minDCF; full theorem chưa đọc trong lớp này nên chỉ định vị, không dạy formula/khẳng định valid clustered case. Nguồn và mức đọc.
6. Không-significance, equivalence và multiplicity#
Equivalence cần margin δ prespecified theo ích lợi thực tế. Với Δ ởpp, null union: hoặc ; TOST bác bỏ cả hai one-sided nulls. Với phù hợp t-model/α0,05,90%CI nằm trọn trong bounds là corresponding criterion;95%CI chứa 0 không chứng minh equivalent. Toy margin±1 pp:90%CI[−0,4;+0,6] đáp ứng; [−1,2;+0,6] không đáp ứng dù contains 0. Đó là interval-logic toy, chưa statistical test trên detector. Lakens 2017 dependent-means/TOST.
20 independent true-null tests mỗi α0,05: probability≥1 false positive=. Holm prespecified family 3 valid p-values [0,01;0,03;0,04]: compare sorted with 0,05/3,0,05/2,0,05. Reject first; stop at 0,03>0,025 nên remaining không reject. Holm 1979 §2 theorem. Không cần independent tests cho Holm union-bound control với valid p-values. Adaptive config feedback/search trên test có thể làm base p-values invalid; multiplicity correction không khôi phục holdout independence.
7. Paper và tự kiểm#
Paper Table 4 báo 3 seeds mỗi comparison, ranges descriptive; In-the-WildFT Table 2 là n=1. Không non-overlap range significance, không same SD equivalence, không seed SD thay generator uncertainty. Chương 06.
- Resample cùng IDs cho A/B có nghĩa bootstrap hai files độc lập rồi align sau được không?
- Vì sao same scores 1 model lặp 100 lần không training uncertainty?
- Muốn CI của EER có được fixed-error-bit bootstrap không?
- Toy group draw g1,g1 có được deduplicate không?
- p0,2 superiority có chứng minh within±1 pp chưa?
Đáp án reasoning
- Không; phải common sampling draw trước, giữ pair/covariance và multiplicity.
- Training weights/random paths không đổi; chỉ resample measurement nếu units khác.
- Không; bits fixed threshold trả lời rule error, cần recompute EER curve.
- Không; duplicate draw là bootstrap weight, khác unwanted duplicate audio trong original protocol.
- Chưa; cần margin/appropriately justified equivalence procedure và đủ precision. Non-rejection có thể chỉ low power.
Đào sâu: tự liệt kê paired group swap assignments cho một fixed metric, ghi strong null/exchangeability, rồi đối chiếu bootstrap interval (sampling) với randomization test (null distribution). Chúng không interchangeable.