# Bài 7 — Uncertainty, paired resampling và phép so đúng câu hỏi

[Trước](06-SPLIT-LEAKAGE-SELECTION.md) · [Tiếp](08-SO-SANH-ABLATION-CO-CHE.md).

## Mục tiêu và tiền đề

Tự thực hiện group draw paired, nói CI conditional trên cái gì, chọn test/null/margin phù hợp. Cần [metrics](02-RANKING-ROC-PR-EER.md), [group aggregation](05-POOLED-MACRO-NHOM.md), [selection](06-SPLIT-LEAKAGE-SELECTION.md). Rates dùng fractions; run EER example ghi % và difference **pp**.

## 1. Không chỉ một nguồn variance

| Random object | Estimand/variation đang đo | Không tạo bằng |
|---|---|---|
| Training run R | Mean metric qua run của recipe đã chốt | Resample scores của 1 checkpoint |
| Evaluation units D | Metric của checkpoint trên population samples | SD qua 3 seeds |
| Recording/speaker cluster G | New cluster samples, within-group dependence | IID crops cùng recording |
| Generator/family/domain | Transfer qua population groups được định nghĩa | Rất nhiều clips của 2 families |
| Search procedure H | Attainable/selected pipeline qua searches/data | Thêm seeds config thắng |

[ Bouthillier 2021 §§2–3/AppendixC.2](https://arxiv.org/html/2103.03098v1) mô hình hóa các nguồn này. Same seed number ở architectures khác không tạo identical random path; pairing cần actual shared experimental units/randomization, không chỉ tên run.

**Seed toy** [6;8;10]% có mean 8%, sample SD=$\sqrt{(4+0+4)/(3−1)}=2$pp. Mean±SD=[6;10]% mô tả spread observed runs, không 95%CI. Nếu iid normal run-level outcomes và estimand mean recipe metric hợp lý, t-CI=$8\pm4,302653·2/\sqrt3≈[3,032;12,968]$%. Với 3 runs normality/tails khó kiểm; không lấy t-interval làm tự động valid. Nó không prediction interval run mới hay eval-sample interval. Bootstrap 3 points không invent information ở tails. [NIST t-CI mean](https://www.itl.nist.gov/div898/handbook/eda/section3/eda352.htm).

## 2. Worked paired group bootstrap có thể liệt kê hết

**Toy hai independent source groups**, mỗi group gồm một real và một derivative spoof; giữ relation cùng group. Scores high=B:

| Group | Label | A | B model |
|---|---|---:|---:|
| g1 | B | 2 | 2 |
| g1 | S | 1 | 3 |
| g2 | B | 1 | 2 |
| g2 | S | 2 | 1 |

Local EER g1: A0,B1 (perfectly inverted); g2: A1,B0. Equal macro EER mỗi model 0,5; $\Delta=M_A-M_B=0$. Resample two whole groups with replacement; **same draws** cho A/B, recompute group EER và equal-weight mean **theo draw multiplicities**:

| Ordered draw | A macro | B macro | Δ |
|---|---:|---:|---:|
| g1,g1 | 0 | 1 | −1 |
| g1,g2 | 0,5 | 0,5 | 0 |
| g2,g1 | 0,5 | 0,5 | 0 |
| g2,g2 | 1 | 0 | +1 |

Exact empirical bootstrap distribution có weights(1/4;1/2;1/4). Không cần 100.000 replicates để có thêm unique group information. Percentile 2,5%/97,5% theo discrete inverse-CDF cho[−1;+1]; đây là **illustration của empirical resampling**, không chứng minh nominal 95%coverage với 2 groups. Trong draw g1,g1, group g1 có multiplicity2; phải giữ weight đó. Deduplicate IDs sau draw sẽ thay estimator.

Nếu chọn **pooled** EER estimand, recompute pooled curve toàn draw thay macro. Trong toy này original pooled EER của cả A và B cũng bằng 0,5: hai estimators tình cờ trùng số, không phải identity. Counterexample score offset ở bài 5 cho per-group EER 0 nhưng pooled EER 0,5. Nếu quantity là actual error tại dev-fixed threshold, giữ threshold fixed mỗi replicate; bootstrap error bits khi đó đúng câu hỏi fixed-rule, không EER. Bootstrap EER/minDCF phải sweep lại trên mỗi replicate, để cả threshold-summary estimation biến thiên.

## 3. Resampling assumptions, stratification và degeneracy

Nonparametric bootstrap lấy empirical sample làm approximation population; cần independent/exchangeable sampling units và đủ support, estimator có regularity thích hợp. Cluster bootstrap giữ dependence **trong** cluster, không tự sửa dependence **giữa** clusters. Unequal cluster size: pooling observations với multiplicities nhắm size-weighted quantity; equal-group macro nhắm group-weighted quantity. Nếu thật sự sampled speakers rồi recordings thì nested design cần rules từng level; đừng tự resample mọi level khi sampling mechanism không có randomness đó. [Field–Welsh 2007 §§2–3,5](https://doi.org/10.1111/j.1467-9868.2007.00593.x); [Shalizi bootstrap procedure](https://www.stat.cmu.edu/~cshalizi/dst/18/lectures/18/lecture-18.html).

Draw 4 IID trials từ 2 B+2 S có probability thiếu class=$2(1/2)^4=1/8$. EER/AUC undefined trên draw ấy. Không quietly discard rồi gọi unchanged unconditional estimand. Prespecify class-stratified design nếu mục tiêu conditional class distributions và group structure cho phép; hoặc use whole groups chứa cả classes như toy; hoặc báo invalid fraction/limitations và sửa design. Không stratify files riêng nếu điều đó phá real–fake parent relation. Nếu finite dataset/IDs là toàn population mục tiêu, descriptive metric không cần sampling CI cho imagined iid population.

## 4. Seed × sample thường crossed, không nested tự động

Mỗi training run được chấm trên cùng eval groups: matrix scores $(r,g,i)$ là crossed factors R×G. Ví dụ estimand $\theta=E_R[M(f_R,D_{new})]$, delta giữa hai procedures. Toy algorithm cho empirical approximation: draw run indices theo run design; draw evaluation groups độc lập theo sampling design; dùng **cùng** group draw cho models và cho mọi run của replicate; tính metric từng run trên draw, rồi average run metrics. Nếu runs là paired experimental blocks, draw run **pairs**; nếu độc lập thì draw mỗi arm riêng. Không pool scores across seeds trừ khi estimand là ensemble/mixture mới.

[Owen–Eckles 2012 v3 §§2–4](https://arxiv.org/html/1106.2125v3) chứng minh product reweighting variance cho means trong crossed random-effects models. **Không** chứng minh percentile bootstrap EER valid cho mọi seed×cluster design. Ví dụ procedure trên là design proposal phải kiểm assumptions/coverage, nhất là ít seeds/families và nonsmooth threshold metrics. Báo separate seed spread và conditional sample interval thường rõ hơn false precision từ combined CI chưa được thẩm định.

## 5. Test chọn theo quantity và null

McNemar: cùng independent test cases, binary correct/incorrect ở **fixed dev-selected rules**. Chỉ discordant pairs hữu ích: $b$=A đúng B sai; $c$=A sai B đúng. Under null equal marginal correctness, conditional trên $n=b+c$, $b\sim Binomial(n,1/2)$. **Toy b7,c1:** exact two-sided $p=2P(X\le1)=2(1+8)/256=18/256≈0,0703125$. A có 7 wins/1 loss vẫn chưa reject ở0,05 theo test này. Same-speaker dependence phá independent-pairs assumption; không dùng McNemar cho EER. Exact conservative/mid-p/asymptotic có tradeoffs; không chọn method sau nhìn p. [Fagerland et al.2013 Methods](https://pmc.ncbi.nlm.nih.gov/articles/PMC3716987/).

Paired t-test/CI: independent run/block **differences**, mean difference dưới assumptions (normal hoặc đủ asymptotic support); không t-test mỗi utterance EER vì utterance không có EER. Paired swap/randomization test: null exchangeability of A/B outcomes trong independent experimental unit; cần swapping unit đúng group và statistic recomputed. Equality of mean alone không tự exchangeability. Label permutation hỏi X–y association khác paired system comparison; không đổi null để có p nhỏ. DeLong là method cho correlated **ROC-AUC**, không EER/minDCF; full theorem chưa đọc trong lớp này nên chỉ định vị, không dạy formula/khẳng định valid clustered case. [Nguồn và mức đọc](research/evidence-sources.md).

## 6. Không-significance, equivalence và multiplicity

Equivalence cần **margin δ prespecified theo ích lợi thực tế**. Với Δ ởpp, null union: $\Delta\le-\delta$ hoặc $\Delta\ge\delta$; TOST bác bỏ cả hai one-sided nulls. Với phù hợp t-model/α0,05,90%CI nằm trọn trong bounds là corresponding criterion;95%CI chứa 0 không chứng minh equivalent. Toy margin±1 pp:90%CI[−0,4;+0,6] đáp ứng; [−1,2;+0,6] không đáp ứng dù contains 0. Đó là interval-logic toy, chưa statistical test trên detector. [Lakens 2017 dependent-means/TOST](https://pmc.ncbi.nlm.nih.gov/articles/PMC5502906/).

20 independent true-null tests mỗi α0,05: probability≥1 false positive=$1−0,95^{20}≈0,641514$. Holm prespecified family 3 valid p-values [0,01;0,03;0,04]: compare sorted with 0,05/3,0,05/2,0,05. Reject first; stop at 0,03>0,025 nên remaining không reject. [Holm 1979 §2 theorem](https://doi.org/10.2307/4615733). Không cần independent tests cho Holm union-bound control với valid p-values. Adaptive config feedback/search trên test có thể làm base p-values invalid; multiplicity correction không khôi phục holdout independence.

## 7. Paper và tự kiểm

Paper Table 4 báo 3 seeds mỗi comparison, ranges descriptive; In-the-WildFT Table 2 là n=1. Không non-overlap range significance, không same SD equivalence, không seed SD thay generator uncertainty. [Chương 06](../06-GIAI-PHAU-PAPER.md).

1. Resample cùng IDs cho A/B có nghĩa bootstrap hai files độc lập rồi align sau được không?
2. Vì sao same scores 1 model lặp 100 lần không training uncertainty?
3. Muốn CI của EER có được fixed-error-bit bootstrap không?
4. Toy group draw g1,g1 có được deduplicate không?
5. p0,2 superiority có chứng minh within±1 pp chưa?

<details>
<summary>Đáp án reasoning</summary>

1. Không; phải common sampling draw trước, giữ pair/covariance và multiplicity.
2. Training weights/random paths không đổi; chỉ resample measurement nếu units khác.
3. Không; bits fixed threshold trả lời rule error, cần recompute EER curve.
4. Không; duplicate **draw** là bootstrap weight, khác unwanted duplicate audio trong original protocol.
5. Chưa; cần margin/appropriately justified equivalence procedure và đủ precision. Non-rejection có thể chỉ low power.

</details>

**Đào sâu:** tự liệt kê $2^G$ paired group swap assignments cho một fixed metric, ghi strong null/exchangeability, rồi đối chiếu bootstrap interval (sampling) với randomization test (null distribution). Chúng không interchangeable.
