ajaudio / studyAUDIO-JEPA · RESEARCH NOTES
7 phút đọc · Toàn văn
Mục lục bài · 8 mục

Bài 7 — Uncertainty, paired resampling và phép so đúng câu hỏi#

Trước · Tiếp.

Mục tiêu và tiền đề#

Tự thực hiện group draw paired, nói CI conditional trên cái gì, chọn test/null/margin phù hợp. Cần metrics, group aggregation, selection. Rates dùng fractions; run EER example ghi % và difference pp.

1. Không chỉ một nguồn variance#

Random objectEstimand/variation đang đoKhông tạo bằng
Training run RMean metric qua run của recipe đã chốtResample scores của 1 checkpoint
Evaluation units DMetric của checkpoint trên population samplesSD qua 3 seeds
Recording/speaker cluster GNew cluster samples, within-group dependenceIID crops cùng recording
Generator/family/domainTransfer qua population groups được định nghĩaRất nhiều clips của 2 families
Search procedure HAttainable/selected pipeline qua searches/dataThêm seeds config thắng

Bouthillier 2021 §§2–3/AppendixC.2 mô hình hóa các nguồn này. Same seed number ở architectures khác không tạo identical random path; pairing cần actual shared experimental units/randomization, không chỉ tên run.

Seed toy [6;8;10]% có mean 8%, sample SD=(4+0+4)/(3−1)=2\sqrt{(4+0+4)/(3−1)}=2pp. Mean±SD=[6;10]% mô tả spread observed runs, không 95%CI. Nếu iid normal run-level outcomes và estimand mean recipe metric hợp lý, t-CI=8±4,302653⋅2/3≈[3,032;12,968]8\pm4,302653·2/\sqrt3≈[3,032;12,968]%. Với 3 runs normality/tails khó kiểm; không lấy t-interval làm tự động valid. Nó không prediction interval run mới hay eval-sample interval. Bootstrap 3 points không invent information ở tails. NIST t-CI mean.

2. Worked paired group bootstrap có thể liệt kê hết#

Toy hai independent source groups, mỗi group gồm một real và một derivative spoof; giữ relation cùng group. Scores high=B:

GroupLabelAB model
g1B22
g1S13
g2B12
g2S21

Local EER g1: A0,B1 (perfectly inverted); g2: A1,B0. Equal macro EER mỗi model 0,5; Δ=MA−MB=0\Delta=M_A-M_B=0. Resample two whole groups with replacement; same draws cho A/B, recompute group EER và equal-weight mean theo draw multiplicities:

Ordered drawA macroB macroΔ
g1,g101−1
g1,g20,50,50
g2,g10,50,50
g2,g210+1

Exact empirical bootstrap distribution có weights(1/4;1/2;1/4). Không cần 100.000 replicates để có thêm unique group information. Percentile 2,5%/97,5% theo discrete inverse-CDF cho[−1;+1]; đây là illustration của empirical resampling, không chứng minh nominal 95%coverage với 2 groups. Trong draw g1,g1, group g1 có multiplicity2; phải giữ weight đó. Deduplicate IDs sau draw sẽ thay estimator.

Nếu chọn pooled EER estimand, recompute pooled curve toàn draw thay macro. Trong toy này original pooled EER của cả A và B cũng bằng 0,5: hai estimators tình cờ trùng số, không phải identity. Counterexample score offset ở bài 5 cho per-group EER 0 nhưng pooled EER 0,5. Nếu quantity là actual error tại dev-fixed threshold, giữ threshold fixed mỗi replicate; bootstrap error bits khi đó đúng câu hỏi fixed-rule, không EER. Bootstrap EER/minDCF phải sweep lại trên mỗi replicate, để cả threshold-summary estimation biến thiên.

3. Resampling assumptions, stratification và degeneracy#

Nonparametric bootstrap lấy empirical sample làm approximation population; cần independent/exchangeable sampling units và đủ support, estimator có regularity thích hợp. Cluster bootstrap giữ dependence trong cluster, không tự sửa dependence giữa clusters. Unequal cluster size: pooling observations với multiplicities nhắm size-weighted quantity; equal-group macro nhắm group-weighted quantity. Nếu thật sự sampled speakers rồi recordings thì nested design cần rules từng level; đừng tự resample mọi level khi sampling mechanism không có randomness đó. Field–Welsh 2007 §§2–3,5; Shalizi bootstrap procedure.

Draw 4 IID trials từ 2 B+2 S có probability thiếu class=2(1/2)4=1/82(1/2)^4=1/8. EER/AUC undefined trên draw ấy. Không quietly discard rồi gọi unchanged unconditional estimand. Prespecify class-stratified design nếu mục tiêu conditional class distributions và group structure cho phép; hoặc use whole groups chứa cả classes như toy; hoặc báo invalid fraction/limitations và sửa design. Không stratify files riêng nếu điều đó phá real–fake parent relation. Nếu finite dataset/IDs là toàn population mục tiêu, descriptive metric không cần sampling CI cho imagined iid population.

4. Seed × sample thường crossed, không nested tự động#

Mỗi training run được chấm trên cùng eval groups: matrix scores (r,g,i)(r,g,i) là crossed factors R×G. Ví dụ estimand θ=ER[M(fR,Dnew)]\theta=E_R[M(f_R,D_{new})], delta giữa hai procedures. Toy algorithm cho empirical approximation: draw run indices theo run design; draw evaluation groups độc lập theo sampling design; dùng cùng group draw cho models và cho mọi run của replicate; tính metric từng run trên draw, rồi average run metrics. Nếu runs là paired experimental blocks, draw run pairs; nếu độc lập thì draw mỗi arm riêng. Không pool scores across seeds trừ khi estimand là ensemble/mixture mới.

Owen–Eckles 2012 v3 §§2–4 chứng minh product reweighting variance cho means trong crossed random-effects models. Không chứng minh percentile bootstrap EER valid cho mọi seed×cluster design. Ví dụ procedure trên là design proposal phải kiểm assumptions/coverage, nhất là ít seeds/families và nonsmooth threshold metrics. Báo separate seed spread và conditional sample interval thường rõ hơn false precision từ combined CI chưa được thẩm định.

5. Test chọn theo quantity và null#

McNemar: cùng independent test cases, binary correct/incorrect ở fixed dev-selected rules. Chỉ discordant pairs hữu ích: bb=A đúng B sai; cc=A sai B đúng. Under null equal marginal correctness, conditional trên n=b+cn=b+c, b∼Binomial(n,1/2)b\sim Binomial(n,1/2). Toy b7,c1: exact two-sided p=2P(X≤1)=2(1+8)/256=18/256≈0,0703125p=2P(X\le1)=2(1+8)/256=18/256≈0,0703125. A có 7 wins/1 loss vẫn chưa reject ở0,05 theo test này. Same-speaker dependence phá independent-pairs assumption; không dùng McNemar cho EER. Exact conservative/mid-p/asymptotic có tradeoffs; không chọn method sau nhìn p. Fagerland et al.2013 Methods.

Paired t-test/CI: independent run/block differences, mean difference dưới assumptions (normal hoặc đủ asymptotic support); không t-test mỗi utterance EER vì utterance không có EER. Paired swap/randomization test: null exchangeability of A/B outcomes trong independent experimental unit; cần swapping unit đúng group và statistic recomputed. Equality of mean alone không tự exchangeability. Label permutation hỏi X–y association khác paired system comparison; không đổi null để có p nhỏ. DeLong là method cho correlated ROC-AUC, không EER/minDCF; full theorem chưa đọc trong lớp này nên chỉ định vị, không dạy formula/khẳng định valid clustered case. Nguồn và mức đọc.

6. Không-significance, equivalence và multiplicity#

Equivalence cần margin δ prespecified theo ích lợi thực tế. Với Δ ởpp, null union: Δ≤−δ\Delta\le-\delta hoặc Δ≥δ\Delta\ge\delta; TOST bác bỏ cả hai one-sided nulls. Với phù hợp t-model/α0,05,90%CI nằm trọn trong bounds là corresponding criterion;95%CI chứa 0 không chứng minh equivalent. Toy margin±1 pp:90%CI[−0,4;+0,6] đáp ứng; [−1,2;+0,6] không đáp ứng dù contains 0. Đó là interval-logic toy, chưa statistical test trên detector. Lakens 2017 dependent-means/TOST.

20 independent true-null tests mỗi α0,05: probability≥1 false positive=1−0,9520≈0,6415141−0,95^{20}≈0,641514. Holm prespecified family 3 valid p-values [0,01;0,03;0,04]: compare sorted with 0,05/3,0,05/2,0,05. Reject first; stop at 0,03>0,025 nên remaining không reject. Holm 1979 §2 theorem. Không cần independent tests cho Holm union-bound control với valid p-values. Adaptive config feedback/search trên test có thể làm base p-values invalid; multiplicity correction không khôi phục holdout independence.

7. Paper và tự kiểm#

Paper Table 4 báo 3 seeds mỗi comparison, ranges descriptive; In-the-WildFT Table 2 là n=1. Không non-overlap range significance, không same SD equivalence, không seed SD thay generator uncertainty. Chương 06.

  1. Resample cùng IDs cho A/B có nghĩa bootstrap hai files độc lập rồi align sau được không?
  2. Vì sao same scores 1 model lặp 100 lần không training uncertainty?
  3. Muốn CI của EER có được fixed-error-bit bootstrap không?
  4. Toy group draw g1,g1 có được deduplicate không?
  5. p0,2 superiority có chứng minh within±1 pp chưa?
Đáp án reasoning
  1. Không; phải common sampling draw trước, giữ pair/covariance và multiplicity.
  2. Training weights/random paths không đổi; chỉ resample measurement nếu units khác.
  3. Không; bits fixed threshold trả lời rule error, cần recompute EER curve.
  4. Không; duplicate draw là bootstrap weight, khác unwanted duplicate audio trong original protocol.
  5. Chưa; cần margin/appropriately justified equivalence procedure và đủ precision. Non-rejection có thể chỉ low power.

Đào sâu: tự liệt kê 2G2^G paired group swap assignments cho một fixed metric, ghi strong null/exchangeability, rồi đối chiếu bootstrap interval (sampling) với randomization test (null distribution). Chúng không interchangeable.

DỪNG LẠI & TỰ KIỂM TRA

Bạn đã giải thích được cơ chế trong bài?

↓ Bản Markdown nguyên gốcGiữ nguyên nội dung · Công thức, bảng và nguồn đầy đủ.Các chat bàn giao được mở trong Codex.

Gõ từ khóa để tìm bài học và đoạn liên quan.