Bài 6 — Split, target access và model selection#
Mục tiêu và tiền đề#
Thiết kế split đúng claim, phân biệt exposure/ancestry với exact leakage, và tự tính selection optimism. Cần passport/protocol, cue/confound, estimand.
1. “Unseen” là một ràng buộc có granularity#
Utterance-disjoint chặn exact file lặp, chưa chặn source-recording crop/re-encode. Speaker-disjoint kiểm identity mới, chưa generator mới. Generator checkpoint-disjoint khác family/vocoder-disjoint. Source/channel/language/time/compositional split đo các shifts khác nhau. Compositional holdout có thể chỉ giữ tổ hợp mới, từng thành phần đã thấy. Exposure từ pretraining kiểm riêng downstream train; nó hợp/không hợp còn tùy claim và protocol. ASVspoof 5 v0.6 §§4.2–4.3 có rules riêng về speaker/resources/test normalization; không áp chúng như luật mọi nghiên cứu.
Worked manifest toy:
| File | Parent recording | Speaker | Process | Channel | Role |
|---|---|---|---|---|---|
| real-u1 | u1 | p1 | genuine | clean | train |
| fake-u1-gA | u1 | p1 | gA/vocoderV | clean | test |
| crop-u1 | u1 | p1 | genuine crop | MP 3 | dev |
| fake-u2-gB | u2 | p2 | gB/vocoderV | clean | test |
| fake-u3-gC | u3 | p3 | gC/vocoderW | MP 3 | test |
File names/byte hashes đều có thể khác. Source-disjoint claim đòi grouping ba descendants u1 vào cùng role. Speaker-disjoint còn nối các recordings khác của p1. GeneratorB mới vẫn dùngV; không vocoder-unseen. C đổi generator/vocoder/channel cùng lúc; kém trênC chưa chọn được nguyên nhân.
Một cách construct split: graph files, cạnh chung parent recording (thêm speaker nếu claim cần); connected components được chia nguyên khối. Group theo mọi yếu tố đồng thời có thể nối graph thành một giant component, không đủ partitions; phải sửa research question/data design thay vì random split rồi tuyên bố disjoint.
flowchart LR U[Source u1] --> R[real-u1] U --> F[fake-u1-gA] U --> C[crop-u1-MP3] P[Speaker p1] --> U R --> G[Group component u1: cùng partition] F --> G C --> G
Xem mã sơ đồ
flowchart LR
U[Source u1] --> R[real-u1]
U --> F[fake-u1-gA]
U --> C[crop-u1-MP3]
P[Speaker p1] --> U
R --> G[Group component u1: cùng partition]
F --> G
C --> GAncestry chung chỉ chỉ đường audit: shared corpus → speaker → transcript → recording → exact/near duplicate là levels khác. Cùng transcript có thể hai bản thu độc lập; khác hash có thể cùng audio qua nén. Đối chiếu manifest/version/source IDs/hashes/fingerprints mới thêm evidence. Kapoor–Narayanan 2022 v1 §2.4 nối nonindependence/test sampling với scientific claim; học liệu không kết luận corpus cụ thể đã overlap khi chưa audit.
2. Vai trò dữ liệu và quyền nhìn target#
Train dùng labels cập nhật weights; development chọn config/checkpoint/calibrator/threshold; final holdout kiểm quy trình đã chốt. Target-unlabelled adaptation, using target metadata, transductive normalization và target-labelled tuning là access regimes khác nhau. Có thể là research protocol hợp lệ nếu disclosed/permitted, nhưng không cùng zero-target-access estimand. Không gọi mọi target exposure gian lận; cũng không gọi test untouched nếu đã dùng feedback chọn design.
EER/minDCF dùng eval labels để tính summary là legitimate benchmark measurement. Dùng summary ấy chọn lần sau LR/head/score polarity làm evaluation feedback thành selection. Fit deploy threshold trên eval rồi chấm chính eval là oracle policy assessment. Ba hành động này cần ba mô tả riêng.
3. Selection winner's curse bằng bốn trường hợp#
Toy: hai configs có cùng true error 10%. Mỗi dev estimator independently nhận 8% hoặc 12% với xác suất 1/2. Chọn config có observed error thấp hơn:
| Dev A | Dev B | Report winner dev |
|---|---|---|
| 8 | 8 | 8 |
| 8 | 12 | 8 |
| 12 | 8 | 8 |
| 12 | 12 | 12 |
Expected winning dev error 9%, mặc dù true error vẫn 10%. Selection đã chọn cả negative estimation noise. Thêm configs có thể tăng cơ hội thắng nhờ noise; uncertainty của winner không chỉ seed variation của config thắng. Independent new holdout có expected 10% trong toy. Cawley–Talbot 2010 §§4.1,5 phân tích finite-sample selection criterion; toy là derivation tự biên soạn.
Nested/group CV có outer test groups không tham gia bất kỳ inner selection/calibration/preprocess fitting nào; outer loop đánh giá learning+selection procedure ở size đó. Fit normalization feature toàn dataset trước split phá boundary, dù labels không dùng; permitted transductive protocol lại là estimand khác. Một final holdout mới cần được khóa trước feedback; đổi tên test sau khi nhìn không khôi phục independence.
4. Historical claim được giới hạn thế nào?#
PDF đã nộp §3.1 (p4) nói exploratory native-input comparison ảnh hưởng lựa chọn finer resolution; §4.1 (p6) nói exploratory variants scored on evaluation data và informed configuration choices. §8 (p12) nói eval sets dùng inform config không held out from development. Main runs vẫn dùng final 6 epoch, A05 monitoring không chọn checkpoint. Đây là reported development history, không re-run/audit toàn log. Chương 06 §§3,6 dẫn notes lịch sử.
Kết luận đủ evidence: một phần evaluation feedback đã tham gia development, vì vậy không mô tả các sets đó như untouched confirmatory holdout. Chưa có audited inventory đủ để nói config nào nhìn từng corpus mấy lần, mọi table bị biased bao nhiêu, hay mọi eval sample đã xuất hiện trong training. Source ancestry không chứng minh exact contamination. Các notebook template không outputs cũng không xác nhận hoặc phủ định history run trên máy khác.
5. Counterexample và tự kiểm#
Train split cùng dataset với eval split có thể valid in-domain evaluation; không vì chung tên corpus mà invalid. Nhưng đổi claim thành unseen-source khi genuine source trùng lại vượt evidence. Đó là mismatch claim/split, cần sửa scope hoặc thiết kế.
- Dời crop-u1 từ dev sang test có giải quyết source-disjoint không?
- EER sweep eval và chọn LR theo eval EER khác nhau ở đâu?
- Vì sao winner dev 9% trong toy không true improvement 1 pp?
- Có thể nói mọi corpus paper đã exact leakage từ §4.1 không?
Đáp án reasoning
- Không; shared parent u1 vẫn nối train/test. Chia component trước selection.
- Một hành động định nghĩa measurement; hành động kia adapt system/hypothesis theo feedback, đổi independence của final assessment.
- Cả configs true 10; minimum noisy estimates thiên xuống. Test mới kiểm procedure đã chọn.
- Không. §4.1 hỗ trợ exploratory selection history; exact/near-duplicate/training exposure cần data evidence riêng.
Đào sâu: tự lập access matrix gồm labels, audio, metadata, statistics cho train/dev/target/final; chốt những gì được dùng ở từng step. Khi adaptive feedback diễn ra, ghi date/decision thay vì chỉ dataset names.