# Bài 4 — Dữ liệu, protocol và baseline đang kiểm điều gì?

[Trước](03-ADAPTATION-TRAINING.md) · [Tiếp](05-DOC-BANG-KET-QUA.md).

## Mục tiêu và tiền đề

Đọc corpus theo label/process/source/exposure; phân biệt metric gate, audio gate và model-wrapper gate. Cần [hộ chiếu lớp 2](../lop-02-deepfake-bai-toan/12-HO-CHIEU-DATASET.md), [cue/confound](../lop-02-deepfake-bai-toan/08-CUE-CONFOUND-DOMAIN-SHIFT.md), [split/access lớp 5](../lop-05-danh-gia-thuc-nghiem/06-SPLIT-LEAKAGE-SELECTION.md), [provenance](../lop-05-danh-gia-thuc-nghiem/09-PROVENANCE-REPRODUCIBILITY.md). Counts là **P**, PDF §4.2/Table 1 trang 7; original/copy sources ở [dossier](research/evidence-sources.md).

## 1. A05 holdout giữ ngoài train cái gì?

ASV2019 LA train có 20 speakers và attacks A01–A06. Paper giữ A05+600B để monitor, còn năm attack IDs trong detector training. A05 là VC dùng VAE và WORLD; A02/A03 cũng dùng WORLD. A06 còn thuộc VC nhưng dùng GMM-UBM/LPCC/MFCC và spectral filtering/overlap-add, khác pipeline A05. Đây là **attack-ID holdout**, chưa phải family/vocoder/speaker holdout. Genuine holdout cấp utterance chưa tự tạo speaker-disjoint dev.

Evaluation có 13 IDs A07–A19 nhưng A16/A19 tái dùng thuật toán A04/A06: **11 unseen algorithms** khác **13 IDs mới**. [ASVspoof2019 v4, Table 1 và §3.1](https://arxiv.org/html/1911.01601v4), PDF §4.2 trang 7.

Toy: train TTS+WORLD, dev VC+WORLD. Nếu detector đọc dấu WORLD, dev chưa kiểm transfer sang neural vocoder mới. Dev vẫn có thể khó do acoustic model/content/channel khác; taxonomy chỉ chỉ ra component chưa được giữ ngoài train.

## 2. Năm corpus thay đổi nhiều trục

| Corpus/copy | Count P | Genuine source | Fake process/conditions | Scope |
|---|---:|---|---|---|
| ASV2019 LA eval | 71.237=7.355B+63.882S | VCTK | 13 TTS/VC IDs, 11 unseen algorithms | Partition khác train; không mọi component mới |
| DFADD HF copy | 3.755=755B+3.000S | VCTK theo PDF | Năm diffusion/flow TTS systems ×600 | Shared source; exact run-copy lineage chưa khóa |
| ADD2022 Track 1 | 109.199=31.334B+77.865S | Mandarin | Full fake, low quality/noise/music | Language/source/channel/generator cùng đổi |
| LibriSeVoc | 18.487 | LibriTTS | Sáu neural vocoder resyntheses | Không phải sáu full TTS pipelines độc lập |
| In-the-Wild | 31.779=19.963B+11.816S | Web recordings | Collected public-figure genuine/fake clips | Curated historical sample; generator lineage chưa đủ |

Operational B/S là nhãn protocol. Genuine truyền qua codec có thể vẫn B; vocoder resynthesis trong LibriSeVoc là S. Label không tự xác nhận consent hoặc factual truth. [LibriSeVoc §4.1](https://arxiv.org/pdf/2304.13085v2), [ADD §§2–4](https://arxiv.org/html/2202.08433v3), [ITW §3/§4.1.2](https://arxiv.org/pdf/2203.16263v5).

DFADD original có discrepancy về D3 VCTK/LibriTTS giữa các đoạn. [Author repo](https://github.com/isjwdu/DFADD) ghi sửa Matcha audio/label mismatch và formats tháng 04/2025. Nó khác benchmark copy `SpeechAntiSpoofingBenchmarks/DFADD` mà PDF liên kết. Chưa có manifest/checksum của executed run để khóa revision. ADD source nêu train/dev speaker split và test unseen utterances; chưa nâng thành test speaker-disjoint nếu nguồn không nói. [Dossier nguồn và unknowns](research/evidence-sources.md).

## 3. Ba gates bắt ba loại lỗi

PDF §4.3 trang 7–8 báo:

| Gate | Giữ/đổi gì | Evidence P | Có thể bắt; chưa kiểm |
|---|---|---|---|
| Recompute raw published scores | Giữ scores; dùng local labels/EER | ADD: 7/8 score sets khớp ≤0,01 pp | ID/label/polarity/metric; chưa đọc waveform |
| Official AASIST rerun | Audio/toolkit/released checkpoint | Năm corpora, Table 1 | Audio/protocol disagreements; chưa kiểm JEPA wrapper |
| Direct vs wrapper score | Cùng detector/waveform, hai paths | Tolerance 10⁻³ | Wrapper implementation; hai paths vẫn có thể cùng assumption sai |

Toy: ID-label join đúng nhưng decoder dùng rendition audio khác. Gate 1 vẫn pass; gate 2 có thể lệch. Wrapper JEPA quên subtract mean thì AASIST gate có thể pass, JEPA direct/wrapper gate có thể fail. Muốn đọc provenance phải biết mỗi gate đã chạm khối nào.

Table 1 AASIST ours/published: ASV 0,83/0,82; DFADD 41,87/41,86; ADD 47,92/47,91; Libri 37,95/37,65; ITW 43,02/43,00. Libri chênh 0,30 pp; chưa gọi exact bitwise replication. Đây là reported evidence, root chưa rerun detector.

## 4. Common evaluation chưa là common training

[Speech DF Arena v1 §3.1](https://arxiv.org/html/2509.02859v1#S3) đánh giá released weights với toolkit/metric/sample rate chung; input duration vẫn khác giữa systems. Nó chưa retrain mọi baseline bằng cùng head/augmentation/search budget. Baselines đa số dùng full ASV2019 train; ours bỏ A05+600B; Whisper-MesoNet dùng ASV2021 DF subset.

Table 3 là system-level context trên DFADD, chưa là encoder ablation hay universal SOTA. Original baseline papers xác lập identity/method; **Arena v1 Table 2** cung cấp các published DFADD scores. [Sổ baseline đầy đủ](research/evidence-sources.md).

**Phản ví dụ:** cùng decoder/metric nhưng hệ A thấy 4 s, hệ B thấy 2,56 s; evidence quan sát chưa bằng nhau. Old DFADD 39,05 tách khỏi submitted 41,87 ở [history](07-LICH-SU-THU-NGHIEM.md).

## 5. Tự kiểm

1. A05 holdout đã kiểm mọi vocoder chưa? Nêu component còn chung.
2. Gate 1 pass có bảo đảm resampling đúng không?
3. Shared VCTK có chứng minh exact train/test overlap không?
4. Vì sao AASIST gate chưa isolate benefit của JEPA objective?

<details>
<summary>Đáp án và reasoning</summary>

1. Chưa: WORLD còn ở A02/A03; A06 còn VC nhưng khác thuật toán.
2. Không: gate dùng scores có sẵn, không đọc audio.
3. Không: cần sample/speaker manifests, ancestry và duplicate audit. Source association khác leakage.
4. Hai complete systems còn khác architecture/input/learning/subset; chưa intervention riêng objective.

</details>

**Đào sâu:** viết exposure vector cho mỗi corpus: pretraining, downstream labels, development access, evaluation feedback và speaker/generator/channel overlap. Corpus có tên khác chưa có nghĩa unseen everything.
