# Sổ hộ chiếu dataset của lớp 2

[Bài7: cách đọc protocol](07-DATASET-VA-PROTOCOL.md) · [Sổ nguồn và mức đọc](09-SO-DANG-KY-NGUON.md) · [Bắt đầu](00-BAT-DAU-LOP-02.md)

Đối chiếu11/10/2026. Đây là **passport từ nguồn công bố**, chưa audit manifests/waveforms của các run trong paper. Bản paper, release và bản copy qua aggregator là ba identity cần khóa riêng. Không dùng mirror statistics thay số chính thức. “Chưa xác định” là unknown thật, không điền bằng giả định.

## P1 — ASVspoof2019 LA

| Trường | Nội dung |
|---|---|
| Identity/version | ASVspoof2019 **Logical Access**; database paper arXiv1911.01601v4,14/07/2020; official2019 plan |
| Task/unit/labels | Countermeasure trong ASV threat model; utterance bona fide/spoof. Target/non-target bona fide của ASV đều bona fide cho binary CM |
| Genuine ancestry | VCTK; đọc speech tiếng Anh; speaker partitions theo protocol |
| Fake pipeline | Train/dev A01–A06: bốn TTS, hai VC. Eval A07–A19. Table1 phân tách frontend/acoustic/duration/voice conversion/waveform stage; không mọi attack đều neural |
| Signal chain | LA post-sensor digital attacks; không phải loa–phòng–micro PA. Chưa audit preprocessing của exact copy chấm paper |
| Split/exposure | Train/dev/eval target-speaker partitions tách theo protocol. Eval attack IDs khác train/dev, nhưng A16/A19 tái dùng thuật toán A04/A06 |
| Generalization | In-corpus eval có11 unknown systems theo định nghĩa paper; không gọi cả13 là thuật toán mới hoặc family/vocoder-disjoint |
| Unknown/limits | Exact corpus checksums và toolkit manifest của run chưa kiểm. Generator identity cần granularity theo Table1, không chỉ A-ID |

Nguồn primary: [database §§2.1–3.1, Table1](https://arxiv.org/pdf/1911.01601), [official evaluation plan](https://www.asvspoof.org/asvspoof2019/asvspoof2019_evaluation_plan.pdf). Paper của mình train subset giữ A05+600 bona fide; xem bài7, không thay official split bằng subset này.

## P2 — ASVspoof2021 LA

| Trường | Nội dung |
|---|---|
| Identity/version | ASVspoof2021 **LA**; official challenge/release, overview arXiv2210.02437v3,22/06/2023 |
| Task/unit/labels | LA countermeasure hỗ trợ ASV; utterance bona fide/spoof; t-DCF là metric ASV-integrated |
| Genuine ancestry | VCTK-related2019LA speakers/source; thêm utterances và channel trong2021 protocol |
| Fake pipeline | Các attack systems2019LA được dùng trong evaluation2021, qua điều kiện truyền dẫn |
| Signal chain | Bảy channel/codec conditions, gồm telephone/VoIP/PSTN; genuine và spoof chịu channel |
| Split/exposure | Challenge dùng ASV2019LA train/dev; không có train/dev2021 mới trong challenge chính |
| Generalization | Đo detection dưới transmission conditions; không tự là generator-family holdout |
| Unknown/limits | Exact subset/phase/metadata của run nếu dùng phải khóa. Không parse trực tiếp PDF plan2021 trong lượt này; dựa official page/release và overview method |

Nguồn: [official2021](https://www.asvspoof.org/index2021.html), [LA release record4837263](https://zenodo.org/records/4837263), [overview §§II–III-A](https://arxiv.org/pdf/2210.02437). Paper hiện tại không báo evaluation2021LA.

## P3 — ASVspoof2021 DF

| Trường | Nội dung |
|---|---|
| Identity/version | ASVspoof2021 **DF**; cùng overview như P2 nhưng task/release khác |
| Task/unit/labels | Standalone speech deepfake detection, **không ASV**; clip bona fide/spoof; metric chính EER |
| Genuine ancestry |2019LA và các nguồn VCC2018/VCC2020; overview nêu DAPS/EMIME trong construction |
| Fake pipeline | TTS/VC từ nhiều systems; pooled set không tương đương một family held out |
| Signal chain | Chín lossy-compression conditions trong corpus; phải xem metadata từng condition |
| Split/exposure | Không release train/dev2021 mới; challenge dùng2019LA training/development và nêu hạn chế data use |
| Generalization | Corpus/source/generator/compression có thể cùng đổi; source bona fide mismatch cần đọc cùng detection |
| Unknown/limits | Không kiểm exact manifest/file mappings. “DF” chỉ protocol này, không tên chung cho LA/PA |

Nguồn: [official DF record4835108](https://zenodo.org/records/4835108), [overview §§II–III-C và analysis](https://arxiv.org/pdf/2210.02437). **PA counterexample:** bona fide recording qua replay là spoof trong physical-access protocol; không suy nhãn từ origin ban đầu. Paper hiện tại không báo evaluation2021DF/PA.

## P4 — ASVspoof5, tách Track1 và Track2

| Trường | Nội dung |
|---|---|
| Identity/version | Official Phase2 evaluation plan **v0.6,28/06/2024**; corpus ASVspoof5, Track và open/closed condition phải ghi |
| Task/unit/labels | Track1: utterance standalone bona fide/spoof. Track2: SASV enrollment–probe trials, target bona fide/non-target bona fide/spoof; chỉ target bona fide chấp nhận |
| Genuine ancestry | MLS English; Phase1 data contributors có thể dùng subset CommonVoice English11.0 để train speaker encoders theo plan §3 |
| Fake pipeline | Spoof systems/attack IDs theo split; exact checkpoint/component provenance chưa audit. Attack-disjoint không tự là family-disjoint |
| Signal chain | Evaluation channel/codec conditions; nhãn bona fide cũng có thể có lossy codec |
| Split/exposure | Speaker partitions và external-data restrictions. Không pool train/dev để train theo plan; dev cho fusion/calibration theo rule. Evaluation trials xử lý độc lập |
| Generalization | Attack/speaker/channel theo protocol. Open training condition không đồng nghĩa open-set attribution output |
| Unknown/limits | Nguồn ngoài phải kiểm cùng speaker/utterance; plan cho LibriSpeech vì speaker-disjoint theo thiết kế. Không suy leakage từ chung LibriVox ancestry. Không kiểm raw manifests/checkpoints |

Nguồn: [plan §§3–4.3, Tables2–3](https://www.asvspoof.org/file/ASVspoof5___Evaluation_Plan_Phase2.pdf), [official repo](https://github.com/asvspoof-challenge/asvspoof5). Subagent có đọc [data/method paper2502.08857v4](https://arxiv.org/html/2502.08857v4); root kiểm các claim lõi trên plan, chưa audit recipes từ paper đó. Paper hiện tại không evaluationASVspoof5/SASV.

<a id="dfadd"></a>

## P5 — DFADD

| Trường | Nội dung |
|---|---|
| Identity/version | DFADD paper2409.08731v1,13/09/2024; official repo ghi **04/2025 sửa Matcha audio–label mismatch và thống nhất format** |
| Task/unit/labels | Utterance TTS deepfake detection; genuine VCTK, spoof speech sinh. “Paired” theo speaker/design, không phải cùng text hoặc same recording |
| Genuine ancestry | VCTK109speakers theo paper. Fake texts:300 câu LJ Speech, tránh trùng prompt VCTK; source/content khác là factor cần ghi |
| Fake pipeline | D1 Grad-TTS, D2 NaturalSpeech2, D3 StyleTTS2; F1 Matcha-TTS, F2 PFlow-TTS. D1/F2 có HiFi-GAN VCTK replacement; D2/F2 có unofficial implementations |
| Signal chain | Paper chuẩn hóa16kHz; exact aggregator preprocessing chưa kiểm |
| Split/exposure | Dev p226/p229, test p227/p228, speakers còn lại train theo paper; các systems có trong split construction, không generator holdout tự động |
| Generalization | Speaker split trong corpus; transfer từ ASV2019 sang DFADD còn phụ thuộc source/decoder/components |
| Unknown/limits | §3.2 nói D2/D3/F1 pretrained VCTK, nhưng §3.2.3 riêng D3 nói StyleTTS2 LibriTTS checkpoint. **D3 ancestry chưa resolve**. Exact release/hash của bản copy paper chưa khóa |

Nguồn: [paper §§3.1–3.3,4.2](https://arxiv.org/html/2409.08731v1), [official repo Updates](https://github.com/isjwdu/DFADD), [author dataset](https://huggingface.co/datasets/isjwdu/DFADD). PDF của mình §4.2 dùng3.755 eval utterances qua [aggregator](https://huggingface.co/datasets/SpeechAntiSpoofingBenchmarks/DFADD), không phải toàn bộ quy mô corpus gốc; URL aggregator ở đây là provenance được PDF báo cáo, chưa audit files/release.

## P6 — LibriSeVoc

| Trường | Nội dung |
|---|---|
| Identity/version | AI-Synthesized Voice Detection Using Neural Vocoder Artifacts,2304.13085v2,27/04/2023; author repo |
| Task/unit/labels | Utterance bona fide/vocoder-resynthesized; nghiên cứu còn có vocoder identification. Task paper hiện tại chỉ binary |
| Genuine ancestry | LibriTTS: original audiobook/text materials liên quan LibriSpeech/LibriVox; same mel/source utterance cho các derivatives |
| Fake pipeline | Sáu vocoders: WaveNet, WaveRNN, WaveGrad, DiffWave, MelGAN, Parallel WaveGAN; self-vocoding chứ không sáu full TTS pipelines độc lập |
| Signal chain | Corpus gốc24kHz; bản copy/detector input có thể resample nên cần ghi riêng |
| Split/exposure | Paper §4.1 nêu non-overlap6:2:2. Speaker-disjoint và grouping mọi derivatives theo source trước khi chia **chưa xác định** từ mô tả đã đọc |
| Generalization | Standard split không tự vocoder-unseen; các vocoders có trong corpus. Cross-source results không isolate vocoder cause |
| Unknown/limits | Manifest/release/checkpoints và pair mapping chưa audit. Shared ancestry không chứng minh exact overlap với train/encoder hoặc ASVspoof5 |

Nguồn: [paper §4.1, Table1](https://arxiv.org/pdf/2304.13085), [author repo](https://github.com/csun22/Synthetic-Voice-Detection-Vocoder-Artifacts), [LibriTTS2019 §§1,3](https://arxiv.org/pdf/1904.02882).

## P7 — In-the-Wild

| Trường | Nội dung |
|---|---|
| Identity/version | Does Audio Deepfake Detection Generalize?,2203.16263; root đọc PDF **v5,27/03/2026**, paper xuất hiện đầu2022. Version paper khác ngày thu corpus |
| Task/unit/labels | Clip English genuine/fake thu từ web; thiết kế dùng external evaluation |
| Genuine ancestry | Public recordings cùng58celebrity/politician identities; matching tương đối style/background/duration với fake |
| Fake pipeline | Publicly advertised deepfake demos/video/audio; generator/model/checkpoint từng file **chưa xác định đầy đủ** |
| Signal chain | Segmentation từ web media, chuyển WAV/downsample16kHz; original source codecs/editing còn có thể để traces |
| Split/exposure | Paper dùng như cross-database eval; không áp một train/dev/test speaker/model-disjoint split chưa được nguồn xác lập |
| Generalization | Transfer tới collected web corpus. Cùng speaker matching không loại hết source/channel confound |
| Unknown/limits | Exact release hashes và generator lineage chưa audit; không gọi temporal holdout hay đại diện mọi generator thương mại2026 |

Nguồn: [paper §3 và4.1.2](https://arxiv.org/pdf/2203.16263), [official dataset page](https://deepfake-total.com/in_the_wild), [author dataset](https://huggingface.co/datasets/mueller91/In-The-Wild). Cần phân biệt publication revision với corpus refresh.

## P8 — ADD2022 Track1 / LF

| Trường | Nội dung |
|---|---|
| Identity/version | ADD2022 **Track1 low-quality full fake**; paper2202.08433v3,02/07/2024; không dùng revision date làm năm challenge |
| Task/unit/labels | Utterance genuine/full fake detection, metric EER; Track2 partial fake và Track3 fake game là task khác |
| Genuine ancestry | Train/dev AISHELL-3 Mandarin; challenge có các nguồn AISHELL khác, không mặc định mọi eval genuine chỉ AISHELL-3 |
| Fake pipeline | TTS/VC systems theo mô tả; exact test generator/component/checkpoint IDs chưa xác định từ phần đọc |
| Signal chain | Noise/background music và quality thấp trong Track1; không cô lập riêng language shift |
| Split/exposure | Train/dev speaker-disjoint được §3.1 xác nhận; có adaptation set. Test gọi unseen utterances; speaker overlap test/adaptation với splits khác **chưa xác định** |
| Generalization | Mandarin+source+speaker/noise/generator có thể cùng đổi; không đủ gọi pure unseen-family hay pure cross-language |
| Unknown/limits | Exact release/manifest/copy trong run paper chưa audit; không dùng adaptation để claim zero target exposure nếu đã truy cập nó |

Nguồn: [challenge paper §§2–4](https://arxiv.org/pdf/2202.08433), [Track1 eval record10843991](https://zenodo.org/records/10843991). Subagent đọc release metadata; root kiểm train/dev split và task trên full text, không tải dataset.

## P9 — PartialSpoof: bản2021 và extended annotations

| Trường | Nội dung |
|---|---|
| Identity/version | Initial Interspeech2021 paper; extended TASLP paper2204.05177v3,30/01/2023; [release v1.2 record5766198](https://zenodo.org/records/5766198) theo subagent |
| Task/unit/labels | Utterance detection và temporal segment labels. Extended version có20/40/80/160/320/640ms; không chuyển mọi con số đó sang bản initial |
| Genuine ancestry | ASVspoof2019LA bona fide; spoof/genuine donor segments từ nguồn benchmark |
| Fake pipeline | VAD candidates, chọn/replacement từ class khác cùng speaker theo extended §III-B, cross-correlation+overlap-add; không phải semantic text editing dataset |
| Signal chain | Amplitude normalization và alignment/crossfade theo recipe; biên không giống mọi partial edit ngoài đời |
| Split/exposure | Kế thừa train/dev/eval nền2019LA; cần derivative/donor mapping để audit thêm disjointness ở cấp source |
| Generalization | Partial fraction, segment duration/resolution và attack pipeline; clip score không chứng minh localization |
| Unknown/limits | Exact annotation release/manifest chưa kiểm. Labels ghi nguồn generated frames; §III-D nêu replacement không theo nghĩa câu/phần âm vị |

Nguồn: [initial2021 §§2–3](https://www.isca-archive.org/interspeech_2021/zhang21ca_interspeech.pdf), [extended §§III-B–III-D,V-C](https://arxiv.org/html/2204.05177v3), [author repo](https://github.com/nii-yamagishilab/PartialSpoof). Paper hiện tại không evaluation/localization trên corpus này.

## P10 — CodecFake **Wu et al.**, không bỏ author/paper ID

| Trường | Nội dung |
|---|---|
| Identity/version | CodecFake2406.07237v1,11/06/2024, Interspeech2024; Wu et al.; author project/codecfake.github.io và rogertseng/CodecFake |
| Task/unit/labels | Utterance genuine/codec resynthesis cho detector; không mọi fake là full audio-LM generation |
| Genuine ancestry | VCTK107speakers theo corpus design; từng codec subset có corresponding source genuine |
| Fake pipeline | Encoder–quantizer–decoder của15pretrained models từ6codec frameworks; codec pretrain sources theo Table1, không đồng nhất |
| Signal chain | Reconstruction/config của từng codec; bitrate/checkpoint khác nhau; preprocessing cụ thể cần khóa |
| Split/exposure | Speaker split train103,dev2,test2; dev p226/p229,test p227/p228. Same codec collection qua standard split không tự codec-heldout |
| Generalization | Study thử transfer codec-trained detection sang codec-TTS; VALL-E evaluation dùng open-source reimplementation, không original Microsoft checkpoint |
| Unknown/limits | Chưa audit all files/config/checksums. Không suy universal ALM fingerprints hoặc codec = spoof theo mọi ứng dụng |

Nguồn: [paper §§2–3, Tables1–2](https://arxiv.org/pdf/2406.07237), [author project](https://codecfake.github.io/), [author dataset](https://huggingface.co/datasets/rogertseng/CodecFake).

**Disambiguation:** Lu et al. [2406.08112v1](https://arxiv.org/html/2406.08112v1) và Xie et al. [2405.04880v3](https://arxiv.org/html/2405.04880v3), [repo](https://github.com/xieyuankun/Codecfake), cũng dùng tên Codecfake. Subagent đọc hai full texts, thấy overlapping authors/methods/counts; quan hệ manifest/release chính xác giữa hai bài chưa kiểm. Root chỉ kiểm identity ở mức trang/repo, **không dùng hai bài đó để bổ sung stats/protocol vào passport Wu hoặc đếm thành hai corpus độc lập**. Đây là giới hạn xác minh được giữ rõ.

## Mẫu passport để tự lập

Sao chép tám trường P1, thêm URL/version/section/người đọc. Dưới mỗi claim disjointness viết evidence tương ứng: metadata protocol hay kết quả file audit. Với exact corpus used in a run, thêm toolkit commit, manifest/hash và mọi filtering/crop/resample. Mẫu này là tài liệu học; lượt này không truy cập corpus/checkpoints hoặc tái chạy evaluation.
