ajaudio / studyAUDIO-JEPA · RESEARCH NOTES
10 phút đọc · Toàn văn
Mục lục bài · 11 mục

Sổ hộ chiếu dataset của lớp 2#

Bài7: cách đọc protocol · Sổ nguồn và mức đọc · Bắt đầu

Đối chiếu11/10/2026. Đây là passport từ nguồn công bố, chưa audit manifests/waveforms của các run trong paper. Bản paper, release và bản copy qua aggregator là ba identity cần khóa riêng. Không dùng mirror statistics thay số chính thức. “Chưa xác định” là unknown thật, không điền bằng giả định.

P1 — ASVspoof2019 LA#

TrườngNội dung
Identity/versionASVspoof2019 Logical Access; database paper arXiv1911.01601v4,14/07/2020; official2019 plan
Task/unit/labelsCountermeasure trong ASV threat model; utterance bona fide/spoof. Target/non-target bona fide của ASV đều bona fide cho binary CM
Genuine ancestryVCTK; đọc speech tiếng Anh; speaker partitions theo protocol
Fake pipelineTrain/dev A01–A06: bốn TTS, hai VC. Eval A07–A19. Table1 phân tách frontend/acoustic/duration/voice conversion/waveform stage; không mọi attack đều neural
Signal chainLA post-sensor digital attacks; không phải loa–phòng–micro PA. Chưa audit preprocessing của exact copy chấm paper
Split/exposureTrain/dev/eval target-speaker partitions tách theo protocol. Eval attack IDs khác train/dev, nhưng A16/A19 tái dùng thuật toán A04/A06
GeneralizationIn-corpus eval có11 unknown systems theo định nghĩa paper; không gọi cả13 là thuật toán mới hoặc family/vocoder-disjoint
Unknown/limitsExact corpus checksums và toolkit manifest của run chưa kiểm. Generator identity cần granularity theo Table1, không chỉ A-ID

Nguồn primary: database §§2.1–3.1, Table1, official evaluation plan. Paper của mình train subset giữ A05+600 bona fide; xem bài7, không thay official split bằng subset này.

P2 — ASVspoof2021 LA#

TrườngNội dung
Identity/versionASVspoof2021 LA; official challenge/release, overview arXiv2210.02437v3,22/06/2023
Task/unit/labelsLA countermeasure hỗ trợ ASV; utterance bona fide/spoof; t-DCF là metric ASV-integrated
Genuine ancestryVCTK-related2019LA speakers/source; thêm utterances và channel trong2021 protocol
Fake pipelineCác attack systems2019LA được dùng trong evaluation2021, qua điều kiện truyền dẫn
Signal chainBảy channel/codec conditions, gồm telephone/VoIP/PSTN; genuine và spoof chịu channel
Split/exposureChallenge dùng ASV2019LA train/dev; không có train/dev2021 mới trong challenge chính
GeneralizationĐo detection dưới transmission conditions; không tự là generator-family holdout
Unknown/limitsExact subset/phase/metadata của run nếu dùng phải khóa. Không parse trực tiếp PDF plan2021 trong lượt này; dựa official page/release và overview method

Nguồn: official2021, LA release record4837263, overview §§II–III-A. Paper hiện tại không báo evaluation2021LA.

P3 — ASVspoof2021 DF#

TrườngNội dung
Identity/versionASVspoof2021 DF; cùng overview như P2 nhưng task/release khác
Task/unit/labelsStandalone speech deepfake detection, không ASV; clip bona fide/spoof; metric chính EER
Genuine ancestry2019LA và các nguồn VCC2018/VCC2020; overview nêu DAPS/EMIME trong construction
Fake pipelineTTS/VC từ nhiều systems; pooled set không tương đương một family held out
Signal chainChín lossy-compression conditions trong corpus; phải xem metadata từng condition
Split/exposureKhông release train/dev2021 mới; challenge dùng2019LA training/development và nêu hạn chế data use
GeneralizationCorpus/source/generator/compression có thể cùng đổi; source bona fide mismatch cần đọc cùng detection
Unknown/limitsKhông kiểm exact manifest/file mappings. “DF” chỉ protocol này, không tên chung cho LA/PA

Nguồn: official DF record4835108, overview §§II–III-C và analysis. PA counterexample: bona fide recording qua replay là spoof trong physical-access protocol; không suy nhãn từ origin ban đầu. Paper hiện tại không báo evaluation2021DF/PA.

P4 — ASVspoof5, tách Track1 và Track2#

TrườngNội dung
Identity/versionOfficial Phase2 evaluation plan v0.6,28/06/2024; corpus ASVspoof5, Track và open/closed condition phải ghi
Task/unit/labelsTrack1: utterance standalone bona fide/spoof. Track2: SASV enrollment–probe trials, target bona fide/non-target bona fide/spoof; chỉ target bona fide chấp nhận
Genuine ancestryMLS English; Phase1 data contributors có thể dùng subset CommonVoice English11.0 để train speaker encoders theo plan §3
Fake pipelineSpoof systems/attack IDs theo split; exact checkpoint/component provenance chưa audit. Attack-disjoint không tự là family-disjoint
Signal chainEvaluation channel/codec conditions; nhãn bona fide cũng có thể có lossy codec
Split/exposureSpeaker partitions và external-data restrictions. Không pool train/dev để train theo plan; dev cho fusion/calibration theo rule. Evaluation trials xử lý độc lập
GeneralizationAttack/speaker/channel theo protocol. Open training condition không đồng nghĩa open-set attribution output
Unknown/limitsNguồn ngoài phải kiểm cùng speaker/utterance; plan cho LibriSpeech vì speaker-disjoint theo thiết kế. Không suy leakage từ chung LibriVox ancestry. Không kiểm raw manifests/checkpoints

Nguồn: plan §§3–4.3, Tables2–3, official repo. Subagent có đọc data/method paper2502.08857v4; root kiểm các claim lõi trên plan, chưa audit recipes từ paper đó. Paper hiện tại không evaluationASVspoof5/SASV.

P5 — DFADD#

TrườngNội dung
Identity/versionDFADD paper2409.08731v1,13/09/2024; official repo ghi 04/2025 sửa Matcha audio–label mismatch và thống nhất format
Task/unit/labelsUtterance TTS deepfake detection; genuine VCTK, spoof speech sinh. “Paired” theo speaker/design, không phải cùng text hoặc same recording
Genuine ancestryVCTK109speakers theo paper. Fake texts:300 câu LJ Speech, tránh trùng prompt VCTK; source/content khác là factor cần ghi
Fake pipelineD1 Grad-TTS, D2 NaturalSpeech2, D3 StyleTTS2; F1 Matcha-TTS, F2 PFlow-TTS. D1/F2 có HiFi-GAN VCTK replacement; D2/F2 có unofficial implementations
Signal chainPaper chuẩn hóa16kHz; exact aggregator preprocessing chưa kiểm
Split/exposureDev p226/p229, test p227/p228, speakers còn lại train theo paper; các systems có trong split construction, không generator holdout tự động
GeneralizationSpeaker split trong corpus; transfer từ ASV2019 sang DFADD còn phụ thuộc source/decoder/components
Unknown/limits§3.2 nói D2/D3/F1 pretrained VCTK, nhưng §3.2.3 riêng D3 nói StyleTTS2 LibriTTS checkpoint. D3 ancestry chưa resolve. Exact release/hash của bản copy paper chưa khóa

Nguồn: paper §§3.1–3.3,4.2, official repo Updates, author dataset. PDF của mình §4.2 dùng3.755 eval utterances qua aggregator, không phải toàn bộ quy mô corpus gốc; URL aggregator ở đây là provenance được PDF báo cáo, chưa audit files/release.

P6 — LibriSeVoc#

TrườngNội dung
Identity/versionAI-Synthesized Voice Detection Using Neural Vocoder Artifacts,2304.13085v2,27/04/2023; author repo
Task/unit/labelsUtterance bona fide/vocoder-resynthesized; nghiên cứu còn có vocoder identification. Task paper hiện tại chỉ binary
Genuine ancestryLibriTTS: original audiobook/text materials liên quan LibriSpeech/LibriVox; same mel/source utterance cho các derivatives
Fake pipelineSáu vocoders: WaveNet, WaveRNN, WaveGrad, DiffWave, MelGAN, Parallel WaveGAN; self-vocoding chứ không sáu full TTS pipelines độc lập
Signal chainCorpus gốc24kHz; bản copy/detector input có thể resample nên cần ghi riêng
Split/exposurePaper §4.1 nêu non-overlap6:2:2. Speaker-disjoint và grouping mọi derivatives theo source trước khi chia chưa xác định từ mô tả đã đọc
GeneralizationStandard split không tự vocoder-unseen; các vocoders có trong corpus. Cross-source results không isolate vocoder cause
Unknown/limitsManifest/release/checkpoints và pair mapping chưa audit. Shared ancestry không chứng minh exact overlap với train/encoder hoặc ASVspoof5

Nguồn: paper §4.1, Table1, author repo, LibriTTS2019 §§1,3.

P7 — In-the-Wild#

TrườngNội dung
Identity/versionDoes Audio Deepfake Detection Generalize?,2203.16263; root đọc PDF v5,27/03/2026, paper xuất hiện đầu2022. Version paper khác ngày thu corpus
Task/unit/labelsClip English genuine/fake thu từ web; thiết kế dùng external evaluation
Genuine ancestryPublic recordings cùng58celebrity/politician identities; matching tương đối style/background/duration với fake
Fake pipelinePublicly advertised deepfake demos/video/audio; generator/model/checkpoint từng file chưa xác định đầy đủ
Signal chainSegmentation từ web media, chuyển WAV/downsample16kHz; original source codecs/editing còn có thể để traces
Split/exposurePaper dùng như cross-database eval; không áp một train/dev/test speaker/model-disjoint split chưa được nguồn xác lập
GeneralizationTransfer tới collected web corpus. Cùng speaker matching không loại hết source/channel confound
Unknown/limitsExact release hashes và generator lineage chưa audit; không gọi temporal holdout hay đại diện mọi generator thương mại2026

Nguồn: paper §3 và4.1.2, official dataset page, author dataset. Cần phân biệt publication revision với corpus refresh.

P8 — ADD2022 Track1 / LF#

TrườngNội dung
Identity/versionADD2022 Track1 low-quality full fake; paper2202.08433v3,02/07/2024; không dùng revision date làm năm challenge
Task/unit/labelsUtterance genuine/full fake detection, metric EER; Track2 partial fake và Track3 fake game là task khác
Genuine ancestryTrain/dev AISHELL-3 Mandarin; challenge có các nguồn AISHELL khác, không mặc định mọi eval genuine chỉ AISHELL-3
Fake pipelineTTS/VC systems theo mô tả; exact test generator/component/checkpoint IDs chưa xác định từ phần đọc
Signal chainNoise/background music và quality thấp trong Track1; không cô lập riêng language shift
Split/exposureTrain/dev speaker-disjoint được §3.1 xác nhận; có adaptation set. Test gọi unseen utterances; speaker overlap test/adaptation với splits khác chưa xác định
GeneralizationMandarin+source+speaker/noise/generator có thể cùng đổi; không đủ gọi pure unseen-family hay pure cross-language
Unknown/limitsExact release/manifest/copy trong run paper chưa audit; không dùng adaptation để claim zero target exposure nếu đã truy cập nó

Nguồn: challenge paper §§2–4, Track1 eval record10843991. Subagent đọc release metadata; root kiểm train/dev split và task trên full text, không tải dataset.

P9 — PartialSpoof: bản2021 và extended annotations#

TrườngNội dung
Identity/versionInitial Interspeech2021 paper; extended TASLP paper2204.05177v3,30/01/2023; release v1.2 record5766198 theo subagent
Task/unit/labelsUtterance detection và temporal segment labels. Extended version có20/40/80/160/320/640ms; không chuyển mọi con số đó sang bản initial
Genuine ancestryASVspoof2019LA bona fide; spoof/genuine donor segments từ nguồn benchmark
Fake pipelineVAD candidates, chọn/replacement từ class khác cùng speaker theo extended §III-B, cross-correlation+overlap-add; không phải semantic text editing dataset
Signal chainAmplitude normalization và alignment/crossfade theo recipe; biên không giống mọi partial edit ngoài đời
Split/exposureKế thừa train/dev/eval nền2019LA; cần derivative/donor mapping để audit thêm disjointness ở cấp source
GeneralizationPartial fraction, segment duration/resolution và attack pipeline; clip score không chứng minh localization
Unknown/limitsExact annotation release/manifest chưa kiểm. Labels ghi nguồn generated frames; §III-D nêu replacement không theo nghĩa câu/phần âm vị

Nguồn: initial2021 §§2–3, extended §§III-B–III-D,V-C, author repo. Paper hiện tại không evaluation/localization trên corpus này.

P10 — CodecFake Wu et al., không bỏ author/paper ID#

TrườngNội dung
Identity/versionCodecFake2406.07237v1,11/06/2024, Interspeech2024; Wu et al.; author project/codecfake.github.io và rogertseng/CodecFake
Task/unit/labelsUtterance genuine/codec resynthesis cho detector; không mọi fake là full audio-LM generation
Genuine ancestryVCTK107speakers theo corpus design; từng codec subset có corresponding source genuine
Fake pipelineEncoder–quantizer–decoder của15pretrained models từ6codec frameworks; codec pretrain sources theo Table1, không đồng nhất
Signal chainReconstruction/config của từng codec; bitrate/checkpoint khác nhau; preprocessing cụ thể cần khóa
Split/exposureSpeaker split train103,dev2,test2; dev p226/p229,test p227/p228. Same codec collection qua standard split không tự codec-heldout
GeneralizationStudy thử transfer codec-trained detection sang codec-TTS; VALL-E evaluation dùng open-source reimplementation, không original Microsoft checkpoint
Unknown/limitsChưa audit all files/config/checksums. Không suy universal ALM fingerprints hoặc codec = spoof theo mọi ứng dụng

Nguồn: paper §§2–3, Tables1–2, author project, author dataset.

Disambiguation: Lu et al. 2406.08112v1 và Xie et al. 2405.04880v3, repo, cũng dùng tên Codecfake. Subagent đọc hai full texts, thấy overlapping authors/methods/counts; quan hệ manifest/release chính xác giữa hai bài chưa kiểm. Root chỉ kiểm identity ở mức trang/repo, không dùng hai bài đó để bổ sung stats/protocol vào passport Wu hoặc đếm thành hai corpus độc lập. Đây là giới hạn xác minh được giữ rõ.

Mẫu passport để tự lập#

Sao chép tám trường P1, thêm URL/version/section/người đọc. Dưới mỗi claim disjointness viết evidence tương ứng: metadata protocol hay kết quả file audit. Với exact corpus used in a run, thêm toolkit commit, manifest/hash và mọi filtering/crop/resample. Mẫu này là tài liệu học; lượt này không truy cập corpus/checkpoints hoặc tái chạy evaluation.

↓ Bản Markdown nguyên gốcGiữ nguyên nội dung · Công thức, bảng và nguồn đầy đủ.Các chat bàn giao được mở trong Codex.

Gõ từ khóa để tìm bài học và đoạn liên quan.