ajaudio / studyAUDIO-JEPA · RESEARCH NOTES
5 phút đọc · Toàn văn
Mục lục bài · 6 mục

Bài 4 — Dữ liệu, protocol và baseline đang kiểm điều gì?#

Trước · Tiếp.

Mục tiêu và tiền đề#

Đọc corpus theo label/process/source/exposure; phân biệt metric gate, audio gate và model-wrapper gate. Cần hộ chiếu lớp 2, cue/confound, split/access lớp 5, provenance. Counts là P, PDF §4.2/Table 1 trang 7; original/copy sources ở dossier.

1. A05 holdout giữ ngoài train cái gì?#

ASV2019 LA train có 20 speakers và attacks A01–A06. Paper giữ A05+600B để monitor, còn năm attack IDs trong detector training. A05 là VC dùng VAE và WORLD; A02/A03 cũng dùng WORLD. A06 còn thuộc VC nhưng dùng GMM-UBM/LPCC/MFCC và spectral filtering/overlap-add, khác pipeline A05. Đây là attack-ID holdout, chưa phải family/vocoder/speaker holdout. Genuine holdout cấp utterance chưa tự tạo speaker-disjoint dev.

Evaluation có 13 IDs A07–A19 nhưng A16/A19 tái dùng thuật toán A04/A06: 11 unseen algorithms khác 13 IDs mới. ASVspoof2019 v4, Table 1 và §3.1, PDF §4.2 trang 7.

Toy: train TTS+WORLD, dev VC+WORLD. Nếu detector đọc dấu WORLD, dev chưa kiểm transfer sang neural vocoder mới. Dev vẫn có thể khó do acoustic model/content/channel khác; taxonomy chỉ chỉ ra component chưa được giữ ngoài train.

2. Năm corpus thay đổi nhiều trục#

Corpus/copyCount PGenuine sourceFake process/conditionsScope
ASV2019 LA eval71.237=7.355B+63.882SVCTK13 TTS/VC IDs, 11 unseen algorithmsPartition khác train; không mọi component mới
DFADD HF copy3.755=755B+3.000SVCTK theo PDFNăm diffusion/flow TTS systems ×600Shared source; exact run-copy lineage chưa khóa
ADD2022 Track 1109.199=31.334B+77.865SMandarinFull fake, low quality/noise/musicLanguage/source/channel/generator cùng đổi
LibriSeVoc18.487LibriTTSSáu neural vocoder resynthesesKhông phải sáu full TTS pipelines độc lập
In-the-Wild31.779=19.963B+11.816SWeb recordingsCollected public-figure genuine/fake clipsCurated historical sample; generator lineage chưa đủ

Operational B/S là nhãn protocol. Genuine truyền qua codec có thể vẫn B; vocoder resynthesis trong LibriSeVoc là S. Label không tự xác nhận consent hoặc factual truth. LibriSeVoc §4.1, ADD §§2–4, ITW §3/§4.1.2.

DFADD original có discrepancy về D3 VCTK/LibriTTS giữa các đoạn. Author repo ghi sửa Matcha audio/label mismatch và formats tháng 04/2025. Nó khác benchmark copy SpeechAntiSpoofingBenchmarks/DFADD mà PDF liên kết. Chưa có manifest/checksum của executed run để khóa revision. ADD source nêu train/dev speaker split và test unseen utterances; chưa nâng thành test speaker-disjoint nếu nguồn không nói. Dossier nguồn và unknowns.

3. Ba gates bắt ba loại lỗi#

PDF §4.3 trang 7–8 báo:

GateGiữ/đổi gìEvidence PCó thể bắt; chưa kiểm
Recompute raw published scoresGiữ scores; dùng local labels/EERADD: 7/8 score sets khớp ≤0,01 ppID/label/polarity/metric; chưa đọc waveform
Official AASIST rerunAudio/toolkit/released checkpointNăm corpora, Table 1Audio/protocol disagreements; chưa kiểm JEPA wrapper
Direct vs wrapper scoreCùng detector/waveform, hai pathsTolerance 10⁻³Wrapper implementation; hai paths vẫn có thể cùng assumption sai

Toy: ID-label join đúng nhưng decoder dùng rendition audio khác. Gate 1 vẫn pass; gate 2 có thể lệch. Wrapper JEPA quên subtract mean thì AASIST gate có thể pass, JEPA direct/wrapper gate có thể fail. Muốn đọc provenance phải biết mỗi gate đã chạm khối nào.

Table 1 AASIST ours/published: ASV 0,83/0,82; DFADD 41,87/41,86; ADD 47,92/47,91; Libri 37,95/37,65; ITW 43,02/43,00. Libri chênh 0,30 pp; chưa gọi exact bitwise replication. Đây là reported evidence, root chưa rerun detector.

4. Common evaluation chưa là common training#

Speech DF Arena v1 §3.1 đánh giá released weights với toolkit/metric/sample rate chung; input duration vẫn khác giữa systems. Nó chưa retrain mọi baseline bằng cùng head/augmentation/search budget. Baselines đa số dùng full ASV2019 train; ours bỏ A05+600B; Whisper-MesoNet dùng ASV2021 DF subset.

Table 3 là system-level context trên DFADD, chưa là encoder ablation hay universal SOTA. Original baseline papers xác lập identity/method; Arena v1 Table 2 cung cấp các published DFADD scores. Sổ baseline đầy đủ.

Phản ví dụ: cùng decoder/metric nhưng hệ A thấy 4 s, hệ B thấy 2,56 s; evidence quan sát chưa bằng nhau. Old DFADD 39,05 tách khỏi submitted 41,87 ở history.

5. Tự kiểm#

  1. A05 holdout đã kiểm mọi vocoder chưa? Nêu component còn chung.
  2. Gate 1 pass có bảo đảm resampling đúng không?
  3. Shared VCTK có chứng minh exact train/test overlap không?
  4. Vì sao AASIST gate chưa isolate benefit của JEPA objective?
Đáp án và reasoning
  1. Chưa: WORLD còn ở A02/A03; A06 còn VC nhưng khác thuật toán.
  2. Không: gate dùng scores có sẵn, không đọc audio.
  3. Không: cần sample/speaker manifests, ancestry và duplicate audit. Source association khác leakage.
  4. Hai complete systems còn khác architecture/input/learning/subset; chưa intervention riêng objective.

Đào sâu: viết exposure vector cho mỗi corpus: pretraining, downstream labels, development access, evaluation feedback và speaker/generator/channel overlap. Corpus có tên khác chưa có nghĩa unseen everything.

DỪNG LẠI & TỰ KIỂM TRA

Bạn đã giải thích được cơ chế trong bài?

↓ Bản Markdown nguyên gốcGiữ nguyên nội dung · Công thức, bảng và nguồn đầy đủ.Các chat bàn giao được mở trong Codex.

Gõ từ khóa để tìm bài học và đoạn liên quan.