ajaudio / studyAUDIO-JEPA · RESEARCH NOTES
17 phút đọc · Toàn văn
Mục lục bài · 5 mục

Sổ nguồn và mức xác minh — lớp 6#

Đối chiếu ngày 11/10/2026. Tài liệu ghi nguồn primary, phạm vi đã đọc, và ranh giới giữa kết quả bản PDF nộp báo cáo với lịch sử notebook/handoff. Nó không xác nhận lại một lần chạy detector, không khóa checksum dataset trong run, và không xác nhận trạng thái chấp nhận paper.

Mức đọc: M = đọc full text các mục/bảng liên quan; O = đọc trang/repo/plan/release chính thức; A = chỉ metadata/abstract hoặc phần được trích dẫn; C = đọc code, không chứng minh code đã chạy; H = handoff/history nội bộ ghi kết quả; P = con số được PDF nộp báo cáo. Các mức này không đồng nghĩa audit toàn bộ paper, checkpoint hay waveform.

1. Bản PDF nộp và các số gốc#

Nguồn là Audio-JEPA Representations for Speech Deepfake Detection, Nguyen Le Nguyen và Tran Van Hoai, bản PDF 14 trang. Đã trích text, đọc protocol/results/discussion/limitations liên quan, và kiểm bằng render trang 7–10: trang 7, trang 8, trang 9, trang 10. Các số dưới đây là reported submitted paper (P), không phải kết quả tái chạy ở lượt này.

Table 1, p.7 báo AASIST official-checkpoint re-run / published EER lần lượt: ASVspoof2019 LA 0.83/0.82; DFADD 41.87/41.86; ADD22 Track 1 47.92/47.91; LibriSeVoc 37.95/37.65; In-the-Wild 43.02/43.00. Section 4.3 p.8 báo score-file checks trên ADD22 khớp Table 2 Arena trong 0.01 điểm cho 7/8 score sets đã kiểm và wrapper/direct score khác không quá 10⁻³; đây vẫn là mô tả của paper, chưa được tái tạo trong lượt này.

Table 2, p.9, EER %; paper nói mean ± sample SD trên ba seed, trừ fine-tuned In-the-Wild có dấu dagger n=1; dấu “—” là chưa chạy:

CorpusJEPA frozenJEPA fine-tunedAASIST re-run
DFADD8.47 ± 0.3510.75 ± 3.4841.87
ASVspoof2019 LA eval17.70 ± 1.447.39 ± 0.690.83
ADD22 Track 1—44.10 ± 2.0447.92
LibriSeVoc47.38 ± 4.1941.96 ± 2.9337.95
In-the-Wild47.55 ± 8.5552.00†43.02

Theo Table 1 p.7 và §4.2 pp.7–8, eval size/provenance được paper báo cáo là: ASVspoof19 LA 7,355 genuine + 63,882 spoof (71,237 tổng; VCTK; 13 eval attacks, trong đó 11 unseen algorithms); DFADD 755+3,000 (3,755; VCTK genuine, 5 TTS systems ×600 spoof); ADD22 Track 1 31,334+77,865 (109,199; Mandarin/noisy); LibriSeVoc 18,487 (LibriTTS source, six vocoders); In-the-Wild 19,963+11,816 (31,779; public web sources). Đây là con số reported trong PDF, không phải audit manifest/waveforms.

Mean ± SD không phải confidence interval. Table 4 p.10 so released Audio-JEPA/AudioMAE checkpoints bằng cùng downstream recipe ở bốn ô: JEPA có mean EER thấp hơn 5.87 pp (ASV fine-tune), 7.55 pp (ASV frozen), 2.57 pp (DFADD frozen); AudioMAE thấp hơn 2.08 pp ở ADD22 fine-tune. Paper đưa ranges qua ba seed để mô tả; §5.2 nói rõ overlap/non-overlap của ranges không phải test significance. Seed đầu ADD22 đảo chiều so với mean. Không suy per-seed values còn thiếu, CI, hoặc equivalence từ SD/ranges.

Section 4.2 p.7–8 báo tập train chính là ASVspoof2019 LA 25,380 files; giữ A05 và 600 bona fide làm holdout giám sát nhưng không dùng chọn checkpoint, train còn 20,980. A05 là VC dùng WORLD; chi tiết thuật toán dưới đây lấy từ source ASVspoof primary. Section 4.4 p.8 nói bảng baseline là so sánh complete systems: training subset của chúng tôi khác full ASVspoof LA vì loại A05 + 600 genuine; kiến trúc, pretraining, recipe cũng khác. Whisper-MesoNet dùng training corpus khác theo paper Arena. Section 6 coi source pattern là association/hypothesis, không causal decomposition. Section 8 nêu so sánh checkpoint không isolate objective.

Evaluation corpus và protocol source#

Nguồn primary / version / mức đọcPhần đã đọc và claim có thể hỗ trợUnknown hoặc ranh giới
Dowerah et al., Speech DF Arena: A Leaderboard for Speech DeepFake Detection Models, arXiv:2509.02859v1, 02/09/2025, M§§2.1–2.2, 3.1–3.2, 4.1/Table 2. Source cố định cho 11 EER DFADD trong Table 3 của PDF. §3.1: dùng released weights; tất cả model train ASVspoof2019, trừ Whisper-MesoNet (subset ASVspoof2021 DF); input 16 kHz, đa số fixed 4 s, WavLM/HuBERT/Wav2Vec2 pad theo utterance dài nhất batch.EER Table 2 là scores của hệ phát hành trong Arena, không phải retraining đồng recipe; bảng không báo seed/n hoặc uncertainty. Leaderboard HF thay đổi theo thời gian, chỉ dùng làm live context, không thay số versioned Table 2.
Wang et al., ASVspoof 2019: A large-scale public database of synthesized, converted and replayed speech, arXiv v4, 14/07/2020, M; 2019 evaluation plan, O§§2.1–2.3, §3.1/Table 1, cụ thể A02–A06/A16/A19. Hỗ trợ quy mô LA, VCTK provenance theo database paper, chia speaker, cấu tạo A05, và algorithm reuse ở evaluation.Holdout một attack ID không đồng nghĩa speaker-, vocoder- hay attack-family-disjoint. Source paper/protocol không khóa exact files/checkpoint/manifests trong run của chúng tôi.
Du et al., DFADD: The Diffusion and Flow-Matching based Audio Deepfake Dataset, arXiv v1, 13/09/2024, M; author repo, O§§3.1–3.3, 4.1–4.2, Tables 1–2. Original design: 109 VCTK speakers, five TTS systems, test-side 600 spoof per system; source paper describes speaker split and sample-rate standardization. Author repository explicitly says [04/2025] fixed Matcha-TTS audio/label mismatch and unified formats.§3.2 says D2/D3/F1 models pretrained on VCTK, while §3.2.3 says D3 StyleTTS2 checkpoint was pretrained on LibriTTS: preserve this unresolved lineage discrepancy. Do not infer whether every copy/run includes the correction without its pinned revision/hash.
DFADD Arena-ready eval-copy commit, OCommit card identifies a test-only package: 3,755 = 755 bona fide + 3,000 spoof, five generators ×600; build note claims raw-byte embed/no re-encode and 16 kHz mono. PDF §4.2 footnote links SpeechAntiSpoofingBenchmarks/DFADD.This is a benchmark maintainer copy, not author repo isjwdu/DFADD; a current card/commit is not proof of the exact earlier main-run audio snapshot. “No re-encode” is the build record statement, not an independent waveform comparison here.
Sun et al., AI-Synthesized Voice Detection Using Neural Vocoder Artifacts, arXiv v2, 27/04/2023, M; author repository, O§4.1, Tables 1–2. Supports LibriTTS-derived self-vocoding corpus, six vocoders, and test count 18,487. Corpus design isolates vocoder conversion more narrowly than six separate full TTS pipelines.Does not establish broad detection of contemporary full TTS systems. File/hash and exact resampling in this project’s evaluation copy not audited.
Yi et al., ADD 2022: The First Audio Deep Synthesis Detection Challenge, arXiv v3, 02/07/2024, M; Track 1 eval release, O§§2–4, Tables 1–2. Separates LF/full fake Track 1 from partial Track 2 and game Track 3; reports Mandarin noisy/low-quality Track 1 and 109,199 test utterances. §3.1 explicitly partitions selected AISHELL-3 train/dev speakers.Track 1 is not a pure language-only shift: source, noise, speaker and TTS conditions can also differ. §3.3 calls test utterances unseen, but sections checked do not assert Track 1 speaker-disjointness against train/dev/adaptation; do not upgrade “unseen utterance” to that claim. Exact release/checkpoint exposure in our run is unknown.
Müller et al., Does Audio Deepfake Detection Generalize?, arXiv v5, 27/03/2026 (paper originated 2022), M§3, §4.1.2 and tables. Supports an author-collected web corpus of public figures, total 37.9 h, 17.2 h fake in the paper’s summary; PDF Table 1 for 19,963 genuine + 11,816 fake clips.Revision date is not corpus collection date. This is a curated historical web sample, not all internet audio/current generators; per-file generator ancestry and exact copy used in this evaluation are unverified.

DFADD 39.05 versus 41.87 is a historical version conflict, not a value to average. The submitted PDF reports 41.87 local AASIST / 41.86 published Arena. Older handoff reports 39.05 on a named HF rendition; it states same 3,755 IDs/labels and that recomputing Arena score files with local labels gives about 41.83, but exact audio bytes, upstream commit and run checksum are not pinned. The author repo’s April-2025 fix makes version control material, but does not by itself prove the precise cause of this old gap. Preserve the old report as history; use submitted PDF values when describing submitted results. See local audio/version check, old-domain handoff, and the class-2 dataset passport.

A05, WORLD overlap and “holdout” scope#

ASVspoof2019 §3.1/Table 1 and the A05 subsections say A05 takes speech input, processes it using WORLD features, maps the voice with a VAE, and synthesizes converted speech with WORLD. A02 and A03 also list WORLD as waveform generator, but are TTS attacks. A06 is also broad-category VC, yet its algorithm is GMM-UBM-based LPCC/MFCC conversion with spectral filtering/overlap-add, not A05’s VAE+WORLD pipeline. The database notes A06/A19 algorithm reuse. So “A05 is a VC attack with WORLD-component overlap” is supported; “A05 is algorithm-family-disjoint from all training attacks” and “A05 equals A06” are not.

For the local experiment, the PDF says A05 plus 600 bona fide were excluded from main downstream training and monitored without checkpoint selection. That protects main checkpoint selection from the specific holdout score, but does not undo earlier exploratory variants/configuration choices informed by evaluation data; paper §4.1 discloses exploratory variants were scored on evaluation data and informed configuration choices, and §8 says those evaluation sets were not held out from model development.

2. Table 3: score authority and original baseline papers#

The score authority for published systems is Arena v1 Table 2, not each original paper’s own benchmark table and not the live leaderboard. Local Table 3 p.10 selects the top eleven Arena open-source rows on DFADD, adds two Audio-JEPA configurations, and gives its own AASIST rerun. Arena supplies each score below; each linked original publication establishes system identity/method, not this particular Arena score.

Local paper Table 3 nameArena Table 2 DFADD EER %Original cited source, authors/year/version/read level
XLSR+SLS7.54Zhang, Wen & Hu, Audio Deepfake Detection with Self-Supervised XLS-R and SLS Classifier, ACM MM 2024, DOI 10.1145/3664647.3681345; OpenReview-hosted paper. Opened method and metadata (M, selected).
TCM8.88Truong et al., Temporal-Channel Modeling in Multi-head Self-Attention for Synthetic Speech Detection, Interspeech 2024; ISCA primary PDF, DOI 10.21437/Interspeech.2024-659. Read abstract and §2 method (M, selected).
XLSR-Mamba10.69Xiao & Das, XLSR-Mamba: A Dual-Column Bidirectional State Space Model for Spoofing Attack Detection, arXiv 2411.10027; IEEE Signal Processing Letters 2025, DOI 10.1109/LSP.2025.3547861. Metadata/abstract and official repo (A/O); Arena supplies EER.
Nes2Net (Arena Table 2 spells it Nes2NetX)11.14Liu et al., Nes2Net: A Lightweight Nested Architecture for Foundation Model Driven Speech Anti-Spoofing, arXiv 2504.05657; IEEE TIFS 2025, DOI 10.1109/TIFS.2025.3626963. Original identity/metadata (A); exact local-vs-Arena label spelling remains a discrepancy.
wav2vec2-AASIST11.92Tak et al., Automatic Speaker Verification Spoofing and Deepfake Detection Using wav2vec 2.0 and Data Augmentation, Odyssey 2022; ISCA primary PDF. Read abstract and relevant §3–4 (M, selected).
RawGAT-ST23.70Tak et al., End-to-End Spectro-Temporal Graph Attention Networks for Speaker Verification Anti-Spoofing and Speech Deepfake Detection, ASVspoof 2021 Workshop; ISCA primary PDF. Read abstract and graph-method sections (M, selected).
Whisper-MesoNet24.11Kawa et al., Improved DeepFake Detection Using Whisper Features, Interspeech 2023, arXiv 2306.01428, DOI 10.21437/Interspeech.2023-1537, ISCA PDF. Paper and Arena identify ASVspoof2021 DF subset training; M/A selected.
RawNet226.19Tak et al., End-to-End Anti-Spoofing with RawNet2, ICASSP 2021; arXiv 2011.01108. Metadata/abstract (A); Arena supplies score.
WavLM-ECAPA29.53Composite identity from Arena Table 1. Components: Chen et al., WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing, arXiv 2110.13900; Desplanques, Thienpondt & Demuynck, ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification, Interspeech 2020 primary PDF. Components identified (A); no separate original paper for this exact Arena anti-spoofing composition located here.
HuBERT-ECAPA34.56Composite identity from Arena Table 1. Components: Hsu et al., HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units, arXiv 2106.07447; same ECAPA source above. Components identified (A); Arena is direct score source.
AASIST41.86Jung et al., AASIST: Audio Anti-Spoofing using Integrated Spectro-Temporal Graph Attention Networks, arXiv 2110.01200, ICASSP 2022; official implementation. Metadata/abstract + official repo (A/O). Local PDF separately reports installed-toolkit checkpoint rerun 41.87.

Arena Table 2 is a system-level score comparison using released systems and Arena’s wrapper. Although §3.1 standardizes sample rate/metric/toolkit, input duration is not equal for every model and systems differ in architecture, pretraining, training subset, weights and recipes. It does not isolate the contribution of the Audio-JEPA encoder/objective. Arena Table 2 gives no per-seed range for these scores. The local paper calls these the “top eleven” open-source DFADD rows; it does not include Arena’s much worse Wav2Vec2-ECAPA DFADD row. The Nes2Net/Nes2NetX naming difference remains unresolved. Whisper’s different training corpus is a caveat, not a reason to substitute another score.

3. Trace lịch sử kết quả và giới hạn provenance#

Các hàng dưới gom những nhánh ảnh hưởng cách đọc paper. Notebook ở workspace cho thấy thiết kế/code hiện có, còn outcome trong handoff giữ là H cho tới khi checkpoint, config, ID manifest, scores/logs và run artifact khớp nhau. Tôi kiểm JSON của 12 notebook được nêu: mọi cell đang có execution_count null và không có output cells. Đây là giới hạn của bản notebook trong workspace, không chứng minh notebook chưa từng chạy ở Colab/Drive; nhiều notebook ghi file kết quả sẽ lưu ngoài checkout.

Nhánh/artifactCấu hình/dataset/metric/n được ghi trong lịch sửMức bằng chứng và cách giữ claim
AASIST head trên JEPA tokens; v27 notebook, design cells 0–6Notebook thiết kế nhánh JEPA+AASIST, cùng MAE/random ablations; handoff nhánh ngoài miền ghi JEPA+AASIST 8.76 ASV19 và 65.78 In-the-Wild, so với model/head chính 6.78 và 52.00 EER. n/checkpoint/manifest không khóa lại ở đây.Notebook local C, no saved outputs; outcome chỉ H. Có thể nói “exploratory handoff reports worse scores”, không quy nguyên nhân cho graph topology/patch count khi chưa cô lập.
Mel theo native Audio-JEPA configHandoff báo A05 2.46 so khoảng 12; ASV19 eval 9.41 so 6.78; In-the-Wild 55.18 so 52.00. Config thay duration/hop/window/frequency range/patch cùng lúc.H, không phải test xác nhận trên A05. A05 đã được nhìn trong exploration; claim cũ dùng nó chọn LR đã rút. EER tệ hơn ngoài A05 không tách tác động mel, duration hay token geometry.
Epoch sweep; v35 notebook và archived v7c controlHandoff nói sweep 2/3/4/6 epoch; DFADD cải thiện đơn điệu trong các config ghi lại; LibriSeVoc tốt nhất 40.82 ở epoch 3 nhưng vẫn cao hơn AASIST 37.95; saturation score giảm khi ít epoch. Per-run seed/config artifact chưa đối chiếu lại.H + C, no local outputs. Hỗ trợ “các handoff sweeps ghi lại không cứu LibriSeVoc”; không đủ kết luận overfitting nói chung không phải nguyên nhân.
Fusion ba seedHandoff báo z-score fusion trên ASV19 7.17 so best single seed 6.78 EER. Handoff ghi ba seed; calibration/raw-score artifact không có trong notebook outputs.H; không suy ensemble diversity/calibration luôn vô ích hay fusion giảm performance trên corpus khác.
Holdout A05 / LR; v18, v17, v13, v7bTài liệu lịch sử ghi bốn lần tương quan ngược chiều, Spearman −1.00 A05 vs CodecFake trên grid đã thử; A05 chọn 1e−4, còn ASV eval tối ưu được ghi ở 3e−5.H; history of withdrawn claims §A3 rút khuyến nghị dùng A05 chọn LR. Không biến một grid exploratory thành luật tổng quát hoặc gọi A05 là validation đại diện domain shift.
Random frozen / AudioMAE control; v38 notebookHandoff ghi DFADD frozen JEPA 8.47±0.35, MAE 11.04±0.75, random 8.66±2.25; ASV frozen JEPA 17.70±1.44, random 14.25±1.48; random fine-tuned ASV 13.68 với n=1.H, no saved outputs locally, ngoài Table 4 PDF. Gợi mở vai trò head capacity/random features; chưa chứng minh pretraining vô ích hoặc checkpoint collapse. Random FT chỉ một seed.
v40 temporal patch sweep; v40 notebookHandoff so patch (4,32) với (8,32): DFADD 9.72±1.07 vs 10.75±3.48; LibriSeVoc 50.01±0.61 vs 41.96±2.93. Handoff nêu ba seed.H + C, notebook không lưu outputs. Chỉ hỗ trợ đúng hai cấu hình/corpora được ghi; không bác bỏ mọi giả thuyết về fine-grained cues.
v41 speech-pretraining preflight; notebook, replacement code, run reportHandoff ghi G1–G5 chạy trên T4, G6 ban đầu lỗi lưu Drive; replacement report ghi bảy gate xanh, đo representation và một training step. Preflight chưa phải completed continued-pretraining experiment.Handoff H báo lượt chạy, code C được kiểm; local notebook không lưu outputs. Tách “reported run” khỏi “verified artifact”; không tuyên bố đã train tiếp 20 epoch hoặc có continued-pretraining downstream result.

DFADD baseline version trail. Local audio/version check says earlier local AASIST 39.05 was scored on another HF audio rendition, with same named 3,755-file list/labels; recomputing EER on Arena released score files with local labels gives about 41.83. The author repository’s April-2025 update confirms a correction existed, but earlier exact HF revision and raw input bytes are not pinned in artifacts read here. Evidence supports a version/rendition discrepancy, not one uniquely proven causal explanation. Keep submitted-paper 41.87 and historical 39.05 separate.

4. v41 representation metrics: code semantics and input mismatch#

Code read at collapse_metrics.py and v41_o_thay_the.py; reported values are in 02-preflight-ket-qua.md §G4b. These describe the metric and source code, not a new measurement:

  • collapse_metrics(H) expects H[B,N,D], flattens token rows to [B·N,D]. rank_eff centers rows of raw H then computes singular values with svdvals. It normalizes singular values σᵢ/Σσ and exponentiates Shannon entropy; it does not use σᵢ² as weights. std_utt instead averages tokens per utterance and L2-normalizes each utterance vector before coordinate std across batch. These are distinct diagnostics.
  • For reported B=64, N=128, D=768, centered row matrix has rank at most min(64×128−1,768)=768, not 63. Therefore reported effective rank 263.1 is numerically admissible; a ≤63 bound would apply only to centering 64 utterance-level rows. It is not evidence of a rank-calculation bug.
  • v41 source reports encode_fixed calls final encoder forward output; it is not a claim that the metric pools all transformer layers or is itself a downstream utterance embedding.
  • Handoff calls G4/G4b “same 64 VoxPopuli clips,” but code does not guarantee same clip IDs: discovery and _take_clips_long reopen streaming datasets independently, and G4b skips clips shorter than 10 s (v41_o_thay_the.py:94–106) whereas the short-crop gate accepts clips at least 2.56 s. Treat as two reported batches under different input-selection rules, not a paired fixed-utterance comparison.
  • “Native” G4b first casts VoxPopuli audio to 16 kHz (v41_o_thay_the.py:97–103), then resamples it to SR_MODEL for mel (v41_o_thay_the.py:82–89). Thus f_max=16 kHz does not establish native content above the 16-kHz-source Nyquist limit of 8 kHz. The comparison is not a causal isolation of resampling, duration, hop, window, patch grid or positional embedding.
  • Handoff-reported native/detection rows are rank_eff 263.1/247.9, cos_cross .6793/.7192, std_utt .0061/.0060. They indicate anisotropy-like measurements for those speech batches under this metric, not global checkpoint collapse, causal EER explanation, or failed representations across the AudioSet distribution. Isotropic random vectors are a reference construction, not a universal healthy-encoder threshold.

5. Tài liệu nội bộ đã dùng để kiểm chéo#

Unknown nếu cần viết claim mạnh hơn#

  1. Exact dataset revision, checksums, manifests/IDs và audio transformations của từng submitted run—đặc biệt DFADD—chưa được khóa trong workspace đã đọc.
  2. Handoff có metric và số run/seed nhưng notebook local không giữ execution outputs; không thể khôi phục per-seed values/config/checkpoint chỉ từ các file này.
  3. Arena Table 2 báo một system score per corpus; baseline seed uncertainty không có ở đó. Table 2/Table 4 của chúng tôi báo SD qua ba seed ở các điều kiện xác định, không phải uncertainty qua unseen generators/speakers.
  4. A05/LR/epoch/fusion/v40/v41/random-feature notes là exploratory hoặc readiness evidence như trên. Preflight thành công không phải kết quả continued-pretraining hay bằng chứng cơ chế tổng quát.
  5. Shared source, cùng benchmark label, cùng category attack, hoặc common downstream recipe không tự là exact-overlap audit hay causal intervention; giữ scope corpus/config/metric của nguồn.
↓ Bản Markdown nguyên gốcGiữ nguyên nội dung · Công thức, bảng và nguồn đầy đủ.Các chat bàn giao được mở trong Codex.

Gõ từ khóa để tìm bài học và đoạn liên quan.