ajaudio / studyAUDIO-JEPA · RESEARCH NOTES
16 phút đọc · Toàn văn
Mục lục bài · 7 mục

Lớp 6 — Hồ sơ nguồn: đường ống, checkpoint và giao thức đánh giá#

Mốc đối chiếu nguồn trực tuyến: 11/10/2026.
Phạm vi: đọc bản PDF SOICT do người dùng cung cấp, quan sát Fig. 1, đối chiếu phần Methods/implementation của hai bài SSL gốc với code/card chính thức và kiểm tra nguồn benchmark Speech DF Arena. Hồ sơ này là sổ chứng cứ cho Chương 06, không phải báo cáo tái lập. Không tải checkpoint/corpus, không chạy detector, không thực thi notebook.

Kết luận đã khóa#

  1. Đường vào của detector trong manuscript là waveform 16 kHz, chuẩn hóa độ dài bởi Speech DF Arena đến 64.600 mẫu, sau đó resample lên 32 kHz và lấy 2,56 giây (81.920 mẫu). Bài báo mô tả log-mel 128 bins, 25 ms window, 10 ms hop, giới hạn 20–8.000 Hz, trừ waveform mean, rồi pad/truncate thời gian về 256 frame. Chia patch 8×32 cho lưới 32×4 = 128 token. Đây là cấu hình downstream; không phải cấu hình đầu vào native của checkpoint.
  2. Patch native Audio-JEPA là 16×16 trên spectrogram 256×128, tương ứng 16 bước thời gian × 8 dải mel. Từ 10 giây/256 frame suy ra bước thời gian danh nghĩa 39,0625 ms; mỗi patch trải 16×39,0625 = 625 ms. AudioMAE native dùng 10 ms hop, vì thế patch thời gian 16 frame trải 160 ms. Recipe downstream chung đổi cả hai sang patch 8×32, span danh nghĩa 80 ms. Bài báo nói rõ đây là comparison giữa hai checkpoint được phát hành dưới một recipe downstream chung, không cô lập objective.
  3. Checkpoint Audio-JEPA được dùng như encoder context đã huấn luyện trước; bài detector không chạy lại pretraining và không dùng predictor hay target encoder. Figure 1(a) là sơ đồ pretraining được vẽ lại từ nguồn Audio-JEPA; 1(b) là pipeline downstream của bài này.
  4. Con số 495.632 backend parameters cộng khớp chính xác từ mô tả paper: 13 trọng số gộp layer + 98.561 tham số ASP + 397.058 tham số classifier. Pooling tạo mean và standard deviation có attention, không có trainable parameters riêng trong phép tổng hợp này.
  5. Paper và artifacts upstream hiện hành không tạo thành một provenance liền mạch cho checkpoint Audio-JEPA. Paper v2, GitHub code snapshot tháng 7/2026, model-card/config revision tháng 7/2026 và checkpoint file commit tháng 5/2025 là các bản khác nhau. Code snapshot cấu hình target-normalization + normalized MSE và scheduler LR/weight-decay khác số trong paper; không thể gán cấu hình code hiện tại cho checkpoint nếu thiếu run manifest.
  6. Cấu hình downstream được paper báo cáo: 25.380 utterance ban đầu; loại attack A05 cùng 600 bona fide, còn 20.980 train utterance. Theo counts trong paper, còn 19.000 spoof và 1.980 bona fide, nên hệ số bona-fide khoảng 19.000/1.980 = 9,596. Batch 32, AdamW, weight decay .05, OneCycle, warm-up 10%, sáu epoch, mixed precision; seed 1234, 2, 3 (trừ cấu hình được đánh dấu một seed). Chọn checkpoint cuối epoch 6, không chọn theo A05 hoặc evaluation.
  7. Không trộn số AASIST toolkit đã công bố với lần rerun trong bài SOICT. Table 1 của SOICT giữ cả hai cột: rerun của họ và số published; chênh lệch LibriSeVoc lớn hơn các corpus còn lại.

1. Phiếu nguồn và mức đọc#

Nguồn / version đã pinPhần đọcClaim mà nguồn hỗ trợMức đọc / giới hạn
Bản PDF SOICT do workspace cung cấp, Audio-JEPA Representations for Speech Deepfake Detection, Nguyen Le Nguyen và Tran Van Hoai, 14 trang§§3.1–3.4, §4.1–4.3; Fig. 1; Tables 1–4; phần Discussion/Limitations cần cho attributionPipeline, backend, training, split, benchmark và cách diễn giải mà manuscript thực sự báo cáoĐọc và trích trực tiếp pypdf; render/quan sát trang 3–6, đặc biệt Fig. 1 ở p.5. Bản cục bộ không kèm DOI/version history; không suy ra trạng thái acceptance/publication từ tên file.
Audio-JEPA paper, arXiv v2, Ludovic Tuncay, Etienne Labbé, Emmanouil Benetos, Thomas Pellegrini, Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning, ICME 2025; revised 26/09/2026§§III–IV, phần data/implementation/objective/evaluationAudioSet pretraining, 32 kHz/10 s/128 mel/256 frames, patch grid, masking, encoder/predictor/teacher, loss và recipe theo paperĐọc Methods và implementation, không chỉ abstract. Bài này không phải run manifest của checkpoint đang dùng trong SOICT.
Audio-JEPA official code, commit ddd97ee9, commit 16/07/2026Config, mel transform, model/loss, training module, masks, scheduler, EMA callbackHành vi được mô tả bởi code/config ở snapshot cụ thể nàyĐây là source snapshot sau checkpoint file; code behavior không chứng minh đã chạy để tạo checkpoint.
Audio-JEPA Hugging Face model card/config, revision c65d33bf; checkpoint file commit d430e4d3Card, file list, config/revision metadataCheckpoint public và thời điểm tách biệt giữa checkpoint file với card/config mới hơnChỉ đọc metadata; không tải, hash hoặc load checkpoint. Card không cung cấp run manifest nối checkpoint với code/config commit.
AudioMAE paper, arXiv v3, Po-Yao Huang, Hu Xu, Juncheng Li, Alexei Baevski, Michael Auli, Wojciech Galuba, Florian Metze, Christoph Feichtenhofer, 12/01/2023§3; §§4.1–4.3; Appendix BAudioMAE preprocessing, native tokenization, masked reconstruction objective, architecture và pretraining recipe do paper báo cáoĐọc method và experimental setup/appendix, không chỉ abstract. AudioSet là cùng nguồn corpus ở mức tên/source, không chứng minh exact file-list trùng Audio-JEPA.
AudioMAE official code, commit bd60e296, repo archived/read-only từ 06/08/2025dataset/pretraining script/model/config-launchHành vi code và launch parameters của snapshotScript ghi 33 epochs trong khi paper ghi 32; đây là artifact mismatch, không bằng chứng phiên bản nào đã sinh checkpoint được dùng trong SOICT.
Speech DF Arena paper, arXiv v1, Sandipana Dowerah, Atharva Kulkarni, Ajinkya Kulkarni, Hoan My Tran, Joonas Kalda, Artem Fedorchenko, Benoit Fauve, Damien Lolive, Tanel Alumäe, Matthew Magimai Doss, 02/09/2025§§2–3; Tables 2–3Nguồn leaderboard/toolkit và số published cho các baselineTable 2 được dùng riêng để lấy số published; không thay cho rerun của bài SOICT.
Speech DF Arena official code, commit 97a4445b, commit 26/02/2026DataModule, evaluate wrapper, metrics, AASIST outputHành vi source ở snapshot toolkitSnapshot mới hơn benchmark paper. Cấu hình tùy chọn của toolkit không được xem là đã bật trong mọi lần chạy; manuscript nói rõ fixed-length configuration mà họ dùng.
Torchaudio Kaldi fbank API, documentation 2.1.2Parameters snip_edges, frame_length/frame_shift, window_typeÝ nghĩa mặc định: snip_edges=True chỉ xuất frame đủ chiều dài; window_type mặc định là povey, có các lựa chọn hanning/povey/etc.Đây là quy ước thư viện ở version docs đã pin; không xác định được version/arguments runtime của từng notebook nếu không có environment manifest.
Local crosswalk: JEPA code/source comparison, lớp 3, JEPA source dossier, lớp 3, training source dossier, lớp 4Các mục Audio-JEPA, source/checkpoint provenance, fixed-length/loss/samplingĐịnh vị tiền đề và lịch sử source review nội bộĐây là chỉ mục nội bộ để lần ngược nguồn. Các claim chính bên dưới gắn với PDF, primary papers, official code hoặc model card trực tiếp.

2. Đường ống của manuscript và hình#

Đầu vào và patching#

PDF §§3.1, 3.4 và 4.1 (pp.3–6) tạo ra chuỗi sau:

BướcManuscript báo cáoÝ nghĩa/giới hạn
Tải clipToolkit 16 kHz; tile hoặc truncate thành 64.600 samples (xấp xỉ 4,04 s)Độ dài này là đầu vào toolkit ở fixed-length mode. Không phải 64.600 samples sau resampling.
Resample/cropResample 16→32 kHz, lấy 2,56 s = 81.920 samples; nếu quá ngắn thì tileViệc tăng sample rate không thêm thông tin ngoài băng tần nguồn 16 kHz; model nhận chiều thời gian tương ứng 32 kHz.
Log-mel128 mel bands, window 25 ms, hop 10 ms, 20–8.000 Hz, subtract waveform meanBài báo cho thông số, không pin version thư viện frontend trong PDF. Static v38 source passes Hanning in its Kaldi fbank call; Torchaudio 2.1.2 docs say snip_edges defaults true and emits only complete frames. For exactly 81.920 samples at 32 kHz, 25 ms window and 10 ms hop, that convention yields 254 valid frames before padding to 256. This is a conditional code/API calculation, not an executed feature-count observation.
Chiều thời gianPad hoặc truncate thành 256 framesBài báo không mô tả chính xác biên frame/centering implementation; không suy số frame trước pad chỉ từ 2,56/0,01.
Patch8 frame × 32 mel bins; grid 32×4 =128 tokens, 768 chiềuFull patch grid trước masking; downstream augmentation là mask trên feature, không phải pretraining patch mask 40–60%.
Trọng số khởi tạoInterpolate bicubic patch-embedding weights từ 16×16 sang 8×32; regenerate 2D sinusoidal positions; không class tokenGiữ 128 token tổng, nhưng thay đổi hình dạng và span patch so với native pretraining.

Figure 1 nằm p.5. Panel (a) là sơ đồ Audio-JEPA pretraining “schematically redrawn from [30]”: context branch thấy visible patches; predictor dự đoán latent ở masked locations; target encoder thấy full input, EMA update và stop-gradient. Panel (b) là downstream của manuscript: 13 readouts của context encoder → learned layer aggregation → attentive statistics pooling → hai-logit classifier → score logit-difference. Caption nói tường minh rằng downstream dùng released context encoder; không lặp pretraining. Figure là sơ đồ khái niệm, không phải trace/log của một lần chạy.

Tính lại backend#

PDF p.4 mô tả đúng các kích thước và số tham số. Phép tính khớp:

  • Layer aggregation: softmax trên 13 scalar = 13 tham số. Khởi tạo logits 0, nên initial weights bằng nhau sau softmax.
  • ASP attention network: Linear(768,128), bias 128, rồi Linear(128,1), bias 1: (768×128+128)+(128×1+1) = 98.561.
  • Attentive mean/std concatenate thành 1.536-D; phép weighted mean/std không thêm tham số được nêu trong paper.
  • Classifier: LayerNorm(1536) có scale+bias = 3.072; Linear(1536,256) có bias = 393.472; Linear(256,2) có bias = 514. Tổng 397.058.
  • Backend: 13+98.561+397.058 = 495.632.

Đây là phép tính từ module shapes được paper báo cáo, không phải đếm state_dict của checkpoint. 495.632 ≈ 0,496 M, xấp xỉ 0,58% của encoder 85,4 M được manuscript nêu. Score được ghi là d = z_bonafide − z_spoof; higher means bona fide theo convention toolkit.

Downstream training do paper báo cáo#

  • Dataset nguồn: ASVspoof2019 LA train, 25.380 utterance (2.580 bona fide, 22.800 spoof).
  • Split: hold out attack A05 và 600 bona fide, để lại 20.980 utterance train. Theo counts, 19.000 spoof /1.980 bona fide = 9,596; paper gọi khoảng 9,6 và dùng trọng số đó cho bona-fide class trong weighted cross-entropy.
  • Frozen: encoder evaluation mode, no gradients, train backend. Fine-tuned: encoder + backend; peak LR encoder 3×10⁻⁵, backend 10⁻³ trong cả hai regime.
  • Cả hai: AdamW, weight decay .05, OneCycle schedule, 10% warm-up, batch 32, mixed precision, đúng sáu epoch.
  • Training augmentation duy nhất được mô tả: một time band tối đa 16/256 frame và một frequency band tối đa 16/128 mel bins, vị trí ngẫu nhiên và sample lại cho mỗi spectrogram. Đây là downstream SpecAugment-style masking.
  • Seed 1234, 2, 3 cho mỗi main config, trừ ô được đánh dấu single-seed. Chấm final six-epoch checkpoint; không early stop, không chọn checkpoint theo A05 hay evaluation. Paper thừa nhận exploratory variants có xem evaluation để inform choices, nên phân biệt điều đó với main checkpoint selection statement.

3. Native pretraining và recipe downstream chung#

Thuộc tínhAudio-JEPA nativeAudioMAE nativeDownstream chung trong SOICT
Input length / sample rate10 s / 32 kHz10 s / mono 16 kHz2,56 s / 32 kHz sau toolkit 16 kHz
Spectrogram128 mel ×256 time bins128 mel ×1.024 time bins128 mel ×256 time bins
Nominal hop10.000 ms /256 =39,0625 ms10 ms10 ms
Patch size16 time ×16 mel16 time ×16 mel8 time ×32 mel
Native patch grid16×8 =12864×8 =51232×4 =128
Temporal patch span16×39,0625 =625 ms16×10 =160 ms8×10 =80 ms
Pretraining targetContinuous masked latent prediction; context/predictor vs EMA target branchReconstruct spectrogram at masked patchesKhông pretraining; supervised bona-fide/spoof CE

Audio-JEPA sources: paper v2 §§III–IV, data config, mel transform, mask config, mask implementation, model loss config. AudioMAE sources: paper §§3, 4.2–4.3 and Appendix B; target time bins in code, frame shift and pad/truncate, ViT-B encoder.

At this Audio-JEPA code snapshot, the transform derives hop as clip_length_ms / target_time_bins and frame_length as 2.5×hop, subtracts waveform mean, calls Kaldi fbank with explicit Hanning and 20 Hz-to-Nyquist limits, then pads/truncates to the configured time-bin count. This is source behavior at commit ddd97ee9, not proof that the May 2025 checkpoint was produced with this transform.

So sánh objective: điều paper cho phép nói#

Audio-JEPA pretraining predicts continuous target-encoder representations at masked positions; it does not reconstruct mel patches. AudioMAE decodes masked spectrogram patches and minimizes input-space reconstruction MSE over unknown patches. Downstream in this manuscript both become ordinary supervised detectors with the same layer aggregation, pooling, classifier, dataset split, and training recipe. Thus “latent prediction beat reconstruction in these released checkpoints under the common recipe” is a bounded result; “the objective alone caused the gap” is not identified by this design. The paper itself notes different pretraining mask ratio, sample rate, optimizer/compute and unequal displacement from native input geometries.

4. Checkpoint and upstream artifact provenance#

Audio-JEPA: paper, code, checkpoint card are different records#

  • Paper v2: reports AudioSet-2M, 1.921.982 clips / 5.338 hours; 32 kHz, 10 s, 128 mel bands, 256 temporal frames; 16×16 spectrogram patches; continuous latent prediction and EMA target. It reports a ViT-Base context encoder and separately predictor/target components. Its stated pretraining setup uses batch 256, AdamW, β=(.9,.95), weight decay .05, LR 1e-6→3e-4 with 1k warm-up steps then cosine decay, and 100,000 steps (about 13 epochs; about 14 hours on four V100s). It masks 40–60% of patches. These are paper-reported settings, not verified checkpoint metadata. The method and paper loss belong to this version.
  • Pinned official code snapshot (16/07/2026): config uses target normalization plus norm_mse. Loss implementation normalizes target tokens and computes normalized-vector MSE-like loss, not the raw average squared-L2 expression as printed in the paper. Config optimizer fields list AdamW lr .0003 and WD .05; its LR scheduler says warmup 1,000 steps, start 1e-6, ref 1e-3, final 0; WD scheduler ref/final are both 1e-6 (config lines). Thus the wired scheduler values differ from the paper's reported 3e-4 peak LR and .05 WD. This establishes a mismatch between current artifacts, not which exact schedule produced the public checkpoint.
  • Model card/config revision c65d33bf (16/07/2026): later than the checkpoint file commit d430e4d3 (22/05/2025). The card’s grid axis labels conflict with its stated 256 time ×128 mel, 16×16 patch and temporal position rate; the dimensions imply 16 temporal ×8 mel tokens. Treat card grid_time/grid_freq as swapped. The card and code snapshot do not link the checkpoint to an exact training run, git SHA, data manifest or logged effective hyperparameters.
  • Downstream use: SOICT explicitly takes context encoder only. In canonical JEPA pretraining, context encoder and predictor get learning gradients, target branch is no-grad and updated by EMA. That distinction describes the upstream pretext task; it does not mean the detector has an EMA teacher or predictor.

Relevant code wiring: context/predictor vs full-input target path, EMA callback. Those are public source behavior only; no checkpoint was loaded or matched to a run.

AudioMAE: paper/code mismatch to preserve#

Paper v3 reports mono 16 kHz audio, 25 ms Hanning window/10 ms hop, 128 mel bands, 10 s yielding 1×1.024×128; non-overlapping 16×16 patches; 80% random masking during pretraining; ViT-B encoder; MSE reconstruction on masked regions; 32 epochs, batch 512, LR 2e-4, 64 V100 GPUs, about 36 hours, random start/cyclic 10 s extraction and magnitude jitter ±6 dB. Direct source locators: paper §§3, 4.2–4.3 and Appendix B.

Official code at commit bd60e296 sets target_time_bins=1024 and dataset frame_shift=10 ms. Official launch script specifies epochs=33, among its run flags; this differs from paper’s 32. Do not silently select one value. The paper says AudioSet has approx. 1.96M unbalanced plus 21K balanced clips and uses their union for pretraining. Audio-JEPA reports its own filtered count; common source name does not establish identical clip membership.

The SOICT paper identifies its imported model using the timm-port string vit_base_patch16_1024_128.audiomae_as2m, with ViT-Base, patch16 and 1024×128 pretraining input. The official AudioMAE repository README links a Google Drive checkpoint, but the source review did not establish a cryptographic mapping between SOICT’s timm-port string, a specific upstream file hash, and the AudioMAE paper run. Mark that lineage unknown.

5. Speech DF Arena: benchmark implementation vs published values#

SOICT §4.1 says the runs use the toolkit fixed-length protocol: load at 16 kHz, tile/truncate to 64.600 samples, score each utterance, and compute EER from bona-fide and spoof score lists. The official repo’s DataModule length handling and audio loading/fixed-length branch show tile/truncate and 16 kHz paths; metric code defines the toolkit EER routine. In that function, the implementation forms the DET arrays, selects the nearest point by argmin(abs(FRR−FAR)) and averages those two rates; it does not interpolate the crossing. The manuscript score is bona-fide logit minus spoof logit, consistent with higher = more bona fide.

Table 1 of the SOICT paper compares two distinct numbers, EER percent:

CorpusAASIST rerun in SOICT pipelinePublished AASIST in Speech DF Arena paper
ASVspoof2019 LA eval0.830.82
DFADD41.8741.86
ADD 2022 Track 147.9247.91
LibriSeVoc37.9537.65
In-the-Wild43.0243.00

The first column is the authors’ rerun and pipeline verification; the second is a quoted published benchmark result. Speech DF Arena Table 2 is the primary locator for published baseline scores. The repo snapshot is later than the leaderboard paper, so code state alone cannot identify the exact toolkit commit used for either published result or SOICT rerun.

6. Limits and open provenance#

  • No public artifact inspected here links the May 2025 Audio-JEPA checkpoint to a specific code SHA/config/data manifest/training log. Do not claim the July 2026 code config is its actual recipe.
  • The PDF names the public AudioMAE timm port, but this review did not match it to a hash of an official checkpoint file or run.
  • Frontend parameters are reported, but the downstream PDF does not pin a library version or fully specify frame centering/window implementation. Exact pre-pad frame count and behavior must remain implementation-dependent; 254 is only the conditional result under the referenced snip_edges=True convention and exact input length.
  • The benchmark paper’s published AASIST score, toolkit’s current code behavior and SOICT’s AASIST rerun are separate evidence items.
  • Local notebooks v25 and v38 were inspected as static JSON only: each has zero executed cell counts and zero stored outputs. They show authored code, not verified runtime outputs. v25 and v38 are distinct snapshots; do not combine the freeze/fine-tune path or other settings from one into the other. Any local code-path statement must be attributed to the relevant snapshot/version and not presented as executed evidence.
  • The local class 3/4 crosswalks help locate prior source reviews; they do not repair missing checkpoint provenance.
↓ Bản Markdown nguyên gốcGiữ nguyên nội dung · Công thức, bảng và nguồn đầy đủ.Các chat bàn giao được mở trong Codex.

Gõ từ khóa để tìm bài học và đoạn liên quan.