ajaudio / studyAUDIO-JEPA · RESEARCH NOTES
8 phút đọc · Toàn văn
Mục lục bài · 4 mục

Sổ nguồn lớp 4 — mức đọc và giới hạn#

Tra cứu ngày 11/10/2026. M: method/equations liên quan được đọc, không toàn văn. C: code/config liên quan được đọc, không thực thi. Meta: landing metadata. Hai agent Luna/max đọc dossiers riêng; root kiểm lại các điểm trọng yếu trong root review. Các toys là suy luận/tính toán tự biên soạn, không kết quả thực nghiệm của paper được dẫn.

Kiến trúc, update và aggregation#

Nguồn primary, tác giả, năm/venueVersion/mức đọcDùng cho / không suy ra
Audio-JEPA Representations for Speech Deepfake Detection, Nguyen Le Nguyen, Tran Van Hoai, SOICT 2026 submission 4308PDF người dùng cung cấp; root metadata trang 1, Method trang 3–6, render trang 4Recipe báo cáo và tensor trace; không chứng minh checkpoint/runtime đã tái lập
ASVspoof 2021 official baselines, ASVspoof organizers/Todisco/Tak, 2021C: LFCC/CQCC MATLAB scripts, RawNet2 YAML/model/README; SHA 9b33f5e…Density score, configs, fixed Sinc; helpers/pretrained artifacts chưa audit
End-to-end anti-spoofing with RawNet2, Hemlata Tak et al., ICASSP 2021v3, M §3/Table 1, §§4.1–4.4; root+agentPaper recipe khác challenge config; fusion gain không chứng minh cue nhân quả
AASIST: Audio Anti-Spoofing using Integrated Spectro-Temporal Graph Attention Networks, Jee-weon Jung et al., ICASSP 2022v1, M §§2–4.2; official code, C model/config/scorer/dataHai nhóm node, graph/readout và raw class-1 score; không tự chứng minh mọi graph head tương đương
AST: Audio Spectrogram Transformer, Yuan Gong, Yu-An Chung, James Glass, Interspeech 2021v3, M §2.1; rootPatch/ViT và image-supervised transfer; không ví dụ thuần speech SSL
Robust Speech Recognition via Large-Scale Weak Supervision, Alec Radford et al., ICML 2023v1, M §2 bởi agent; root mở nguồnWhisper supervision khác SSL; ASR performance không chứng minh forensic suitability
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning, Ludovic Tuncay et al., ICME 2025v2, agent M; upstream SHA ddd97ee…; root nối lớp 3Pretraining vs downstream scorer; không gán current upstream defaults cho checkpoint cũ
Parameter-Efficient Transfer Learning for NLP, Neil Houlsby et al., ICML 2019M §2/2.1, rootBottleneck adapter update path; NLP evidence không chuyển thành audio gain
LoRA: Low-Rank Adaptation of Large Language Models, Edward J. Hu et al., ICLR 2022v2 M §§4.1–4.2; agent C loralibRank update/init/merge; rank update không là output embedding width/rank
Distilling the Knowledge in a Neural Network, Geoffrey Hinton, Oriol Vinyals, Jeff Dean, 2015 reportRoot M §2Soft teacher targets/temperature; teacher không tự là ground truth
Tent: Fully Test-Time Adaptation by Entropy Minimization, Dequan Wang et al., ICLR 2021v3, root M §§2–3Target-only access, entropy/norm affine updates; không biến thành frozen benchmark protocol
Deep contextualized word representations, Matthew Peters et al., NAACL 2018M §3.2, rootLearned task-specific scalar layer mixture; NLP source không bảo đảm layer importance
Attentive Statistics Pooling for Deep Speaker Embedding, Koji Okabe, Takafumi Koshinaka, Koichi Shinoda, Interspeech 2018Root+agent M §3, eqs. 3–6Weighted mean/population std; speaker verification evidence, không universal spoof gain
ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification, Brecht Desplanques, Jenthe Thienpondt, Kris Demuynck, Interspeech 2020Root M §3.1Channel-dependent attention, trục normalization; không mặc định ASP nào cũng dùng weights theo channel
Attention-based Deep Multiple Instance Learning, Maximilian Ilse, Jakub Tomczak, Max Welling, ICML 2018Root M §2/2.3Bag/instance assumptions; utterance attention chưa là segment supervision

Loss, augmentation và data/domain#

Nguồn primary, tác giả, năm/venueVersion/mức đọcDùng cho / không suy ra
Focal Loss for Dense Object Detection, Tsung-Yi Lin et al., ICCV 2017v2 M §3.2/root; agent §§3–4, Detectron SHA 04155a0…Relative hard-example emphasis; không tự phân biệt noise và informative hard cases
ArcFace: Additive Angular Margin Loss for Deep Face Recognition, Jiankang Deng, Jia Guo, Niannan Xue, Stefanos Zafeiriou, CVPR 2019Root v3 §2.1; agent dossier v4 là bản mở rộng khácAngular margin toy; không đồng nhất version/tác giả của bản v4 với v3
One-Class Learning Towards Synthetic Voice Spoofing Detection, You Zhang et al., IEEE SPL 2021Root M v2 §II.B/eq. 3OC-Softmax là design cụ thể; không mọi one-class loss hoặc margin đều cùng cơ chế
Supervised Contrastive Learning, Prannay Khosla et al., NeurIPS 2020Agent M v5 §§3.1–3.2, eq. 2; SupContrast SHA 66a8fe5…; root mở PDFPositive/negative sets và grouping; image gain không chứng minh binary fake grouping tối ưu
Decoupled Weight Decay Regularization, Ilya Loshchilov, Frank Hutter, ICLR 2019M §2, root mở paperAdamW decay semantics; tên optimizer không đủ recipe
RawBoost: A Raw Data Boosting and Augmentation Method applied to Automatic Speaker Verification Anti-Spoofing, Hemlata Tak et al., ICASSP 2022v2 root §3/agent §§2–4; official SHA 4f161a8…Waveform distortion/noise families; không đã chạy trên Audio-JEPA ở lượt này
SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition, Daniel Park et al., Interspeech 2019Agent M v3 §§2–3; root mở PDFFeature masks/warp; project chỉ báo cáo masks, không full original policy
mixup: Beyond Empirical Risk Minimization, Hongyi Zhang et al., ICLR 2018Agent M v2 §§2–3/pseudocode; root mở nguồnConvex input/label interpolation; soft target không fake-duration fraction
An Initial Investigation for Detecting Partially Spoofed Audio, Lin Zhang et al., Interspeech 2021v2 root §§2–3/agent §§2–4.2.3Partial segment/utterance labels và crop risk; chưa xác nhận lỗi partial crop trong Audio-JEPA
A Data-Centric Approach to Generalizable Speech Deepfake Detection, Wen Huang, Yuchen Mao, Yanmin Qian, ACL 2026, pp. 17520–17539, DOI 10.18653/v1/2026.acl-long.796Final Meta checked; agent trích final PDF tạm; root+agent M arXiv v3 §§3–5/Algorithms 1–2; agent Appendix CDOSS domain composition/cap/temperature/ratio; không proof final/arXiv giống mọi result, không Group DRO
Domain-Adversarial Training of Neural Networks, Yaroslav Ganin et al., JMLR 17(59), 2016Root M §4/eqs. 10–17; agent MGRL update and source/target access; discriminator yếu chưa proof independence
Distributionally Robust Neural Networks for Group Shifts: On the Importance of Regularization for Worst-Case Generalization, Shiori Sagawa et al., ICLR 2020v2 root §§2–3/agent §§3.1–3.3; official SHA cbbc1c5…Worst-group objective/regularization; không guarantee unseen conditional shift
Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results, Antti Tarvainen, Harri Valpola, NeurIPS 2017Agent M §2; root mở nguồnEMA target/consistency; views phải giữ desired label/evidence
AI-Synthesized Voice Detection Using Neural Vocoder Artifacts, Chengzhe Sun et al., CVPRW 2023Agent M v2 §§3–4; root mở nguồnBinary+vocoder task, shared-source self-vocoding; coverage sáu vocoders không toàn forensic world

Scoring và diagnostics#

Nguồn primary, tác giả, năm/venueVersion/mức đọcDùng cho / không suy ra
On Calibration of Modern Neural Networks, Chuan Guo et al., ICML 2017Root M §4Positive temperature scaling; calibration chưa tăng discrimination
An Image is Worth 16×16 Words: Transformers for Image Recognition at Scale, Alexey Dosovitskiy et al., ICLR 2021Method §3.1, nguồn đã đọc trong lớp 1/3; root mở lạiFull self-attention context; chunking chưa tự là causal streaming
Designing and Interpreting Probes with Control Tasks, John Hewitt, Percy Liang, EMNLP-IJCNLP 2019Agent M §§2–4; root PDF fetch lỗi, dùng dossierProbe capacity/control; decodability chưa detector reliance
Attention is not Explanation, Sarthak Jain, Byron Wallace, NAACL 2019Root M §4.2/agent methodPermutation/counterfactual attention với fixed h; NLP findings không phủ định mọi attention audio
Sanity Checks for Saliency Maps, Julius Adebayo et al., NeurIPS 2018Root M §§3–4/agent full paperModel/data randomization tests; qua check chưa causal proof
Similarity of Neural Network Representations Revisited, Simon Kornblith et al., ICML 2019Root M §§2–3/Table 1CKA invariances/geometry; không tự nói label/nuisance usefulness

API và nguồn bổ sung chỉ nằm trong dossiers#

PyTorch 2.14 autograd, CE, BCE, BatchNorm, Dropout, sampler, STFT: root đọc các contract/công thức liên quan; agent độc lập đọc CE/BCE/sampler 2.9 và stable grad-mode APIs. Docs version không xác nhận notebook runtime version.

Dossier training còn có AM-Softmax, uncertainty-weighted multitask, speed perturbation, Teffic-Audio 2026 và silence-shortcut study; metadata/section/limits nằm ngay từng mục dossier. Chúng không là danh sách kỹ thuật đã chạy trong dự án. Dossier kiến trúc chỉ lấy được fragment ASVspoof 2021 evaluation plan §§7.2/8, nên không dùng như audit toàn protocol.

Không gán full-text read cho nguồn chỉ được mở landing/abstract, không nâng source observation thành runtime fact. Nếu cần triển khai/tái lập sau này, phải kiểm exact version, helpers, model artifact và protocol; lượt này chỉ hoàn thiện học liệu.

↓ Bản Markdown nguyên gốcGiữ nguyên nội dung · Công thức, bảng và nguồn đầy đủ.Các chat bàn giao được mở trong Codex.

Gõ từ khóa để tìm bài học và đoạn liên quan.