Sổ nguồn lớp 4 — mức đọc và giới hạn#
Tra cứu ngày 11/10/2026. M: method/equations liên quan được đọc, không toàn văn. C: code/config liên quan được đọc, không thực thi. Meta: landing metadata. Hai agent Luna/max đọc dossiers riêng; root kiểm lại các điểm trọng yếu trong root review. Các toys là suy luận/tính toán tự biên soạn, không kết quả thực nghiệm của paper được dẫn.
Kiến trúc, update và aggregation#
| Nguồn primary, tác giả, năm/venue | Version/mức đọc | Dùng cho / không suy ra |
|---|---|---|
| Audio-JEPA Representations for Speech Deepfake Detection, Nguyen Le Nguyen, Tran Van Hoai, SOICT 2026 submission 4308 | PDF người dùng cung cấp; root metadata trang 1, Method trang 3–6, render trang 4 | Recipe báo cáo và tensor trace; không chứng minh checkpoint/runtime đã tái lập |
| ASVspoof 2021 official baselines, ASVspoof organizers/Todisco/Tak, 2021 | C: LFCC/CQCC MATLAB scripts, RawNet2 YAML/model/README; SHA 9b33f5e… | Density score, configs, fixed Sinc; helpers/pretrained artifacts chưa audit |
| End-to-end anti-spoofing with RawNet2, Hemlata Tak et al., ICASSP 2021 | v3, M §3/Table 1, §§4.1–4.4; root+agent | Paper recipe khác challenge config; fusion gain không chứng minh cue nhân quả |
| AASIST: Audio Anti-Spoofing using Integrated Spectro-Temporal Graph Attention Networks, Jee-weon Jung et al., ICASSP 2022 | v1, M §§2–4.2; official code, C model/config/scorer/data | Hai nhóm node, graph/readout và raw class-1 score; không tự chứng minh mọi graph head tương đương |
| AST: Audio Spectrogram Transformer, Yuan Gong, Yu-An Chung, James Glass, Interspeech 2021 | v3, M §2.1; root | Patch/ViT và image-supervised transfer; không ví dụ thuần speech SSL |
| Robust Speech Recognition via Large-Scale Weak Supervision, Alec Radford et al., ICML 2023 | v1, M §2 bởi agent; root mở nguồn | Whisper supervision khác SSL; ASR performance không chứng minh forensic suitability |
| Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning, Ludovic Tuncay et al., ICME 2025 | v2, agent M; upstream SHA ddd97ee…; root nối lớp 3 | Pretraining vs downstream scorer; không gán current upstream defaults cho checkpoint cũ |
| Parameter-Efficient Transfer Learning for NLP, Neil Houlsby et al., ICML 2019 | M §2/2.1, root | Bottleneck adapter update path; NLP evidence không chuyển thành audio gain |
| LoRA: Low-Rank Adaptation of Large Language Models, Edward J. Hu et al., ICLR 2022 | v2 M §§4.1–4.2; agent C loralib | Rank update/init/merge; rank update không là output embedding width/rank |
| Distilling the Knowledge in a Neural Network, Geoffrey Hinton, Oriol Vinyals, Jeff Dean, 2015 report | Root M §2 | Soft teacher targets/temperature; teacher không tự là ground truth |
| Tent: Fully Test-Time Adaptation by Entropy Minimization, Dequan Wang et al., ICLR 2021 | v3, root M §§2–3 | Target-only access, entropy/norm affine updates; không biến thành frozen benchmark protocol |
| Deep contextualized word representations, Matthew Peters et al., NAACL 2018 | M §3.2, root | Learned task-specific scalar layer mixture; NLP source không bảo đảm layer importance |
| Attentive Statistics Pooling for Deep Speaker Embedding, Koji Okabe, Takafumi Koshinaka, Koichi Shinoda, Interspeech 2018 | Root+agent M §3, eqs. 3–6 | Weighted mean/population std; speaker verification evidence, không universal spoof gain |
| ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification, Brecht Desplanques, Jenthe Thienpondt, Kris Demuynck, Interspeech 2020 | Root M §3.1 | Channel-dependent attention, trục normalization; không mặc định ASP nào cũng dùng weights theo channel |
| Attention-based Deep Multiple Instance Learning, Maximilian Ilse, Jakub Tomczak, Max Welling, ICML 2018 | Root M §2/2.3 | Bag/instance assumptions; utterance attention chưa là segment supervision |
Loss, augmentation và data/domain#
| Nguồn primary, tác giả, năm/venue | Version/mức đọc | Dùng cho / không suy ra |
|---|---|---|
| Focal Loss for Dense Object Detection, Tsung-Yi Lin et al., ICCV 2017 | v2 M §3.2/root; agent §§3–4, Detectron SHA 04155a0… | Relative hard-example emphasis; không tự phân biệt noise và informative hard cases |
| ArcFace: Additive Angular Margin Loss for Deep Face Recognition, Jiankang Deng, Jia Guo, Niannan Xue, Stefanos Zafeiriou, CVPR 2019 | Root v3 §2.1; agent dossier v4 là bản mở rộng khác | Angular margin toy; không đồng nhất version/tác giả của bản v4 với v3 |
| One-Class Learning Towards Synthetic Voice Spoofing Detection, You Zhang et al., IEEE SPL 2021 | Root M v2 §II.B/eq. 3 | OC-Softmax là design cụ thể; không mọi one-class loss hoặc margin đều cùng cơ chế |
| Supervised Contrastive Learning, Prannay Khosla et al., NeurIPS 2020 | Agent M v5 §§3.1–3.2, eq. 2; SupContrast SHA 66a8fe5…; root mở PDF | Positive/negative sets và grouping; image gain không chứng minh binary fake grouping tối ưu |
| Decoupled Weight Decay Regularization, Ilya Loshchilov, Frank Hutter, ICLR 2019 | M §2, root mở paper | AdamW decay semantics; tên optimizer không đủ recipe |
| RawBoost: A Raw Data Boosting and Augmentation Method applied to Automatic Speaker Verification Anti-Spoofing, Hemlata Tak et al., ICASSP 2022 | v2 root §3/agent §§2–4; official SHA 4f161a8… | Waveform distortion/noise families; không đã chạy trên Audio-JEPA ở lượt này |
| SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition, Daniel Park et al., Interspeech 2019 | Agent M v3 §§2–3; root mở PDF | Feature masks/warp; project chỉ báo cáo masks, không full original policy |
| mixup: Beyond Empirical Risk Minimization, Hongyi Zhang et al., ICLR 2018 | Agent M v2 §§2–3/pseudocode; root mở nguồn | Convex input/label interpolation; soft target không fake-duration fraction |
| An Initial Investigation for Detecting Partially Spoofed Audio, Lin Zhang et al., Interspeech 2021 | v2 root §§2–3/agent §§2–4.2.3 | Partial segment/utterance labels và crop risk; chưa xác nhận lỗi partial crop trong Audio-JEPA |
| A Data-Centric Approach to Generalizable Speech Deepfake Detection, Wen Huang, Yuchen Mao, Yanmin Qian, ACL 2026, pp. 17520–17539, DOI 10.18653/v1/2026.acl-long.796 | Final Meta checked; agent trích final PDF tạm; root+agent M arXiv v3 §§3–5/Algorithms 1–2; agent Appendix C | DOSS domain composition/cap/temperature/ratio; không proof final/arXiv giống mọi result, không Group DRO |
| Domain-Adversarial Training of Neural Networks, Yaroslav Ganin et al., JMLR 17(59), 2016 | Root M §4/eqs. 10–17; agent M | GRL update and source/target access; discriminator yếu chưa proof independence |
| Distributionally Robust Neural Networks for Group Shifts: On the Importance of Regularization for Worst-Case Generalization, Shiori Sagawa et al., ICLR 2020 | v2 root §§2–3/agent §§3.1–3.3; official SHA cbbc1c5… | Worst-group objective/regularization; không guarantee unseen conditional shift |
| Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results, Antti Tarvainen, Harri Valpola, NeurIPS 2017 | Agent M §2; root mở nguồn | EMA target/consistency; views phải giữ desired label/evidence |
| AI-Synthesized Voice Detection Using Neural Vocoder Artifacts, Chengzhe Sun et al., CVPRW 2023 | Agent M v2 §§3–4; root mở nguồn | Binary+vocoder task, shared-source self-vocoding; coverage sáu vocoders không toàn forensic world |
Scoring và diagnostics#
| Nguồn primary, tác giả, năm/venue | Version/mức đọc | Dùng cho / không suy ra |
|---|---|---|
| On Calibration of Modern Neural Networks, Chuan Guo et al., ICML 2017 | Root M §4 | Positive temperature scaling; calibration chưa tăng discrimination |
| An Image is Worth 16×16 Words: Transformers for Image Recognition at Scale, Alexey Dosovitskiy et al., ICLR 2021 | Method §3.1, nguồn đã đọc trong lớp 1/3; root mở lại | Full self-attention context; chunking chưa tự là causal streaming |
| Designing and Interpreting Probes with Control Tasks, John Hewitt, Percy Liang, EMNLP-IJCNLP 2019 | Agent M §§2–4; root PDF fetch lỗi, dùng dossier | Probe capacity/control; decodability chưa detector reliance |
| Attention is not Explanation, Sarthak Jain, Byron Wallace, NAACL 2019 | Root M §4.2/agent method | Permutation/counterfactual attention với fixed h; NLP findings không phủ định mọi attention audio |
| Sanity Checks for Saliency Maps, Julius Adebayo et al., NeurIPS 2018 | Root M §§3–4/agent full paper | Model/data randomization tests; qua check chưa causal proof |
| Similarity of Neural Network Representations Revisited, Simon Kornblith et al., ICML 2019 | Root M §§2–3/Table 1 | CKA invariances/geometry; không tự nói label/nuisance usefulness |
API và nguồn bổ sung chỉ nằm trong dossiers#
PyTorch 2.14 autograd, CE, BCE, BatchNorm, Dropout, sampler, STFT: root đọc các contract/công thức liên quan; agent độc lập đọc CE/BCE/sampler 2.9 và stable grad-mode APIs. Docs version không xác nhận notebook runtime version.
Dossier training còn có AM-Softmax, uncertainty-weighted multitask, speed perturbation, Teffic-Audio 2026 và silence-shortcut study; metadata/section/limits nằm ngay từng mục dossier. Chúng không là danh sách kỹ thuật đã chạy trong dự án. Dossier kiến trúc chỉ lấy được fragment ASVspoof 2021 evaluation plan §§7.2/8, nên không dùng như audit toàn protocol.
Không gán full-text read cho nguồn chỉ được mở landing/abstract, không nâng source observation thành runtime fact. Nếu cần triển khai/tái lập sau này, phải kiểm exact version, helpers, model artifact và protocol; lượt này chỉ hoàn thiện học liệu.