# Sổ nguồn lớp 4 — mức đọc và giới hạn

Tra cứu ngày 11/10/2026. **M**: method/equations liên quan được đọc, không toàn văn. **C**: code/config liên quan được đọc, không thực thi. **Meta**: landing metadata. Hai agent Luna/max đọc dossiers riêng; root kiểm lại các điểm trọng yếu trong [root review](research/root-source-review.md). Các toys là suy luận/tính toán tự biên soạn, không kết quả thực nghiệm của paper được dẫn.

## Kiến trúc, update và aggregation

| Nguồn primary, tác giả, năm/venue | Version/mức đọc | Dùng cho / không suy ra |
|---|---|---|
| [Audio-JEPA Representations for Speech Deepfake Detection](<C:/Users/LENOVO/Downloads/SOICT_2026_paper_4308.pdf>), Nguyen Le Nguyen, Tran Van Hoai, SOICT 2026 submission 4308 | PDF người dùng cung cấp; root metadata trang 1, Method trang 3–6, render trang 4 | Recipe báo cáo và tensor trace; không chứng minh checkpoint/runtime đã tái lập |
| [ASVspoof 2021 official baselines](https://github.com/asvspoof-challenge/2021/tree/9b33f5eca887bd3bf629e3fb79428cd9bb7ac972/LA), ASVspoof organizers/Todisco/Tak, 2021 | C: LFCC/CQCC MATLAB scripts, RawNet2 YAML/model/README; SHA 9b33f5e… | Density score, configs, fixed Sinc; helpers/pretrained artifacts chưa audit |
| [End-to-end anti-spoofing with RawNet2](https://arxiv.org/html/2011.01108v3), Hemlata Tak et al., ICASSP 2021 | v3, M §3/Table 1, §§4.1–4.4; root+agent | Paper recipe khác challenge config; fusion gain không chứng minh cue nhân quả |
| [AASIST: Audio Anti-Spoofing using Integrated Spectro-Temporal Graph Attention Networks](https://arxiv.org/pdf/2110.01200v1), Jee-weon Jung et al., ICASSP 2022 | v1, M §§2–4.2; [official code](https://github.com/clovaai/aasist/tree/a04c9863f63d44471dde8a6abcb3b082b07cd1), C model/config/scorer/data | Hai nhóm node, graph/readout và raw class-1 score; không tự chứng minh mọi graph head tương đương |
| [AST: Audio Spectrogram Transformer](https://arxiv.org/pdf/2104.01778), Yuan Gong, Yu-An Chung, James Glass, Interspeech 2021 | v3, M §2.1; root | Patch/ViT và image-supervised transfer; không ví dụ thuần speech SSL |
| [Robust Speech Recognition via Large-Scale Weak Supervision](https://arxiv.org/html/2212.04356v1), Alec Radford et al., ICML 2023 | v1, M §2 bởi agent; root mở nguồn | Whisper supervision khác SSL; ASR performance không chứng minh forensic suitability |
| [Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning](https://arxiv.org/abs/2507.02915v2), Ludovic Tuncay et al., ICME 2025 | v2, agent M; upstream SHA ddd97ee…; root nối [lớp 3](../lop-03-ssl-jepa/research/jepa-code-doi-chieu.md) | Pretraining vs downstream scorer; không gán current upstream defaults cho checkpoint cũ |
| [Parameter-Efficient Transfer Learning for NLP](https://proceedings.mlr.press/v97/houlsby19a/houlsby19a.pdf), Neil Houlsby et al., ICML 2019 | M §2/2.1, root | Bottleneck adapter update path; NLP evidence không chuyển thành audio gain |
| [LoRA: Low-Rank Adaptation of Large Language Models](https://arxiv.org/html/2106.09685v2), Edward J. Hu et al., ICLR 2022 | v2 M §§4.1–4.2; agent C [loralib](https://github.com/microsoft/LoRA/blob/c4593f060e6a368d7bb5af5273b8e42810cdef90/loralib/layers.py) | Rank update/init/merge; rank update không là output embedding width/rank |
| [Distilling the Knowledge in a Neural Network](https://arxiv.org/pdf/1503.02531), Geoffrey Hinton, Oriol Vinyals, Jeff Dean, 2015 report | Root M §2 | Soft teacher targets/temperature; teacher không tự là ground truth |
| [Tent: Fully Test-Time Adaptation by Entropy Minimization](https://arxiv.org/html/2006.10726v3), Dequan Wang et al., ICLR 2021 | v3, root M §§2–3 | Target-only access, entropy/norm affine updates; không biến thành frozen benchmark protocol |
| [Deep contextualized word representations](https://aclanthology.org/N18-1202.pdf), Matthew Peters et al., NAACL 2018 | M §3.2, root | Learned task-specific scalar layer mixture; NLP source không bảo đảm layer importance |
| [Attentive Statistics Pooling for Deep Speaker Embedding](https://www.isca-archive.org/interspeech_2018/okabe18_interspeech.pdf), Koji Okabe, Takafumi Koshinaka, Koichi Shinoda, Interspeech 2018 | Root+agent M §3, eqs. 3–6 | Weighted mean/population std; speaker verification evidence, không universal spoof gain |
| [ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification](https://arxiv.org/pdf/2005.07143), Brecht Desplanques, Jenthe Thienpondt, Kris Demuynck, Interspeech 2020 | Root M §3.1 | Channel-dependent attention, trục normalization; không mặc định ASP nào cũng dùng weights theo channel |
| [Attention-based Deep Multiple Instance Learning](https://proceedings.mlr.press/v80/ilse18a/ilse18a.pdf), Maximilian Ilse, Jakub Tomczak, Max Welling, ICML 2018 | Root M §2/2.3 | Bag/instance assumptions; utterance attention chưa là segment supervision |

## Loss, augmentation và data/domain

| Nguồn primary, tác giả, năm/venue | Version/mức đọc | Dùng cho / không suy ra |
|---|---|---|
| [Focal Loss for Dense Object Detection](https://arxiv.org/html/1708.02002v2), Tsung-Yi Lin et al., ICCV 2017 | v2 M §3.2/root; agent §§3–4, Detectron SHA 04155a0… | Relative hard-example emphasis; không tự phân biệt noise và informative hard cases |
| [ArcFace: Additive Angular Margin Loss for Deep Face Recognition](https://arxiv.org/html/1801.07698v3#S2.SS1), Jiankang Deng, Jia Guo, Niannan Xue, Stefanos Zafeiriou, CVPR 2019 | Root **v3 §2.1**; agent dossier v4 là bản mở rộng khác | Angular margin toy; không đồng nhất version/tác giả của bản v4 với v3 |
| [One-Class Learning Towards Synthetic Voice Spoofing Detection](https://arxiv.org/pdf/2010.13995v2), You Zhang et al., IEEE SPL 2021 | Root M v2 §II.B/eq. 3 | OC-Softmax là design cụ thể; không mọi one-class loss hoặc margin đều cùng cơ chế |
| [Supervised Contrastive Learning](https://arxiv.org/pdf/2004.11362), Prannay Khosla et al., NeurIPS 2020 | Agent M v5 §§3.1–3.2, eq. 2; SupContrast SHA 66a8fe5…; root mở PDF | Positive/negative sets và grouping; image gain không chứng minh binary fake grouping tối ưu |
| [Decoupled Weight Decay Regularization](https://arxiv.org/pdf/1711.05101), Ilya Loshchilov, Frank Hutter, ICLR 2019 | M §2, root mở paper | AdamW decay semantics; tên optimizer không đủ recipe |
| [RawBoost: A Raw Data Boosting and Augmentation Method applied to Automatic Speaker Verification Anti-Spoofing](https://arxiv.org/pdf/2111.04433v2), Hemlata Tak et al., ICASSP 2022 | v2 root §3/agent §§2–4; official SHA 4f161a8… | Waveform distortion/noise families; không đã chạy trên Audio-JEPA ở lượt này |
| [SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition](https://arxiv.org/pdf/1904.08779), Daniel Park et al., Interspeech 2019 | Agent M v3 §§2–3; root mở PDF | Feature masks/warp; project chỉ báo cáo masks, không full original policy |
| [mixup: Beyond Empirical Risk Minimization](https://arxiv.org/html/1710.09412v2), Hongyi Zhang et al., ICLR 2018 | Agent M v2 §§2–3/pseudocode; root mở nguồn | Convex input/label interpolation; soft target không fake-duration fraction |
| [An Initial Investigation for Detecting Partially Spoofed Audio](https://arxiv.org/pdf/2104.02518v2), Lin Zhang et al., Interspeech 2021 | v2 root §§2–3/agent §§2–4.2.3 | Partial segment/utterance labels và crop risk; chưa xác nhận lỗi partial crop trong Audio-JEPA |
| [A Data-Centric Approach to Generalizable Speech Deepfake Detection](https://aclanthology.org/2026.acl-long.796/), Wen Huang, Yuchen Mao, Yanmin Qian, ACL 2026, pp. 17520–17539, DOI 10.18653/v1/2026.acl-long.796 | Final Meta checked; agent trích final PDF tạm; root+agent M [arXiv v3](https://arxiv.org/html/2512.18210v3) §§3–5/Algorithms 1–2; agent Appendix C | DOSS domain composition/cap/temperature/ratio; không proof final/arXiv giống mọi result, không Group DRO |
| [Domain-Adversarial Training of Neural Networks](https://jmlr.org/papers/volume17/15-239/15-239.pdf), Yaroslav Ganin et al., JMLR 17(59), 2016 | Root M §4/eqs. 10–17; agent M | GRL update and source/target access; discriminator yếu chưa proof independence |
| [Distributionally Robust Neural Networks for Group Shifts: On the Importance of Regularization for Worst-Case Generalization](https://arxiv.org/html/1911.08731v2), Shiori Sagawa et al., ICLR 2020 | v2 root §§2–3/agent §§3.1–3.3; official SHA cbbc1c5… | Worst-group objective/regularization; không guarantee unseen conditional shift |
| [Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results](https://arxiv.org/pdf/1703.01780), Antti Tarvainen, Harri Valpola, NeurIPS 2017 | Agent M §2; root mở nguồn | EMA target/consistency; views phải giữ desired label/evidence |
| [AI-Synthesized Voice Detection Using Neural Vocoder Artifacts](https://arxiv.org/pdf/2304.13085), Chengzhe Sun et al., CVPRW 2023 | Agent M v2 §§3–4; root mở nguồn | Binary+vocoder task, shared-source self-vocoding; coverage sáu vocoders không toàn forensic world |

## Scoring và diagnostics

| Nguồn primary, tác giả, năm/venue | Version/mức đọc | Dùng cho / không suy ra |
|---|---|---|
| [On Calibration of Modern Neural Networks](https://proceedings.mlr.press/v70/guo17a/guo17a.pdf), Chuan Guo et al., ICML 2017 | Root M §4 | Positive temperature scaling; calibration chưa tăng discrimination |
| [An Image is Worth 16×16 Words: Transformers for Image Recognition at Scale](https://arxiv.org/pdf/2010.11929), Alexey Dosovitskiy et al., ICLR 2021 | Method §3.1, nguồn đã đọc trong lớp 1/3; root mở lại | Full self-attention context; chunking chưa tự là causal streaming |
| [Designing and Interpreting Probes with Control Tasks](https://aclanthology.org/D19-1275/), John Hewitt, Percy Liang, EMNLP-IJCNLP 2019 | Agent M §§2–4; root PDF fetch lỗi, dùng dossier | Probe capacity/control; decodability chưa detector reliance |
| [Attention is not Explanation](https://aclanthology.org/N19-1357.pdf), Sarthak Jain, Byron Wallace, NAACL 2019 | Root M §4.2/agent method | Permutation/counterfactual attention với fixed h; NLP findings không phủ định mọi attention audio |
| [Sanity Checks for Saliency Maps](https://papers.nips.cc/paper/8160-sanity-checks-for-saliency-maps.pdf), Julius Adebayo et al., NeurIPS 2018 | Root M §§3–4/agent full paper | Model/data randomization tests; qua check chưa causal proof |
| [Similarity of Neural Network Representations Revisited](https://proceedings.mlr.press/v97/kornblith19a/kornblith19a.pdf), Simon Kornblith et al., ICML 2019 | Root M §§2–3/Table 1 | CKA invariances/geometry; không tự nói label/nuisance usefulness |

## API và nguồn bổ sung chỉ nằm trong dossiers

[PyTorch 2.14 autograd](https://docs.pytorch.org/docs/2.14/notes/autograd.html), [CE](https://docs.pytorch.org/docs/2.14/generated/torch.nn.CrossEntropyLoss.html), [BCE](https://docs.pytorch.org/docs/2.14/generated/torch.nn.BCEWithLogitsLoss.html), [BatchNorm](https://docs.pytorch.org/docs/2.14/generated/torch.nn.BatchNorm1d.html), [Dropout](https://docs.pytorch.org/docs/2.14/generated/torch.nn.Dropout.html), [sampler](https://docs.pytorch.org/docs/2.14/data.html#torch.utils.data.WeightedRandomSampler), [STFT](https://docs.pytorch.org/docs/2.14/generated/torch.stft.html): root đọc các contract/công thức liên quan; agent độc lập đọc CE/BCE/sampler 2.9 và stable grad-mode APIs. Docs version không xác nhận notebook runtime version.

Dossier training còn có AM-Softmax, uncertainty-weighted multitask, speed perturbation, Teffic-Audio 2026 và silence-shortcut study; metadata/section/limits nằm ngay từng mục [dossier](research/training-sources.md). Chúng không là danh sách kỹ thuật đã chạy trong dự án. Dossier kiến trúc chỉ lấy được **fragment** ASVspoof 2021 evaluation plan §§7.2/8, nên không dùng như audit toàn protocol.

Không gán full-text read cho nguồn chỉ được mở landing/abstract, không nâng source observation thành runtime fact. Nếu cần triển khai/tái lập sau này, phải kiểm exact version, helpers, model artifact và protocol; lượt này chỉ hoàn thiện học liệu.
