# Sổ đăng ký nguồn lớp 6

Đối chiếu11/10/2026. Sổ chỉ rõ nguồn hỗ trợ claim nào và đã đọc tới đâu. Không dùng abstract để kết luận method, không xem code/card hiện tại là run manifest của released weights. **M:** đọc method/setup liên quan; **O:** official repo/card/protocol; **A:** metadata/abstract; **P/C/H/T:** nhãn bằng chứng trong [điểm vào](00-BAT-DAU-LOP-06.md).

## 1. Nguồn trung tâm

| Tên/tác giả/năm/version/link | Phần đã đọc | Claim được hỗ trợ; giới hạn |
|---|---|---|
| *Audio-JEPA Representations for Speech Deepfake Detection*, Nguyen Le Nguyen, Tran Van Hoai; [PDF workspace 14 trang](<C:/Users/LENOVO/Downloads/SOICT_2026_paper_4308.pdf>), bản nộp 2026 | Root đọc đủ 14 trang; trực tiếp kiểm render pp.3/4/5/7/9/10: Fig. 1/captions, Tables1–4/footnotes | P cho recipe/results/limitations; chưa rerun, chưa acceptance check |
| *Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning*, Tuncay, Labbé, Benetos, Pellegrini; ICME 2025, [arXiv v2 revised 26/09/2026](https://arxiv.org/pdf/2507.02915v2) | M §§III–IV, objective/input/setup; root đọc method, agent bổ sung version/setup | Upstream targets/geometry/reported recipe; chưa exact checkpoint run |
| [Audio-JEPA official code ddd97ee9](https://github.com/LudovicTuncay/Audio-JEPA/tree/ddd97ee9b88c00572f59bf2a00eb42120e6a4dea),16/07/2026 | C/O config/mel/loss/EMA/scheduler | Source behavior tại SHA; differences với paper được giữ |
| [Audio-JEPA HF card/config c65d33bf](https://huggingface.co/ltuncay/Audio-JEPA/tree/c65d33bfdef48ccfada785493f3cc0db7409c06f),16/07/2026; [weight-file commit d430e4d3](https://huggingface.co/ltuncay/Audio-JEPA/commit/d430e4d32d27d22f1f0b1b5853711605129539ff),22/05/2025 | O metadata/card/config; không tải/load/hash weights | Card sau file; grid labels đảo; exact run-to-checkpoint unknown |
| *Masked Autoencoders that Listen*, Huang et al.; NeurIPS 2022, [v3 12/01/2023](https://arxiv.org/html/2207.06405v3), [v2 root đọc](https://arxiv.org/html/2207.06405v2) | M §3/§§4.1–4.3; agent AppendixB | Input reconstruction/geometry/reported budget; chưa sole objective effect |
| [AudioMAE official code bd60e296](https://github.com/facebookresearch/AudioMAE/tree/bd60e29651285f80d32a6405082835ad26e6f19f),13/05/2023 | C/O dataset/model/launch script | Script33 epochs vs paper 32; timm-port exact hash lineage chưa khóa |
| *Speech DF Arena*, Dowerah et al.; [arXiv v1 02/09/2025](https://arxiv.org/html/2509.02859v1) | M §§2–3, Table 2; root kiểm training/input/scores | Authority published baseline scores; hệ phát hành, chưa matched retraining |
| [Arena official code 97a4445b](https://github.com/Speech-Arena/speech_df_arena/tree/97a4445bc82638d7b0ebd62cc4d21f1d55ffb824),26/02/2026 | C/O datamodule/wrapper/metrics; root đọc metric trực tiếp | Nearest FAR/FRR averaging, no interpolation tại SHA; exact main-run commit chưa khóa |
| [Torchaudio Kaldi fbank 2.1.2](https://docs.pytorch.org/audio/2.1/generated/torchaudio.compliance.kaldi.fbank.html) | O frame_length/frame_shift/snip_edges/window | Conditional 254valid frames với args đọc; chưa observed runtime count |
| [PyTorch CrossEntropyLoss2.14](https://docs.pytorch.org/docs/2.14/generated/torch.nn.CrossEntropyLoss.html), docs đọc11/10/2026 | O weighted hard-label mean semantics, mục class indices/reduction; version runtime chưa pin | Denominator sum target weights; chưa xác nhận environment của main run |

Full author lists/config locators/commits và differences ở [pipeline dossier](research/pipeline-sources.md). Đọc v2/v3 MAE được ghi riêng; không tự gán revision mới hơn cho checkpoint.

## 2. Dữ liệu và baselines

| Primary source/version | Read scope | Claim/unknown |
|---|---|---|
| Wang et al., ASVspoof2019 database; [v4 14/07/2020](https://arxiv.org/html/1911.01601v4), [official plan](https://www.asvspoof.org/asvspoof2019/asvspoof2019_evaluation_plan.pdf) | M/O construction/Table 1/A05/A06/A16/A19 | Attack-ID/components/reuse; không khóa local holdout manifests |
| Du et al., DFADD; [v1 13/09/2024](https://arxiv.org/html/2409.08731v1), [author repo](https://github.com/isjwdu/DFADD), [HF eval-copy commit](https://huggingface.co/datasets/SpeechAntiSpoofingBenchmarks/DFADD/commit/c578c836da3b522b27d3dd85f89309f1737e5d31) | M/O §§3–4, release/card correction | Author khác eval aggregator; D3 ancestry discrepancy và exact main-run copy unresolved |
| Sun et al., *AI-Synthesized Voice Detection Using Neural Vocoder Artifacts*; [v2 27/04/2023](https://arxiv.org/pdf/2304.13085v2) | M §4.1, Tables1–2 | LibriTTS/sáu vocoders/test 18.487; chưa full-TTS generality |
| Yi et al., ADD2022; [v3 02/07/2024](https://arxiv.org/html/2202.08433v3), [Track1 release](https://zenodo.org/records/10843991) | M/O §§2–4 | Track1 tiếng Mandarin/noisy; test speaker-disjoint chưa được đoạn đọc khẳng định |
| Müller et al., *Does Audio Deepfake Detection Generalize?*; [v5 27/03/2026](https://arxiv.org/pdf/2203.16263v5), paper originated 2022 | M §3/§4.1.2/tables | Historical curated web sample; revision date khác collection date |

[Evidence dossier §2](research/evidence-sources.md) có đủ **11 original baseline references** của Table 3: XLSR+SLS, TCM, XLSR-Mamba, Nes2Net(X), wav2vec2-AASIST, RawGAT-ST, Whisper-MesoNet, RawNet2, WavLM-ECAPA, HuBERT-ECAPA, AASIST. Nó ghi authors/year/version/link/read level, phân biệt sources chỉ đọc A/O với sources đọc M. Original papers xác lập method identity; các DFADD scores lấy **Arena v1 Table 2**, không lấy chéo original result tables. Nes2Net/Nes2NetX spelling còn discrepancy.

## 3. Nguồn local và root review

[Evidence dossier §§3–5](research/evidence-sources.md) ghi lịch sử và read scope notebook. [Root review](research/root-source-review.md) ghi những điều root trực tiếp kiểm, tiếp nhận có scope, corrections và unknowns. [Gap analysis](research/00-LO-HONG-BAN-DAU.md) được lập trước lesson writing. [QA script](scripts/kiem-tra-hoc-lieu.py) chỉ kiểm numerical toys/text/links/hashes; [validation](research/validation.json) không phải detector reproducibility report.

Giữ phân biệt full-paper read của PDF local với selected-section read của sources online. Không nâng mọi reference lên full-text audit, exact-overlap audit, checkpoint authentication hoặc executed evidence.
