Bài 3 — Adaptation, gradient và đơn vị huấn luyện#
Mục tiêu và tiền đề#
Phân biệt forward/update, loss weight/sampling, epoch/step và monitoring/selection. Cần gradient lớp 1, JEPA step lớp 3, adaptation, loss/optimization. Recipe P: PDF §3.3 trang 4–6, §4.1 trang 6; code C: dossier pipeline.
1. Hai stage, hai targets#
Upstream: context encoder/predictor dự đoán targets tại masked positions do full-input EMA target encoder tạo. Downstream: context encoder → back-end → logits B/S → CE. Teacher và predictor không tham gia gradient của detector này.
flowchart LR X[Log-mel] --> F[Context encoder θ] F --> H[Layer fusion ASP classifier ψ] H --> Z[Logits] Y[B/S labels] --> L[Weighted CE] Z --> L L -. gradient khi full .-> F L -. gradient cả hai regimes .-> H
Xem mã sơ đồ
flowchart LR
X[Log-mel] --> F[Context encoder θ]
F --> H[Layer fusion ASP classifier ψ]
H --> Z[Logits]
Y[B/S labels] --> L[Weighted CE]
Z --> L
L -. gradient khi full .-> F
L -. gradient cả hai regimes .-> HFrozen khóa learning của θ, vẫn học ψ. v38 dùng encoder.eval(), requires_grad=False, no_grad/detach. eval() riêng chỉ đổi train/eval behaviors phù hợp, không chặn gradient. Full mở updates encoder; peak LR encoder 3×10⁻⁵, head 10⁻³. Frozen head vẫn phi tuyến qua ASP/tanh/softmax/std/LN/GELU, chưa phải linear probe.
2. Counts và weighted CE#
P §4.2 trang 7: 25.380=2.580B+22.800S; holdout A05+600B; train còn 20.980. Suy ra T: A05 có 3.800S; train còn 1.980B+19.000S. Weight B=19000/1980≈9,59596, weight S=1. Đây chưa là raw-manifest audit.
v38 dùng weights [1,ns/nb], y=1 cho B. Với hard labels, default weighted mean của PyTorch CE là:
Toy batch 1B+1S, mỗi loss là 0,693147: L vẫn 0,693147, nhưng contributions B:S≈9,596:1. Batch chỉ có một B sẽ hủy scalar weight qua denominator. Trung bình các minibatches chưa bằng global balancing của full dataset. Weight không tạo genuine sources còn thiếu.
Dưới ideal weighted-CE optimum, odds bị nhân w_B/w_S so với train posterior; cân contributions chưa có nghĩa calibrated deployment probability. Paper dùng d cho EER ranking, chưa báo calibration/DCF. Xác suất/loss lớp 1.
3. Sáu epochs chưa tự xác định successful updates#
P: sáu epochs, batch 32, AdamW WD 0,05, OneCycle warmup 10%, AMP, seeds 1234/2/3, final checkpoint. C v38 có drop_last=True:
| Đại lượng T cho template | Giá trị |
|---|---|
| Batches mỗi epoch | floor(20980/32)=655 |
| Examples mỗi epoch | 655×32=20.960 |
| Bỏ mỗi shuffled epoch | 20 |
| Scheduled iterations sáu epochs | 3.930 |
| Example exposures | 125.760 |
| Nếu drop_last=False | 656 batches/epoch, 3.936 iterations |
Đây là arithmetic cho template, chưa xác nhận main-run steps hoặc unique examples seen. AMP có thể skip optimizer update khi overflow; cần logs để đếm successful updates. OneCycle peak khác constant LR; warmup 10% của 3930 xấp xỉ 393 schedule steps, exact transition tùy implementation. Seeds chưa đảm bảo toàn bộ data ordering/augmentation/workers deterministic.
4. Augmentation tác động vào evidence#
P: một time band ≤16/256=6,25%; một mel band ≤16/128=12,5%. C v38 sample widths 0..16. Nếu cả hai ở maximum, union grid cells bị zero là:
Đây khác 40–60% patch removal của upstream. Mask có thể giữ nhãn B/S nhưng che forensic cue; label-preserving chưa đảm bảo evidence-preserving. Phản ví dụ: toy có nhãn nằm ở một narrow band; mask che band giữ nhãn file nhưng tạo input collision giữa classes. Augmentation lớp 4.
5. Monitoring và model selection là hai quyết định#
Main runs dùng final six-epoch checkpoint, monitor A05 nhưng không early stop (P §4.1). Paper cũng báo exploratory evaluation đã ảnh hưởng config choices. Không dùng A05 chọn epoch main run chưa khôi phục evaluation thành untouched test. Cần trình bày development history, chưa suy từ đó fraud hay unbiased final-performance estimate. Selection lớp 5, history.
6. Tự kiểm#
eval()đã đủ khóa updates chưa? Vì sao còn dùngno_grad?- Weight ≈9,6 dùng counts trước hay sau holdout?
- Sáu epochs có 3.936 successful updates không? Nêu assumptions.
- “A05 không chọn checkpoint” có chứng minh evaluation chưa ảnh hưởng recipe không?
Đáp án và reasoning
- Chưa.
no_gradtránh graph;requires_gradvà optimizer branch kiểm parameter learning. - Sau holdout: 19.000S/1.980B, chưa phải 22.800/2.580.
- Template drop_last có 3.930 scheduled batches. 3.936 chỉ khi giữ last; successful AMP updates cần logs.
- Không. Checkpoint selection và recipe selection là hai stages; paper §4.1/§8 công bố exploratory access.
Đào sâu: viết contract gồm sample unit, batch/update/epoch, augmentation draws, seed, successful-step count và checkpoint rule. Đây là đọc provenance, chưa yêu cầu chạy GPU.