# Bài 3 — Adaptation, gradient và đơn vị huấn luyện

[Trước](02-WAVEFORM-TOI-SCORE.md) · [Tiếp](04-DU-LIEU-PROTOCOL-BASELINE.md).

## Mục tiêu và tiền đề

Phân biệt forward/update, loss weight/sampling, epoch/step và monitoring/selection. Cần [gradient lớp 1](../lop-01-nen-tang/08-MANG-NORON-ATTENTION-TOI-UU.md), [JEPA step lớp 3](../lop-03-ssl-jepa/06-JEPA-TRAINING-STEP.md), [adaptation](../lop-04-ky-thuat-detector/03-ADAPTATION-GRADIENT-PEFT.md), [loss/optimization](../lop-04-ky-thuat-detector/05-LOSS-SAMPLING-OPTIMIZATION.md). Recipe **P**: PDF §3.3 trang 4–6, §4.1 trang 6; code **C**: [dossier pipeline](research/pipeline-sources.md).

## 1. Hai stage, hai targets

Upstream: context encoder/predictor dự đoán targets tại masked positions do full-input EMA target encoder tạo. Downstream: context encoder → back-end → logits B/S → CE. Teacher và predictor không tham gia gradient của detector này.

~~~mermaid
flowchart LR
 X[Log-mel] --> F[Context encoder θ]
 F --> H[Layer fusion ASP classifier ψ]
 H --> Z[Logits]
 Y[B/S labels] --> L[Weighted CE]
 Z --> L
 L -. gradient khi full .-> F
 L -. gradient cả hai regimes .-> H
~~~

Frozen khóa learning của θ, vẫn học ψ. v38 dùng `encoder.eval()`, `requires_grad=False`, `no_grad`/detach. `eval()` riêng chỉ đổi train/eval behaviors phù hợp, không chặn gradient. Full mở updates encoder; peak LR encoder 3×10⁻⁵, head 10⁻³. Frozen head vẫn phi tuyến qua ASP/tanh/softmax/std/LN/GELU, chưa phải linear probe.

## 2. Counts và weighted CE

P §4.2 trang 7: 25.380=2.580B+22.800S; holdout A05+600B; train còn 20.980. Suy ra **T**: A05 có 3.800S; train còn 1.980B+19.000S. Weight B=19000/1980≈9,59596, weight S=1. Đây chưa là raw-manifest audit.

v38 dùng weights `[1,ns/nb]`, y=1 cho B. Với hard labels, default weighted mean của [PyTorch CE](https://docs.pytorch.org/docs/stable/generated/torch.nn.CrossEntropyLoss.html) là:

$$L=\frac{\sum_iw_{y_i}[-\log q_{y_i}(x_i)]}{\sum_iw_{y_i}}.$$

Toy batch 1B+1S, mỗi loss là 0,693147: L vẫn 0,693147, nhưng contributions B:S≈9,596:1. Batch chỉ có một B sẽ hủy scalar weight qua denominator. Trung bình các minibatches chưa bằng global balancing của full dataset. Weight không tạo genuine sources còn thiếu.

Dưới ideal weighted-CE optimum, odds bị nhân w_B/w_S so với train posterior; cân contributions chưa có nghĩa calibrated deployment probability. Paper dùng d cho EER ranking, chưa báo calibration/DCF. [Xác suất/loss lớp 1](../lop-01-nen-tang/06-XAC-SUAT-LOSS-DETECTION.md).

## 3. Sáu epochs chưa tự xác định successful updates

P: sáu epochs, batch 32, AdamW WD 0,05, OneCycle warmup 10%, AMP, seeds 1234/2/3, final checkpoint. C v38 có `drop_last=True`:

| Đại lượng T cho template | Giá trị |
|---|---:|
| Batches mỗi epoch | floor(20980/32)=655 |
| Examples mỗi epoch | 655×32=20.960 |
| Bỏ mỗi shuffled epoch | 20 |
| Scheduled iterations sáu epochs | 3.930 |
| Example exposures | 125.760 |
| Nếu drop_last=False | 656 batches/epoch, 3.936 iterations |

Đây là arithmetic cho template, chưa xác nhận main-run steps hoặc unique examples seen. AMP có thể skip optimizer update khi overflow; cần logs để đếm successful updates. OneCycle peak khác constant LR; warmup 10% của 3930 xấp xỉ 393 schedule steps, exact transition tùy implementation. Seeds chưa đảm bảo toàn bộ data ordering/augmentation/workers deterministic.

## 4. Augmentation tác động vào evidence

P: một time band ≤16/256=6,25%; một mel band ≤16/128=12,5%. C v38 sample widths 0..16. Nếu cả hai ở maximum, union grid cells bị zero là:

$$1-(1-16/256)(1-16/128)=17,96875\%.$$

Đây khác 40–60% patch removal của upstream. Mask có thể giữ nhãn B/S nhưng che forensic cue; label-preserving chưa đảm bảo evidence-preserving. **Phản ví dụ:** toy có nhãn nằm ở một narrow band; mask che band giữ nhãn file nhưng tạo input collision giữa classes. [Augmentation lớp 4](../lop-04-ky-thuat-detector/06-AUGMENTATION-LABEL-EVIDENCE.md).

## 5. Monitoring và model selection là hai quyết định

Main runs dùng final six-epoch checkpoint, monitor A05 nhưng không early stop (P §4.1). Paper cũng báo exploratory evaluation đã ảnh hưởng config choices. Không dùng A05 chọn epoch main run chưa khôi phục evaluation thành untouched test. Cần trình bày development history, chưa suy từ đó fraud hay unbiased final-performance estimate. [Selection lớp 5](../lop-05-danh-gia-thuc-nghiem/06-SPLIT-LEAKAGE-SELECTION.md), [history](07-LICH-SU-THU-NGHIEM.md).

## 6. Tự kiểm

1. `eval()` đã đủ khóa updates chưa? Vì sao còn dùng `no_grad`?
2. Weight ≈9,6 dùng counts trước hay sau holdout?
3. Sáu epochs có 3.936 successful updates không? Nêu assumptions.
4. “A05 không chọn checkpoint” có chứng minh evaluation chưa ảnh hưởng recipe không?

<details>
<summary>Đáp án và reasoning</summary>

1. Chưa. `no_grad` tránh graph; `requires_grad` và optimizer branch kiểm parameter learning.
2. Sau holdout: 19.000S/1.980B, chưa phải 22.800/2.580.
3. Template drop_last có 3.930 scheduled batches. 3.936 chỉ khi giữ last; successful AMP updates cần logs.
4. Không. Checkpoint selection và recipe selection là hai stages; paper §4.1/§8 công bố exploratory access.

</details>

**Đào sâu:** viết contract gồm sample unit, batch/update/epoch, augmentation draws, seed, successful-step count và checkpoint rule. Đây là đọc provenance, chưa yêu cầu chạy GPU.
