ajaudio / studyAUDIO-JEPA · RESEARCH NOTES
4 phút đọc · Toàn văn
Mục lục bài · 7 mục

Bài 3 — Adaptation, gradient và đơn vị huấn luyện#

Trước · Tiếp.

Mục tiêu và tiền đề#

Phân biệt forward/update, loss weight/sampling, epoch/step và monitoring/selection. Cần gradient lớp 1, JEPA step lớp 3, adaptation, loss/optimization. Recipe P: PDF §3.3 trang 4–6, §4.1 trang 6; code C: dossier pipeline.

1. Hai stage, hai targets#

Upstream: context encoder/predictor dự đoán targets tại masked positions do full-input EMA target encoder tạo. Downstream: context encoder → back-end → logits B/S → CE. Teacher và predictor không tham gia gradient của detector này.

flowchart LR
 X[Log-mel] --> F[Context encoder θ]
 F --> H[Layer fusion ASP classifier ψ]
 H --> Z[Logits]
 Y[B/S labels] --> L[Weighted CE]
 Z --> L
 L -. gradient khi full .-> F
 L -. gradient cả hai regimes .-> H
Xem mã sơ đồ
flowchart LR
 X[Log-mel] --> F[Context encoder θ]
 F --> H[Layer fusion ASP classifier ψ]
 H --> Z[Logits]
 Y[B/S labels] --> L[Weighted CE]
 Z --> L
 L -. gradient khi full .-> F
 L -. gradient cả hai regimes .-> H

Frozen khóa learning của θ, vẫn học ψ. v38 dùng encoder.eval(), requires_grad=False, no_grad/detach. eval() riêng chỉ đổi train/eval behaviors phù hợp, không chặn gradient. Full mở updates encoder; peak LR encoder 3×10⁻⁵, head 10⁻³. Frozen head vẫn phi tuyến qua ASP/tanh/softmax/std/LN/GELU, chưa phải linear probe.

2. Counts và weighted CE#

P §4.2 trang 7: 25.380=2.580B+22.800S; holdout A05+600B; train còn 20.980. Suy ra T: A05 có 3.800S; train còn 1.980B+19.000S. Weight B=19000/1980≈9,59596, weight S=1. Đây chưa là raw-manifest audit.

v38 dùng weights [1,ns/nb], y=1 cho B. Với hard labels, default weighted mean của PyTorch CE là:

L=∑iwyi[−log⁡qyi(xi)]∑iwyi.L=\frac{\sum_iw_{y_i}[-\log q_{y_i}(x_i)]}{\sum_iw_{y_i}}.

Toy batch 1B+1S, mỗi loss là 0,693147: L vẫn 0,693147, nhưng contributions B:S≈9,596:1. Batch chỉ có một B sẽ hủy scalar weight qua denominator. Trung bình các minibatches chưa bằng global balancing của full dataset. Weight không tạo genuine sources còn thiếu.

Dưới ideal weighted-CE optimum, odds bị nhân w_B/w_S so với train posterior; cân contributions chưa có nghĩa calibrated deployment probability. Paper dùng d cho EER ranking, chưa báo calibration/DCF. Xác suất/loss lớp 1.

3. Sáu epochs chưa tự xác định successful updates#

P: sáu epochs, batch 32, AdamW WD 0,05, OneCycle warmup 10%, AMP, seeds 1234/2/3, final checkpoint. C v38 có drop_last=True:

Đại lượng T cho templateGiá trị
Batches mỗi epochfloor(20980/32)=655
Examples mỗi epoch655×32=20.960
Bỏ mỗi shuffled epoch20
Scheduled iterations sáu epochs3.930
Example exposures125.760
Nếu drop_last=False656 batches/epoch, 3.936 iterations

Đây là arithmetic cho template, chưa xác nhận main-run steps hoặc unique examples seen. AMP có thể skip optimizer update khi overflow; cần logs để đếm successful updates. OneCycle peak khác constant LR; warmup 10% của 3930 xấp xỉ 393 schedule steps, exact transition tùy implementation. Seeds chưa đảm bảo toàn bộ data ordering/augmentation/workers deterministic.

4. Augmentation tác động vào evidence#

P: một time band ≤16/256=6,25%; một mel band ≤16/128=12,5%. C v38 sample widths 0..16. Nếu cả hai ở maximum, union grid cells bị zero là:

1−(1−16/256)(1−16/128)=17,96875%.1-(1-16/256)(1-16/128)=17,96875\%.

Đây khác 40–60% patch removal của upstream. Mask có thể giữ nhãn B/S nhưng che forensic cue; label-preserving chưa đảm bảo evidence-preserving. Phản ví dụ: toy có nhãn nằm ở một narrow band; mask che band giữ nhãn file nhưng tạo input collision giữa classes. Augmentation lớp 4.

5. Monitoring và model selection là hai quyết định#

Main runs dùng final six-epoch checkpoint, monitor A05 nhưng không early stop (P §4.1). Paper cũng báo exploratory evaluation đã ảnh hưởng config choices. Không dùng A05 chọn epoch main run chưa khôi phục evaluation thành untouched test. Cần trình bày development history, chưa suy từ đó fraud hay unbiased final-performance estimate. Selection lớp 5, history.

6. Tự kiểm#

  1. eval() đã đủ khóa updates chưa? Vì sao còn dùng no_grad?
  2. Weight ≈9,6 dùng counts trước hay sau holdout?
  3. Sáu epochs có 3.936 successful updates không? Nêu assumptions.
  4. “A05 không chọn checkpoint” có chứng minh evaluation chưa ảnh hưởng recipe không?
Đáp án và reasoning
  1. Chưa. no_grad tránh graph; requires_grad và optimizer branch kiểm parameter learning.
  2. Sau holdout: 19.000S/1.980B, chưa phải 22.800/2.580.
  3. Template drop_last có 3.930 scheduled batches. 3.936 chỉ khi giữ last; successful AMP updates cần logs.
  4. Không. Checkpoint selection và recipe selection là hai stages; paper §4.1/§8 công bố exploratory access.

Đào sâu: viết contract gồm sample unit, batch/update/epoch, augmentation draws, seed, successful-step count và checkpoint rule. Đây là đọc provenance, chưa yêu cầu chạy GPU.

DỪNG LẠI & TỰ KIỂM TRA

Bạn đã giải thích được cơ chế trong bài?

↓ Bản Markdown nguyên gốcGiữ nguyên nội dung · Công thức, bảng và nguồn đầy đủ.Các chat bàn giao được mở trong Codex.

Gõ từ khóa để tìm bài học và đoạn liên quan.