ajaudio / studyAUDIO-JEPA · RESEARCH NOTES
5 phút đọc · Toàn văn
Mục lục bài · 8 mục

Bài 10 — Efficiency, quality–cost frontier và claim có scope#

Trước · Điểm vào.

Mục tiêu và tiền đề#

Tính latency/RTF/exposures/cache, đọc frontier trong scope và tổng hợp claim khớp evidence. Cần forward vs update, inference/windows, estimand, comparison.

1. Cost không chỉ parameter count#

Total parameters quyết định một phần weight storage/compute, trainable parameters quyết định một phần gradients/optimizer. Frozen encoder vẫn forward; FLOPs còn sequence length/architecture, latency còn kernels/device/I/O/batch/precision; peak memory còn activations/workspace/cache/optimizer. Không suy wall-time ratio từ parameter ratio.

Paper báo 85,4 M encoder+0,495632 M backend≈85,9 M total; frozen trainable≈0,5 M. Weight storage 85,9 M×2 bytes=171,8 MB decimal FP16, chưa activations/runtime buffers/checkpoint metadata; full-training optimizer/gradients/master weights có thể thêm costs tùy implementation. Đây là arithmetic từ reported sizes, chưa memory measurement. Chương 06 §§2–3.

2. Worked timing example và system boundary#

Toy measurement specification: same device/software/precision, batch 1,4 s audio, warm-up xong, synchronized timer, include decode/resample/mel/transfer/forward. End-to-end 0,12 s trong đó forward 0,08 s. RTF=processing/audio duration=0,12/4=0,03; forward-only RTF 0,02 là quantity khác. Single sequential throughput≈1/0,12=8,33 clips/s và 33,33 audio-seconds/s nếu no queue/parallelism. Không suy batch 8 throughput từ batch 1 latency; cần đo batch 8 riêng.

Window 2,56 s thu xong mới chạy 0,12 s cho first output: acquisition+processing≥2,68 s trong toy. RTF<1 không có nghĩa streaming delay 0,03 s. p95 latency cần đủ repetitions/workload để estimate tails; report percentile method/count và load/queue conditions. Async accelerator clock cần synchronization để elapsed time chứa completed work. PyTorch Benchmark tutorial, docs 2.14 có warm-up/threadpool/synchronization facilities; version đọc là web docs, không version đã chạy project. Không chỉ stopwatch quanh enqueue.

3. Cache đổi cả boundary và storage#

Cache frozen features có thể bỏ encoder forward khỏi repeated head-training loops, nhưng encoder extraction phải tính one-time cost. Cache 13 layer×128 token×768 dim×4 bytes=5.111.808 bytes/clip≈4,875 MiB. Với 20.980 clips≈107,2457 GB decimal, chưa indexing/metadata. Aggregated 1536 dimfp 32 chỉ 6.144 bytes/clip nhưng có thể khóa learned layer/pooling choices; caching sau learned modules sẽ đổi what remains trainable. Caching tokens cũng không lưu được fresh waveform augmentation nếu encoder input đổi, trừ recompute/config-specific cache.

Feature-input throughput không end-to-end audio detector throughput. Cache warm/cold policy/I/O/storage/checkpoint bytes cần disclosed. Không gọi frozen cheap khi chỉ đo cached head mà baseline chạy full waveform pipeline.

4. Exposure và compute: same epochs trả lời khác same budget#

Toy datasets 100 h/300 h,20 epochs đọc mỗi giờ một lần per epoch: cumulative audio exposures 2.000 h/6.000 h. Dataset diversity và training exposures đều đổi. Fixed 2.000 h exposure:100 h×20 epochs vs 300 h×6⅔epochs, có thể khác convergence/reuse frequency. Fixedsteps cũng chưa fixed compute khi clip lengths/tokens/batch khác; fixed accelerator-hours còn optimizer/hardware throughput. Exact samples/tokens/audio-seconds drawn giúp đọc sampling thực tế, không dùng “epochs” thay toàn budget. Bouthillier 2021 pipeline-budget framing; exposure calculation là toy giả định no sampling imbalance.

Training cost toàn research include pretraining/access cost, downstream winners+search/ablation/failed runs, feature extraction, evaluation và storage nếu claim là total research budget. Deployment cost có thể exclude sunk training nhưng phải explicit. Cùng download checkpoint không nghĩa equal pretraining compute hoặc causal objective isolation.

5. Frontier trong một scope#

Toy same-device batch 1, lower EER/latency đều better:

SystemEER%Latency s
A80,10
B70,20
C90,15
D70,10

D dominatesA/B/C: no worse hai dimensions, strictly better ít nhất một. D là Pareto frontier trong bốn points đo này. Chưa “globally Pareto optimal”; chưa xét actDCF/memory/train cost hoặc confidence intervals. Nếu latency intervals overlap, point-estimate domination chưa strong uncertain-frontier claim. Một config 0,5 Mtrainable chưa đủ frontier; cần quality/cost pairs trong common scope. MLPerf inference rules là primary example quality target gắn scenario/workload, link master chưa pinned nên chỉ dùng framing, không claim tuân thủ benchmark.

6. Viết claim năm phần#

Setting → intervention/contrast → observation → uncertainty → scope. Ví dụ reported, không rerun: “Dưới common downstream recipe, hai released ViT checkpoints được so ở bốn corpus/regime comparisons. Audio-JEPA có EER mean thấp hơn AudioMAE ở3/4; mỗi comparison 3 seeds có sampleSD/ranges. Đây là checkpoint-system comparison; pretraining recipes còn khác, không identify objective cause.” Chương 06 Table 4/§5.

Claim “efficient” cần measured budget/boundary; “robust” cần shift axis/operating metric/uncertainty; “mechanism” cần contrasts/interventions loại competing explanations. Kết quả hữu ích có thể là method, evidence về mechanism, protocol/data hoặc reliability/cost. Novelty, clarity và evidence cần tương ứng câu hỏi; không checklist nào bảo đảm acceptance. Lớp 5 giúp tự đánh giá claim, không chọn venue/direction.

7. Bài tập tổng hợp#

  1. Detector frozen 0,5 Mtrainable có inference model 0,5 M không?
  2. Toy time 0,12 s/4 s: RTF và initial delay windows 2,56 s là gì?
  3. 100→300 h same 20 epochs có isolate diversity không?
  4. FrontierD có chứng minh optimal mọi devices/memory constraints không?
  5. Viết lại “JEPA giữ artifacts tốt hơn MAE và nhẹ hơn” thành hai questions có evidence requirements.
Đáp án reasoning
  1. Không; encoder 85,4 M vẫn forward, total≈85,9 M.
  2. RTF 0,03; first output≥2,68 s theo toy acquisition+compute. Ratio/cadence/delay khác nhau.
  3. Không; cumulative exposures 3×. Fixed exposure hỏi khác optimal-after-convergence; disclose choices.
  4. Không; frontier chỉ four measured points trong defined hardware/metric dimensions, uncertainty còn giới hạn.
  5. Artifact retention: cue definition, representation/decision interventions và matched objective/data/budget với controls. Cost: end-to-end timing/memory/training/search và quality under same scope. PaperEER/params hiện tại chưa trả lời cả hai mạnh như claim.

Đào sâu: lấy một abstract giả định, tô mỗi sentence theo lý thuyết/reported/source verified/toy/inference. Cho mỗi empirical claim viết estimand, source chain, uncertainty và một counterexample. Đây là tự kiểm cuối lớp; không đánh dấu đạt khi chưa có lời giải người học.

DỪNG LẠI & TỰ KIỂM TRA

Bạn đã giải thích được cơ chế trong bài?

↓ Bản Markdown nguyên gốcGiữ nguyên nội dung · Công thức, bảng và nguồn đầy đủ.Các chat bàn giao được mở trong Codex.

Gõ từ khóa để tìm bài học và đoạn liên quan.