ajaudio / studyAUDIO-JEPA · RESEARCH NOTES
4 phút đọc · Toàn văn
Mục lục bài · 6 mục

Bài 1 — Quyết định trước metric: đang đo điều gì?#

Điểm vào · Tiếp: ranking.

Mục tiêu và tiền đề#

Tự tính conditional error rates, nối chúng với population risk và viết estimand cụ thể. Cần xác suất/Bayes, task/label và score hai logits.

1. Hành động và population#

Một clip score 0,8 chưa cho biết hệ tốt. Phải biết label operational, polarity, ai được accept và threshold. B là genuine hợp lệ; S là spoof theo threat model, không mọi âm thanh được công cụ xử lý. Một trial ở đây là một utterance, một quyết định cho toàn utterance. Model có thể chỉ quan sát crop: nhãn và observation unit cần phân biệt.

Đặt aτ(x)=1[s(x)≥τ]a_\tau(x)=\mathbf1[s(x)\ge\tau]. Với NB,NS>0N_B,N_S>0:

P^miss=∑i:yi=B1[si<τ]NB,P^fa=∑i:yi=S1[si≥τ]NS.\widehat P_{miss}=\frac{\sum_{i:y_i=B}\mathbf1[s_i<\tau]}{N_B},\qquad \widehat P_{fa}=\frac{\sum_{i:y_i=S}\mathbf1[s_i\ge\tau]}{N_S}.

Miss chia cho số B, false acceptance chia cho số S. Rates có điều kiện trên lớp; số spoof benchmark không tự là deployment prior. ASVspoof 5 Phase 2 v0.6 Appendix 11.1 cho convention này.

Estimand là đại lượng muốn biết: chẳng hạn “expected cost của rule đã chọn trên dev khi gặp các cuộc gọi tiếng Việt ở distribution P, conditional trên checkpoint này”. Nó khác “EER của finite eval list” hoặc “expected EER qua training runs”. Ước lượng nào cũng cần scope, dù bảng chỉ có một số.

2. Worked example: bốn scores tới population risk#

Toy tự biên soạn, không data dự án. B=[0,8;0,4], S=[0,6;0,1], τ=0,5\tau=0,5.

TrialLabelScoreDecisionError
b1B0,8AcceptKhông
b2B0,4RejectMiss
s1S0,6AcceptFalse acceptance
s2S0,1RejectKhông

Pmiss=1/2P_{miss}=1/2, Pfa=1/2P_{fa}=1/2. Empirical accuracy 2/4=50%. B-positive precision=recall=1/2. Reject là đúng khi trial S, sai khi B: không gọi mọi reject là false reject.

Nếu deployment πB=0,99,πS=0,01\pi_B=0,99,\pi_S=0,01 và conditional rates giữ nguyên, expected accuracy vẫn 50%. Với Cmiss=1,Cfa=100C_{miss}=1,C_{fa}=100, raw risk=0,99(0,5)+100(0,01)(0,5)=0,9950,99(0,5)+100(0,01)(0,5)=0,995 cost/trial. Risk có thể >1 nếu costs lớn. Chuyển rate benchmark sang population này cần conditional score distributions phù hợp; domain shift có thể làm cả hai rates đổi.

Ví dụ khác: luôn accept đạt 99% accuracy trong population 99%B, false acceptance spoof 100%, cost 1. Rule với Pmiss=0,01,Pfa=0,10P_{miss}=0,01,P_{fa}=0,10 đạt 98,91% accuracy nhưng cost 0,1099. Metric không có costs/priors có thể xếp rules theo thứ tự khác risk. Đây là law of total probability, không đo risk triển khai.

3. Measurement contract trước bảng kết quả#

Viết: population/trial unit; label/threat model; score/polarity; decision/threshold source; metric/weights/denominators; variation/scope. Ví dụ reported paper:

Population: ASV2019 LA eval partition, utterance trials, exact release/hash cần xác nhận.
Label: protocol B/S; high=B score d=zB-zS.
Observation: toolkit fixed length rồi model crop2,56s.
Metric: EER%, eval-label sweep; scorer/version/tie convention ghi riêng.
Selection: final6epoch; configuration development history disclosed.
Scope: benchmark + variation qua3runs; chưa low-FPR deployment claim.

Không điền unknown thành version đã audit. Chương 06 §§2–4/PDF §§3–4 là reported setup; checkpoint provenance vẫn mở ở hồ sơ lớp 3.

4. Counterexample#

Với balanced labels, (Pmiss,Pfa)=(0,1)(P_{miss},P_{fa})=(0,1) và (0,5;0,5)(0,5;0,5) đều accuracy 50%, nhưng costs/prior khác sẽ cho risk khác. “Cùng EER” cũng chưa nói cùng operating risk. Source/attack mix thu có chủ đích không tự random sample của mọi deepfake; finite benchmark estimand hợp lệ, inference ra web mới cần sampling/support assumptions. Tăng decimals không thêm scope.

5. Tự kiểm#

  1. Toy bốn scores, τ=0,6\tau=0,6: s1 bằng threshold bị xử lý thế nào, rates bao nhiêu?
  2. Deployment 99%B có nên thay denominator NSN_S thành tổng trials?
  3. Viết claim “đáng tin cậy” thành estimand có threshold source và population.
Đáp án reasoning
  1. s1 accept vì ≥; b2 miss. Rates 0,5/0,5. Threshold ngay trên 0,6 mới loại s1.
  2. Không. Đổi denominator thành tổng là joint error frequency trong sample mix; conditional rate cần NSN_S. Expected error ghép với prior sau, chỉ chuyển khi conditional distributions phù hợp.
  3. Ví dụ “Risk tại threshold fit trên dev, test tiếng Việt clean protocol X, priors/costs cụ thể, checkpoint Y”. Chưa có risk/uncertainty thì đây mới là câu hỏi đo.

Đào sâu: suy Accuracy=1−πBPmiss−πSPfaAccuracy=1-\pi_BP_{miss}-\pi_SP_{fa}. NIST SRE 2010 cost model là official example khác task, giữ target/non-target semantics riêng.

DỪNG LẠI & TỰ KIỂM TRA

Bạn đã giải thích được cơ chế trong bài?

↓ Bản Markdown nguyên gốcGiữ nguyên nội dung · Công thức, bảng và nguồn đầy đủ.Các chat bàn giao được mở trong Codex.

Gõ từ khóa để tìm bài học và đoạn liên quan.