# Báo cáo enrich — Lớp 5: đánh giá, thực nghiệm và bằng chứng

Ngày 11/10/2026. Đã phát triển chương 05 thành mười bài học sâu, có derivation, worked examples, counterexamples, bài tập/đáp án reasoning và links tới tiền đề lớp 1–4. **Trạng thái biên soạn khác trạng thái người học:** câu đầu M1 đang chờ trả lời; lớp 5 chưa được đánh dấu học xong.

[Điểm vào](00-BAT-DAU-LOP-05.md) · [nguồn](11-SO-DANG-KY-NGUON.md) · [tiến độ/handoff](13-TIEN-DO-HOC-VA-BAN-GIAO.md).

## Từ lỗ hổng tới năng lực kiểm được

| Khoảng mỏng trước enrich | Nội dung đã thêm | Worked case/counterexample | Tự kiểm |
|---|---|---|---|
| Decision chưa có estimand/population/unit | Bài 1: measurement contract, conditional denominators, priors/costs | 99%accuracy always-accept có risk cao; finite threshold tradeoff | Viết action/class/unit/population/threshold source |
| EER/ROC/PR chỉ tên/định nghĩa | Bài 2: tie-block sweep/endpoints, pair AUC, AP vs trapezoid, DET, finite interpolation, low-FPR support | AUC 13/18, AP 34/45, PR area 133/180; no exact EER point; official all-tie behavior | Tính bằng tay và khóa convention |
| DCF/Cllr chưa có Bayes derivation/calibration example | Bài 3: raw/normalized, Bayes threshold, weighted-CE offset, stable softplus, PAV oracle | Shift+10 giữ EER/minDCF nhưng actDCF/Cllr khác; EER-optimal threshold không minDCF-optimal | Phân biệt posterior/LLR/neural score và dev-fit/eval oracle |
| tDCF/SASV/localization chỉ taxonomy | Bài 4: three trial types, conditional tandem products, modern normalizer, interval/range denominators | aDCF 0.5625; tandem raw 0.195; joint spoofFA 0–50% với same marginals; timelineF 1=0.5 | Nêu task, assumptions, boundary/collar/matching |
| Pooled/macro chỉ trực giác | Bài 5: offsets, shared genuine references, weights/worst-group | Mỗi groupEER 0 nhưng pooledEER 0.5; ranking model đảo khi weights đổi | Chốt shared/global hay local threshold, weights |
| Split chưa nối development access/selection | Bài 6: ancestry graph, connected groups, access regimes, nested design, reported history | Ba descendants u1 khác bytes nhưng shared parent; winner dev 9% khi true 10% | Phân biệt exact contamination với ancestry và test feedback |
| Uncertainty chưa thao tác | Bài 7: conditional seed CI, paired group draws, crossed factors, degenerate draws, matched tests/TOST/Holm | Bốn ordered draws với Δ−1/0/+1; McNemar exactp 18/256; equivalence bounds và Holm stop | Nêu random unit, estimand, null, CI assumptions/margin |
| Mechanism chưa tách contrasts | Bài 8: fixed vs tuned recipe, objective/recipe confounds, inference/retrain ablations | Probe source 100% nhưng head không reliance; redundant cue retraining counterexample | Nêu intervention và competing explanations |
| Provenance chỉ checklist | Bài 9: chain từ waveform tới rounded table, manifests, reproduction levels | Positional join làm AUC 0.75 thành 1 dù shapes/row counts đúng | Join IDs/labels và giới hạn baseline/wrapper gates |
| Efficiency/claim còn tên axes | Bài 10: timing boundary, RTF/delay/cache/exposures, finite frontier, claim năm phần | RTF 0.03 khác 2.68 s initial delay; cache 4.875 MiB; exposures 3× | Viết measured cost/scope/uncertainty, không suy từ trainable count |

## Tài sản đã tạo và phạm vi chỉnh sửa

Folder lớp 5 có start, mười bài, ledger 11, report 12, progress 13; ba hồ sơ research (hai agent và root), initial-gap note, baseline snapshot, pinned scorer source, [script QA](scripts/kiem-tra-hoc-lieu.py), [machine output](research/validation.json) và hai [hình metrics](assets/metric-toy.png)/[cost-bootstrap](assets/cost-bootstrap-toy.png).

Ngoài folder lớp 5:
- Global 00 thêm đúng một đoạn liên kết lớp 5.
- Chương 05 giữ map; thêm deep links theo mục và sửa ambiguity về EER finite sample, PAV oracle, reported development history và seed×sample crossed design.
- Snapshot SHA-256 kiểm 241 files; chỉ hai map 00/05 được phép đổi. Không sửa lớp 1–4, chương 06–09, manuscript/notebook/recipe/data trong snapshot. Coordinator 09 không được sửa, không gửi message sang chat khác.

Hai agent **Luna max** được giao nghiên cứu riêng metrics/evidence theo yêu cầu, chỉ ghi dossier trong research. Root dùng file handoff, đọc thêm nguồn primary và quyết định claims cuối. [Root review](research/root-source-review.md) ghi những corrections và phần trực tiếp kiểm.

## Kiểm tra và cách chạy lại

Script chỉ chạy finite toys và material QA; không model/corpus. Dùng environment đã có ở workspace: tmp/lop 01-runtime/Scripts/python.exe, NumPy 2.2.6, Matplotlib 3.10.7; Python exact version ghi trong JSON. Không cài dependency mới.

~~~powershell
& '.\tmp\lop01-runtime\Scripts\python.exe' 'JEPA Research/on-tap-2026-10/lop-05-danh-gia-thuc-nghiem/scripts/kiem-tra-hoc-lieu.py'
~~~

Kết quả cuối: **passed=true**, 48 nhóm phép tính đạt, 139 local links tồn tại, cấu trúc 20 Markdown files đạt và snapshot 241 files chỉ có hai thay đổi được phép. Chi tiết xem JSON:
- **48 nhóm numerical assertions**: sweep/independent AUC area, AP/PR area, interpolated EER, min/actDCF, Cllr, modern tDCF, local rates, group-weight/crossed-design toys, exact paired draws, McNemar/Holm, mechanism/ID-join/cache/exposure/frontier.
- Imported pinned official scorer chỉ để kiểm all-tie EER, shifted-score actDCF và finite Cllr. Không xác nhận scorer Arena dùng trong paper.
- Local file links được kiểm tồn tại, Markdown fences/table column counts/details/math delimiters/encoding được kiểm cấu trúc. Không coi tồn tại link là semantic audit toàn nội dung, và không crawl mọi remote URL.
- Baseline hashes kiểm phạm vi protected; output allowed_changes chỉ gồm 00/05.
- Hai PNG đã xem trực tiếp; plots không cắt nhãn/trục, convention B-positive và DET clipping rõ. PDF reported Tables 2/4 được đối chiếu với render p9–10.

Các lỗi số phát hiện trong quá trình biên soạn đã sửa trước bàn giao: PR trapezoid 133/180; minDCF 2/3 thay 29/30; paired pooled toy coincidence 0.5 thay suy sai 1/3; Cllr/cache rounding. Những checks này xác minh phép tính minh họa, không xác minh result của paper.

## Facts còn thiếu và claims không được nâng mức

Paper EER/recipe/sizes/development history ở lớp này là **reported**. Chưa audit raw per-ID scores, toàn search history, exact lineage/sample exposure, runtime wrappers/checkpoints hay calibration/latency/memory ngoài miền. Không mới kết luận significance, objective causality, exact leakage, localization quality hoặc detector Pareto frontier thật.

Field–Welsh chỉ agent đọc PDF method; DeLong full method chưa đọc; Dietterich OCR không đủ; Geirhos chỉ abstract; MLPerf master chưa pin. PAV và BOSARIS có sample-optimality scope; không theorem generalization. Seed×group EER resampling là design proposal với assumptions/coverage còn phải kiểm, không automatic valid CI.

Học liệu giữ mạch từ lý thuyết → phép đo → design → claim. Empirical questions chưa đủ evidence được ghi giới hạn, không chuyển thành training/download/rerun trong nhiệm vụ enrich này. Không chọn đề tài, GPU hay venue.
