# Bài 9 — Truy vết từ audio, checkpoint và score tới con số

[Trước](08-SO-SANH-ABLATION-CO-CHE.md) · [Tiếp](10-EFFICIENCY-CLAIM.md).

## Mục tiêu và tiền đề

Xây evidence chain của một table cell, kiểm join IDs/labels và phân biệt reproduced metric với reproduced algorithm/result. Cần [score/sweep](02-RANKING-ROC-PR-EER.md), [split/selection](06-SPLIT-LEAKAGE-SELECTION.md), [recipe/checkpoint scope](../lop-03-ssl-jepa/14-HO-SO-RECIPE.md).

## 1. Một hàng kết quả là sản phẩm nhiều transformations

~~~mermaid
flowchart LR
 A[Audio release và protocol IDs] --> B[Decode resample crop pad]
 C[Checkpoint hash branch config] --> D[Encoder pooling head]
 B --> D
 E[Train selection manifest] --> C
 D --> F[Raw score theo ID và polarity]
 A --> G[Join ID label group]
 F --> G
 G --> H[Metric code threshold aggregation]
 H --> I[Table mean SD rounding]
~~~

Manifest không chỉ file path. Ghi dataset/release/protocol digest, audio ancestry/crop valid-length, checkpoint SHA/branch/position geometry, code/config/dependencies, train subset/seed/randomization/exposures, checkpoint selection history, calibration/threshold source, per-ID raw score và metric version/conventions. Hash chứng minh byte identity, không semantic correctness/disjointness. [Pineau et al.2021 §5/Appendix checklist](https://www.jmlr.org/papers/volume22/20-303/20-303.pdf) thúc đẩy metric/error-bars/central tendency reporting; chain và schema này là engineering synthesis của học liệu, không checklist certification.

## 2. Worked example: join sai làm AUC đẹp hơn

**Toy canonical labels/scores**:

| ID | Label | True score |
|---|---|---:|
| u1 | B | 3 |
| u2 | S | 2 |
| u3 | B | 1 |
| u4 | S | 0 |

Correct AUC: B3 thắng 2 S; B1 thắngS 0 nhưng thuaS 2 →3/4. Score file có order(u2,u4,u1,u3) và scores(2,0,3,1). Nếu positional zip với protocol labels(B,S,B,S), nhận Bscores(2,3),Sscores(0,1) →AUC 1. File có đủ 4 rows, score shapes khớp, không NaN, vẫn sai metric vì join.

Correct procedure: keyed join ID; assert unique score IDs và unique protocol IDs; assert exact expected sets (extra/missing policies explicit); compare labels from canonical protocol; retain groups and deterministic ordering; validate finite values/polarity; compute metric. Không allow duplicated score ID u2 silently overwrite; duplicate **bootstrap draw weights** khác duplicated original score rows. Với SASV trial, key có thể gồm enrollmentID+testID+trialID, không utteranceID alone.

Kiểm score-label correlation có thể phát hiện anomaly, nhưng không đảo dấu cho đẹp số khi protocol mapping chưa đúng. Filename hash khớp không chứng minh labels aligned. Một “metric reproduced to 0,01” dùng cùng sai labels ở cả hai paths cũng không độc lập xác nhận semantics.

## 3. Bốn mức reproduce khác nhau

| Evidence | Hỗ trợ | Chưa hỗ trợ |
|---|---|---|
| Vẽ lại figure từ table | Plot/arithmetic/rounding | Raw-score identity/model behavior |
| Recompute EER từ raw scores+protocol | Metric/list join theo version đó | Audio frontend/training reproducibility |
| Load checkpoint+direct/wrapper same waveform | Input/output wrapper agreement trong cases đã kiểm | Every audio/domain/length path hoặc right pretrained branch |
| Retrain algorithm từ recipe qua runs | Procedure reproducibility trong scope software/data/hardware | Broad generator/generalization hoặc mechanism |

Shape load thành công là necessary compatibility ở tensor sizes, chưa correct position axes/checkpoint branch. Wrapper/direct score agreement $10^{-3}$ trên waveform có nghĩa hai paths gần nhau cho input ấy; phải có shared preprocessor/exact shape/state/crop checks. Hai paths share cùng bug có thể vẫn match. Baseline re-run gần published number là system pipeline evidence hữu ích; model-specific wrappers/new geometry vẫn cần riêng tests. Baseline lệch là tín hiệu điều tra version/preprocess/label/weight/metric, chưa proof protocol fault.

## 4. Paper và unresolved provenance

PDF §4.3 reports: recompute released scores/labels, AASIST re-run và direct-wrapper comparison. Những gates có purposes khác nhau như bảng trên. Lớp này **đọc report**, không tự chạy lại các gates. PDF xác nhận context encoder, prefix/crop, weightedCE và final 6 epoch; exact released pretraining run recipe/checkpoint lineage vẫn unresolved ở lớp 3. [Chương 06 §§2–4](../06-GIAI-PHAU-PAPER.md).

Notes v38–v41 có thể useful history; notebook không outputs không thành evidence runs đã hoàn tất trong lượt này. Raw scores/selection manifests mới giúp kiểm historical metric; không audit toàn history chỉ để học lớp 5. Source/default khác effective run config cần phân biệt. Frozen eval-mode/backend training states ảnh hưởng dropout/BN/caching; hash checkpoint alone không khóa preprocessing.

## 5. Counterexample và tự kiểm

Correctly reproduced score file có thể thuộc wrong checkpoint prefix; cùng EER không giải quyết provenance. Hai scores vectors khác ranking có thể có cùng EER; “number matches” không chứng minh ID-score match.

1. AUC 1 trong positional join toy có đủ evidence better model không?
2. SHA identical có chứng minh no near-duplicate overlap không?
3. Baseline match có certify mọi wrapper đúng không?
4. Re-render table đúng số là reproduce result hay figure?

<details>
<summary>Đáp án reasoning</summary>

1. Không; mapping labels sai. Correct keyed AUC 0,75.
2. Không; SHA xác nhận same bytes cho artifact được hash. Re-encode khác bytes vẫn same recording; absence of same hash không absence overlap.
3. Không; gate kiểm baseline path trong scope cases, new wrappers/branches/crops còn riêng.
4. Figure/arithmetic reproduction. Result-level cần data/score generation/protocol chain; algorithm-level cần training procedure/runs.

</details>

**Đào sâu:** tự viết minimal manifest schema cho một Table 4 cell, thêm selection log và roles để một người khác biết data nào đã được nhìn. Không lấy file path thay checkpoint content hash, không lấy hash thay mechanism proof.
