Campaign Status · Key Paths
7. Campaign Status & Next Steps
The work is now run as an evidence-driven campaign: every line of attack is opened against a measured weakness, gated by the GT instrument (§4.3), and closed — or reported as a negative — with numbers.
Scoreboard: 3 promotions (v7, v11, v14d — current champion, promoted 2026-07-05 by user visual verdict from the seed sweep), 5 evidence-based rejections (v10 rolled back, v12/v13/v15/v17b stopped pre-promotion), 4 honest negatives banked (the synthetic-foreshortening fine-tune data attempt, the C4 HaMeR fine-tune, the v13 attribution result, and the M0 memory null), and a benchmark hardened to GT + strongest-baseline + industry-baseline comparisons (frozen GT-2D/GT-3D instruments, mandatory wild-stability gate, Dyn-HaMR from its production-hardened fork with pinned exact K and a corrected harness, MediaPipe scored on both 2D instruments — §5.5).
M1/M2 — the direct-MANO head: sequence-native 2.5D from the trunk, and it beats the crop pipeline on lab GT. Following the roadmap theses (sequence-first; 2D supervises 2.5D), a flag-gated 4.8M head on the FROZEN Stage-1 temporal trunk predicts per-frame MANO rot6d (+ clip-constant betas — GT betas are constant within a video) trained with HaMeR-ratio keypoint-dominant losses (3D joints : 2D reprojection through a weak-persp training camera : weak param regularizers). Results on the shared 6 GT windows (root-aligned): RA 10.9 / PA 5.8 mm vs the HaMeR pipeline's 16.2 / 7.8 — 33% / 26% better, with lab val held after mixing in wild 2D-reprojection self-supervision (M2). Two honest boundaries: (1) these windows are lab-domain; (2) the head's weak-persp camera is a training scaffold — its wild overlay misplaces hands, so the production form (M3, in progress) pairs head gesture with the existing PnP translation against Stage-1 2D. If M3 holds on wild, HaMeR is demoted from runtime component to teacher and the second ViT-H forward disappears (~0.38 → ~0.15-0.2 s/frame). Discovered debt: the synth renderer discards MANO params at render time — re-rendering with params saved unlocks unlimited free 3D supervision (roadmap).
v14d promoted (2026-07-05) — the seed sweep pays out. The champion-recipe re-roll with the best presence-safe data order (v14b/c/d/e sweep) was promoted after the full two-leg gauntlet AND the user's visual verdict on the lineage side-by-sides: frontier tips +2.0pt / PCK05 +0.9 / lab mean5 .6626 / HInt external recall +2.7 / GT-3D parity / guitar-R 100% / wild present .989 — with the honest, user-acknowledged limitations recorded in the registry: 3D absolute accuracy (the ARCTIC ×1.17 depth residual → v18 z-refine) and occasional dropped frames / flicker on hard wild footage. Also this cycle: v17b rejected (replay_v8m synth in the 2D mix — best-ever wild stability but presence spec .865: truncation/glove synth teaches claiming hands from fragments; the fix is present-label handling for truncated clips, and v8m remains the MANO head's training set where it measurably helped), and v17 (Ego-Exo4D mix) banked as a component: best-ever HInt external recall (.800) and wild jitter (18.05), accuracy within noise — slated for the combined v17c recipe.
v15 (pre-residual local head) — the fourth rejection, and it forced the campaign's biggest methodology upgrade. v15 implemented the loss-level-coupling fix the v14 control pointed to: primary losses on the pure global outputs (verified bit-identical trunk gradients to a no-local-branch model), auxiliary loss training only the residual corrector. Result: lab mean5 +0.4pt but tips only +0.25pt, frontier PCK@0.05 −1pt, presence spec 0.935 vs 0.994, and a
wild-stability gate failure (jitter p95 +17.7%). The instructive part: stripping the local head at eval changes nothing — the degradation lives in the trunk weights despite provably identical per-batch gradients. The only remaining channel is data order: fresh-module weight init consumes RNG, so every local-head run trained on a differently-shuffled stream.
The seed controls (v14b/v14c) — measuring the gates' own noise. We re-ran the exact champion recipe with only the RNG stream perturbed (--rng-burn), no architecture change. The result rewrites how every past and future single-run comparison must be read:
| v11/v14 (order A) | v14b (order B) | v14c (order C) | noisy? | |
|---|---|---|---|---|
| frontier PCK@0.05 | 0.511/0.512 | ~0.51 | 0.523 | ±1.2pt |
| tip-PCK@0.05 | 0.348 | — | 0.366 | ±1.9pt |
| presence spec / false-extra | 0.994 / 0.006 | 0.971 / 0.032 | 0.971 / 0.032 | spec ±0.02, sfe 5× |
| wild jitter p95 | 18.83 | 18.66 | 18.04 | stable ±4% |
| wild teleports /min | 53.6 | 37.5 | 45.5 | ±40% |
| lab mean5 | 0.6592 | — | 0.6624 | ±0.3pt |
Consequences, applied retroactively and honestly: (1) run-to-run variance (±1–2 pt) is the same magnitude as most of the architectural "wins" this campaign chased — v12's +1.2 pt and v13's +0.9 pt tip gains sit inside the seed band and are hereby downgraded from "real effect" to "unresolved against noise"; the v15 rejection stands because jitter p95 is the one gate the controls show to be stable (±4%) and v15 broke it by +17.7%. (2) presence spec 0.994 was the lucky order, 0.971/0.032 is the norm — the presence block's promotion threshold now uses the measured band, and teleports/min is demoted from gate to indicator (±40% noise). (3) v14c — the same champion recipe, re-rolled — beats the champion on nearly every axis (tips +1.9, PCK+1.2, lab +0.3, jitter −4%, teleports −15%, wild-vs-MP drift +2) and is under full-gauntlet evaluation as a promotion candidate: the seed lottery is now a deliberate instrument, not an accident.
Done (this cycle):
- 1. GT evaluation instrument — GT-2D frontier eval (tip-PCK first-class) + GT-3D eval (27 frozen windows, real K) + SHA256-frozen protocol (§4.3), since enriched with dual-convention PA (Sim(3)/SE(3)), jerk, per-finger, continuity, temporal-mode fairness stamps, and the per-clip depth-bias probe — the instrument whose "implausible" Dyn readings led us to audit our own harness, find the proc-res-K bug, and re-run the entire baseline fairly (§5 correction). Post-correction it shows our +83 mm far bias is now the largest quality gap in the comparison, in the corrected baseline's favor (§5.3) — the roadmap item below. MediaPipe-agreement demoted from referee to drift indicator; it was our own teacher.
- 2. Presence recall closed — hysteresis + sustained mid-confidence latch + joint-based off-frame test (§1.6): guitar-R 0% → 100%, medical 100%, ego_wire 98%, OOD specificity held.
- 3. Fast-motion articulation closed — velocity-gated per-DOF smoothing + fingertip data-term upweight (§3.1): fingertip lag −24–32% (guitar 7.3 → 5.0 px), jitter unchanged (still 3.8–75× lower than Dyn-HaMR).
- 4. WILD teacher re-distillation — closed end-to-end. The wild set was relabeled with HaMeR-2D, handedness fixed from COCO annotation boxes (~5% wrong labels found), ~10% ambiguous frames dropped (§2.5) — and the retrain on those labels is v11_wildhamer_36k, promoted to champion as the first model through the full two-leg gauntlet: every GT-frontier axis up (PCK@0.10 0.829, tips 0.348, all five subsets), all lab sets up (mean5 0.659), calibration improved, wild-stability gate passed with improvement (§4.3 part 3).
- 5. v7 promoted; v10 promoted on lab GT, then rolled back on wild evidence — the first honest negative, and the strongest test the eval process had passed at the time. v7 (finger-splay synth) set the GT baselines (PCK@0.10 0.817, TIP-PCK@0.05 0.328, RA-MPJPE 18.4 mm, hand length +0.9 cm). The local-finger-head experiment (v10) beat it on every lab-GT axis at 36k (PCK@0.10 0.822, tips 0.344, truncated +0.043) and was promoted — then the user's visual check caught wild-tail regressions within hours, the A/B audit confirmed them (teleport spikes +60%, phantom hands at conf 0.96+,
local_2d/audit/v10_wild_regression/), and v10 was rolled back the same day (§4.3). Net outcome of the episode: an honest negative-in-the-wild in the lineage, an anti-teleport guard in the tracker (94 → 76 teleports/min, recall kept), and a mandatory wild-stability promotion gate. The system caught and corrected its own bad promotion because the visual check and the A/B protocol worked. - 6. v12 rejected pre-promotion — the second honest negative, and proof the hardened gauntlet works without a rollback. v12 (local-finger head + v11's HaMeR labels) won GT crop accuracy everywhere — tips +1.2pt, lab +0.2pt, wild under both referees — and was rejected anyway: flicker hard-gate +22%, and the freeze audit caught its −27% teleport "win" being manufactured by the anti-teleport guard (a frozen phantom held 0.67 s during ~300 px/frame motion — the exact gate-gaming the code review predicted as H1). Diagnosis: local-head gradients degrade presence specificity (v10+v12 consistent pattern, single-false-extra ~3×). Full verdict: §4.3 part 4; audit:
local_2d/audit/v12_promotion/. - 7. Tracking guard hardened (M4/M5) + light root-trajectory optimization default-ON (§1.6): guard before the off-frame test (flicker halved 0.201 → 0.110), conf ≥ 0.5 anchors, 0.4 s hold cap, engagement telemetry; plus the 0.2 s/clip root-only clip-level solve (teleports −21.5%, jitter p95 −19.8%, GT PCK unchanged, drum-stroke amplitude 97.7% preserved) — Dyn-HaMR's global-consistency idea absorbed at ~1/1000 cost. Campaign stability arc: jitter p95 24.4 → 18.83 px, flicker 0.20 → 0.110/s, teleports 94 → 53.6/min (−23% / −45% / −43%).
- 8. v13 rejected — the third evidence-based rejection, and the cleanest attribution result of the campaign (§4.3 part 5). The presence-isolated local head kept the GT tip win under full gradient isolation (tips +0.9pt, near-camera +1.4/+2.7pt — trunk co-adaptation refuted; the gain size itself was later shown to sit inside the ±1.9pt seed band, §7), passed wild stability (flicker −39%) with a clean freeze audit (guard engagement below v11's) — and still failed GT-frontier presence harder than v12 (single-false-extra 0.070 vs 0.006, phantom second hands monotone 1 → 3 → 11 across v11 → v12 → v13). Because the local path is gradient-isolated, this refutes the v10/v12 gradient-leakage diagnosis; COCO calibration (ECE 0.053, v11-level) never saw the failure — only the GT-frontier presence block did.
- 9. HaMeR ego-foreshortening fine-tune (C4) rejected — an honest negative with a reframing (§3.4). The head-only fine-tune fixed foreshortening on GT crops (FS subset −6.2 mm, orientation error halved, normal subset also up) but regressed the wild ego clip that motivated it (ego_wire worst-window reproj 20.5 → 29 px). Diagnosis: a crop-domain gap (trained on GT-joint boxes, deployed on palm-anchored Stage-1 boxes — any retry trains on deployment-style crops), and the ego_wire worst-window being substantially a Stage-1 2D failure that no HaMeR checkpoint can fix — reframed as a Stage-1 data frontier. The checkpoint remains opt-in via
hamer_ckpt.
Running (verdict pending — will be reported either way):
- 1. M0 — memory-conditioned presence (SAM2-inspired): DONE, verdict NULL (rigorous). Per-hand EMA+anchor appearance tokens conditioning the present head (+3.1M params, zero-init, coords frozen — the cleanest presence-pathway isolation yet). Phantom axes did not move (single-false-extra .0063→.0063, spec .9941→.9941), and the decisive force-empty ablation showed the memory tokens were causally inert: the real side-gains (+2.3pt wild recall, a truncated-hand recovery on conductor) came from the presence-head finetune with SINGLE/TRUNC oversampling, not from memory. Root cause is structural: the confidence-gated write (conf≥0.6) guarantees memory only exists where the head is already confident — empty exactly on the phantom/absent frames it was meant to judge. The memory line continues only with a redesigned write path (anchors from confirmed track spans, occlusion/exit-heavy training, a harder-to-shortcut readout) and a mandatory no-memory control. One instructive gate artifact: honest-confidence increases can raise measured teleports by bypassing the coord-freezing latch — recalibration × heuristic interactions remain the recurring trap.
- 2. v14 — the attribution control: DONE, verdict CLEAN. v11 + 36k continued steps with no local head, same data and seed as v12/v13 → presence frontier identical to v11 (single-false-extra 0.0063, spec 0.9941). Continued-training drift is ruled out; the phantom inflation is loss-level coupling from the residual heads (§4.3 part 5). The local-head line is unblocked with a concrete fix: pre-residual losses for the global heads.
Next (branching on the v14 verdict):
- 1. Control dirty → training-side fix: continued-training runs need presence anchoring (EWC-style regularization toward the champion's presence behavior, or fresh-start training from the v7-recipe baseline instead of stacking +36k warm-starts).
- 2. Control clean → trainer-side true isolation: the loss-level coupling gets removed in the trainer (global heads must not see residual-absorbed errors), then the local head re-runs the gauntlet. Either way the local-head line is blocked, not dead — v13 refuted the gradient-leakage mechanism; the seed controls left the gain itself unresolved against noise (§7), and presence safety remains the blocking issue.
- 3. Tip-heatmap auxiliary head — still queued: the other attack on the tip frontier, independent of the residual-head line.
- 4. Gate the freeze metrics — guard engagement/min and max-hold duration join the wild-stability gate as hard thresholds once champion baselines settle, so a guard-manufactured stability "win" (v12's pattern) is caught by a number, not an audit.
- 5. Scale the temporal MANO student on §2.2 perfect-label replay data — the direct fix for the 34.6 mm underfit, then quantify Stage-3 3D MPJPE on replay clips with real 3D GT.
- 6. Real ego labels — the hardest frontier: hallucination when fingers are fully occluded, the truncated frame-edge reach where GT-2D says we are weakest (PCK@0.10 0.278 even after v11's gain), and — per the C4 reframing — the ego cable-clutter windows where Stage-1 2D itself scatters.
Appendix · Key Paths
- Stage-1 model
hand2d/model_bidir_vit.py(HandBidirViT) · losstrain_unified.compute_loss_unified· inferencehand2d/infer.py(persistence latch:persistent_track) - Temporal MANO student
hand2d/model_lift_temporal.py· 2D-constrained solvehand2d/lift_v2.py(solve_hand_2dfit) - GT evaluation instrument
hand2d/eval_frontier.py(GT-2D, tip-PCK + frontier subsets) ·hand2d/eval_gt3d.py(GT-3D, 27 windows, real K) · wild-stability gatehand2d/eval_wild_stability.py(6 wild clips, mandatory inpromote_best.py) · frozen protocolhand2d/_bench/protocol_v1.json(59 SHA256'd inputs) · audits: v10 rollbacklocal_2d/audit/v10_wild_regression/· v11 promotionlocal_2d/audit/v11_promotion/· v12 rejection + freeze auditlocal_2d/audit/v12_promotion/· v13 rejection + attributionlocal_2d/audit/v13_promotion/+local_2d/frontier_eval/results_v13_36k.json· C4 fine-tune negativelocal_2d/hamer_ft/evals/· stability baselineslocal_2d/audit/wild_stability/· traj-opt validationlocal_2d/audit/traj_opt2d/ - Benchmark / 4-panel renderer
hand2d/benchmark.py(velocity-gated smoothing + tip upweight + default-ON root trajectory optimization live here) · gallery generatorhand2d/gen_model_gallery.py - Synthetic
hand2d/synth_hands.py· replayhand2d/synth_replay.py· v5 occlusionhand2d/render_replay_clips_v5.py - WILD relabel
hand2d/hamer_relabel_wild.py→local_2d/wild_hamer/(HaMeR-2D teacher, COCO-box handedness, MP cross-check) - Best model registry
hand2d/best_model.json→local_2d/best/best.pt(v11_wildhamer_36k — v7 recipe + HaMeR wild teacher, first two-leg-gauntlet champion; v10 rolled back 2026-07, kept in the lineage as lab-better / wild-worse; v12 rejected pre-promotion at the wild-stability gate; v13 rejected pre-promotion at the GT-frontier presence block — attribution control v14 running) - Report generator
hand2d/gen_report_site.py(this site) · sourcelocal_2d/REPORT_stage1_2d.md