EgoHand4D · Technical Report

Campaign Status · Key Paths

7. Campaign Status & Next Steps

The work is now run as an evidence-driven campaign: every line of attack is opened against a measured weakness, gated by the GT instrument (§4.3), and closed — or reported as a negative — with numbers.

Scoreboard: 3 promotions (v7, v11, v14d — current champion, promoted 2026-07-05 by user visual verdict from the seed sweep), 5 evidence-based rejections (v10 rolled back, v12/v13/v15/v17b stopped pre-promotion), 4 honest negatives banked (the synthetic-foreshortening fine-tune data attempt, the C4 HaMeR fine-tune, the v13 attribution result, and the M0 memory null), and a benchmark hardened to GT + strongest-baseline + industry-baseline comparisons (frozen GT-2D/GT-3D instruments, mandatory wild-stability gate, Dyn-HaMR from its production-hardened fork with pinned exact K and a corrected harness, MediaPipe scored on both 2D instruments — §5.5).

M1/M2 — the direct-MANO head: sequence-native 2.5D from the trunk, and it beats the crop pipeline on lab GT. Following the roadmap theses (sequence-first; 2D supervises 2.5D), a flag-gated 4.8M head on the FROZEN Stage-1 temporal trunk predicts per-frame MANO rot6d (+ clip-constant betas — GT betas are constant within a video) trained with HaMeR-ratio keypoint-dominant losses (3D joints : 2D reprojection through a weak-persp training camera : weak param regularizers). Results on the shared 6 GT windows (root-aligned): RA 10.9 / PA 5.8 mm vs the HaMeR pipeline's 16.2 / 7.8 — 33% / 26% better, with lab val held after mixing in wild 2D-reprojection self-supervision (M2). Two honest boundaries: (1) these windows are lab-domain; (2) the head's weak-persp camera is a training scaffold — its wild overlay misplaces hands, so the production form (M3, in progress) pairs head gesture with the existing PnP translation against Stage-1 2D. If M3 holds on wild, HaMeR is demoted from runtime component to teacher and the second ViT-H forward disappears (~0.38 → ~0.15-0.2 s/frame). Discovered debt: the synth renderer discards MANO params at render time — re-rendering with params saved unlocks unlimited free 3D supervision (roadmap).

v14d promoted (2026-07-05) — the seed sweep pays out. The champion-recipe re-roll with the best presence-safe data order (v14b/c/d/e sweep) was promoted after the full two-leg gauntlet AND the user's visual verdict on the lineage side-by-sides: frontier tips +2.0pt / PCK05 +0.9 / lab mean5 .6626 / HInt external recall +2.7 / GT-3D parity / guitar-R 100% / wild present .989 — with the honest, user-acknowledged limitations recorded in the registry: 3D absolute accuracy (the ARCTIC ×1.17 depth residual → v18 z-refine) and occasional dropped frames / flicker on hard wild footage. Also this cycle: v17b rejected (replay_v8m synth in the 2D mix — best-ever wild stability but presence spec .865: truncation/glove synth teaches claiming hands from fragments; the fix is present-label handling for truncated clips, and v8m remains the MANO head's training set where it measurably helped), and v17 (Ego-Exo4D mix) banked as a component: best-ever HInt external recall (.800) and wild jitter (18.05), accuracy within noise — slated for the combined v17c recipe.

v15 (pre-residual local head) — the fourth rejection, and it forced the campaign's biggest methodology upgrade. v15 implemented the loss-level-coupling fix the v14 control pointed to: primary losses on the pure global outputs (verified bit-identical trunk gradients to a no-local-branch model), auxiliary loss training only the residual corrector. Result: lab mean5 +0.4pt but tips only +0.25pt, frontier PCK@0.05 −1pt, presence spec 0.935 vs 0.994, and a

wild-stability gate failure (jitter p95 +17.7%). The instructive part: stripping the local head at eval changes nothing — the degradation lives in the trunk weights despite provably identical per-batch gradients. The only remaining channel is data order: fresh-module weight init consumes RNG, so every local-head run trained on a differently-shuffled stream.

The seed controls (v14b/v14c) — measuring the gates' own noise. We re-ran the exact champion recipe with only the RNG stream perturbed (--rng-burn), no architecture change. The result rewrites how every past and future single-run comparison must be read:

v11/v14 (order A)v14b (order B)v14c (order C)noisy?
frontier PCK@0.050.511/0.512~0.510.523±1.2pt
tip-PCK@0.050.3480.366±1.9pt
presence spec / false-extra0.994 / 0.0060.971 / 0.0320.971 / 0.032spec ±0.02, sfe 5×
wild jitter p9518.8318.6618.04stable ±4%
wild teleports /min53.637.545.5±40%
lab mean50.65920.6624±0.3pt

Consequences, applied retroactively and honestly: (1) run-to-run variance (±1–2 pt) is the same magnitude as most of the architectural "wins" this campaign chased — v12's +1.2 pt and v13's +0.9 pt tip gains sit inside the seed band and are hereby downgraded from "real effect" to "unresolved against noise"; the v15 rejection stands because jitter p95 is the one gate the controls show to be stable (±4%) and v15 broke it by +17.7%. (2) presence spec 0.994 was the lucky order, 0.971/0.032 is the norm — the presence block's promotion threshold now uses the measured band, and teleports/min is demoted from gate to indicator (±40% noise). (3) v14c — the same champion recipe, re-rolled — beats the champion on nearly every axis (tips +1.9, PCK+1.2, lab +0.3, jitter −4%, teleports −15%, wild-vs-MP drift +2) and is under full-gauntlet evaluation as a promotion candidate: the seed lottery is now a deliberate instrument, not an accident.

Done (this cycle):

  1. 1. GT evaluation instrument — GT-2D frontier eval (tip-PCK first-class) + GT-3D eval (27 frozen windows, real K) + SHA256-frozen protocol (§4.3), since enriched with dual-convention PA (Sim(3)/SE(3)), jerk, per-finger, continuity, temporal-mode fairness stamps, and the per-clip depth-bias probe — the instrument whose "implausible" Dyn readings led us to audit our own harness, find the proc-res-K bug, and re-run the entire baseline fairly (§5 correction). Post-correction it shows our +83 mm far bias is now the largest quality gap in the comparison, in the corrected baseline's favor (§5.3) — the roadmap item below. MediaPipe-agreement demoted from referee to drift indicator; it was our own teacher.
  2. 2. Presence recall closed — hysteresis + sustained mid-confidence latch + joint-based off-frame test (§1.6): guitar-R 0% → 100%, medical 100%, ego_wire 98%, OOD specificity held.
  3. 3. Fast-motion articulation closed — velocity-gated per-DOF smoothing + fingertip data-term upweight (§3.1): fingertip lag −24–32% (guitar 7.3 → 5.0 px), jitter unchanged (still 3.8–75× lower than Dyn-HaMR).
  4. 4. WILD teacher re-distillation — closed end-to-end. The wild set was relabeled with HaMeR-2D, handedness fixed from COCO annotation boxes (~5% wrong labels found), ~10% ambiguous frames dropped (§2.5) — and the retrain on those labels is v11_wildhamer_36k, promoted to champion as the first model through the full two-leg gauntlet: every GT-frontier axis up (PCK@0.10 0.829, tips 0.348, all five subsets), all lab sets up (mean5 0.659), calibration improved, wild-stability gate passed with improvement (§4.3 part 3).
  5. 5. v7 promoted; v10 promoted on lab GT, then rolled back on wild evidence — the first honest negative, and the strongest test the eval process had passed at the time. v7 (finger-splay synth) set the GT baselines (PCK@0.10 0.817, TIP-PCK@0.05 0.328, RA-MPJPE 18.4 mm, hand length +0.9 cm). The local-finger-head experiment (v10) beat it on every lab-GT axis at 36k (PCK@0.10 0.822, tips 0.344, truncated +0.043) and was promoted — then the user's visual check caught wild-tail regressions within hours, the A/B audit confirmed them (teleport spikes +60%, phantom hands at conf 0.96+, local_2d/audit/v10_wild_regression/), and v10 was rolled back the same day (§4.3). Net outcome of the episode: an honest negative-in-the-wild in the lineage, an anti-teleport guard in the tracker (94 → 76 teleports/min, recall kept), and a mandatory wild-stability promotion gate. The system caught and corrected its own bad promotion because the visual check and the A/B protocol worked.
  6. 6. v12 rejected pre-promotion — the second honest negative, and proof the hardened gauntlet works without a rollback. v12 (local-finger head + v11's HaMeR labels) won GT crop accuracy everywhere — tips +1.2pt, lab +0.2pt, wild under both referees — and was rejected anyway: flicker hard-gate +22%, and the freeze audit caught its −27% teleport "win" being manufactured by the anti-teleport guard (a frozen phantom held 0.67 s during ~300 px/frame motion — the exact gate-gaming the code review predicted as H1). Diagnosis: local-head gradients degrade presence specificity (v10+v12 consistent pattern, single-false-extra ~3×). Full verdict: §4.3 part 4; audit: local_2d/audit/v12_promotion/.
  7. 7. Tracking guard hardened (M4/M5) + light root-trajectory optimization default-ON (§1.6): guard before the off-frame test (flicker halved 0.201 → 0.110), conf ≥ 0.5 anchors, 0.4 s hold cap, engagement telemetry; plus the 0.2 s/clip root-only clip-level solve (teleports −21.5%, jitter p95 −19.8%, GT PCK unchanged, drum-stroke amplitude 97.7% preserved) — Dyn-HaMR's global-consistency idea absorbed at ~1/1000 cost. Campaign stability arc: jitter p95 24.4 → 18.83 px, flicker 0.20 → 0.110/s, teleports 94 → 53.6/min (−23% / −45% / −43%).
  8. 8. v13 rejected — the third evidence-based rejection, and the cleanest attribution result of the campaign (§4.3 part 5). The presence-isolated local head kept the GT tip win under full gradient isolation (tips +0.9pt, near-camera +1.4/+2.7pt — trunk co-adaptation refuted; the gain size itself was later shown to sit inside the ±1.9pt seed band, §7), passed wild stability (flicker −39%) with a clean freeze audit (guard engagement below v11's) — and still failed GT-frontier presence harder than v12 (single-false-extra 0.070 vs 0.006, phantom second hands monotone 1 → 3 → 11 across v11 → v12 → v13). Because the local path is gradient-isolated, this refutes the v10/v12 gradient-leakage diagnosis; COCO calibration (ECE 0.053, v11-level) never saw the failure — only the GT-frontier presence block did.
  9. 9. HaMeR ego-foreshortening fine-tune (C4) rejected — an honest negative with a reframing (§3.4). The head-only fine-tune fixed foreshortening on GT crops (FS subset −6.2 mm, orientation error halved, normal subset also up) but regressed the wild ego clip that motivated it (ego_wire worst-window reproj 20.5 → 29 px). Diagnosis: a crop-domain gap (trained on GT-joint boxes, deployed on palm-anchored Stage-1 boxes — any retry trains on deployment-style crops), and the ego_wire worst-window being substantially a Stage-1 2D failure that no HaMeR checkpoint can fix — reframed as a Stage-1 data frontier. The checkpoint remains opt-in via hamer_ckpt.

Running (verdict pending — will be reported either way):

  1. 1. M0 — memory-conditioned presence (SAM2-inspired): DONE, verdict NULL (rigorous). Per-hand EMA+anchor appearance tokens conditioning the present head (+3.1M params, zero-init, coords frozen — the cleanest presence-pathway isolation yet). Phantom axes did not move (single-false-extra .0063→.0063, spec .9941→.9941), and the decisive force-empty ablation showed the memory tokens were causally inert: the real side-gains (+2.3pt wild recall, a truncated-hand recovery on conductor) came from the presence-head finetune with SINGLE/TRUNC oversampling, not from memory. Root cause is structural: the confidence-gated write (conf≥0.6) guarantees memory only exists where the head is already confident — empty exactly on the phantom/absent frames it was meant to judge. The memory line continues only with a redesigned write path (anchors from confirmed track spans, occlusion/exit-heavy training, a harder-to-shortcut readout) and a mandatory no-memory control. One instructive gate artifact: honest-confidence increases can raise measured teleports by bypassing the coord-freezing latch — recalibration × heuristic interactions remain the recurring trap.
  2. 2. v14 — the attribution control: DONE, verdict CLEAN. v11 + 36k continued steps with no local head, same data and seed as v12/v13 → presence frontier identical to v11 (single-false-extra 0.0063, spec 0.9941). Continued-training drift is ruled out; the phantom inflation is loss-level coupling from the residual heads (§4.3 part 5). The local-head line is unblocked with a concrete fix: pre-residual losses for the global heads.

Next (branching on the v14 verdict):

  1. 1. Control dirty → training-side fix: continued-training runs need presence anchoring (EWC-style regularization toward the champion's presence behavior, or fresh-start training from the v7-recipe baseline instead of stacking +36k warm-starts).
  2. 2. Control clean → trainer-side true isolation: the loss-level coupling gets removed in the trainer (global heads must not see residual-absorbed errors), then the local head re-runs the gauntlet. Either way the local-head line is blocked, not dead — v13 refuted the gradient-leakage mechanism; the seed controls left the gain itself unresolved against noise (§7), and presence safety remains the blocking issue.
  3. 3. Tip-heatmap auxiliary head — still queued: the other attack on the tip frontier, independent of the residual-head line.
  4. 4. Gate the freeze metrics — guard engagement/min and max-hold duration join the wild-stability gate as hard thresholds once champion baselines settle, so a guard-manufactured stability "win" (v12's pattern) is caught by a number, not an audit.
  5. 5. Scale the temporal MANO student on §2.2 perfect-label replay data — the direct fix for the 34.6 mm underfit, then quantify Stage-3 3D MPJPE on replay clips with real 3D GT.
  6. 6. Real ego labels — the hardest frontier: hallucination when fingers are fully occluded, the truncated frame-edge reach where GT-2D says we are weakest (PCK@0.10 0.278 even after v11's gain), and — per the C4 reframing — the ego cable-clutter windows where Stage-1 2D itself scatters.

Appendix · Key Paths