# Rerun Comparison 2026-03-11

## Scope

Compared two states in detached worktrees on the same machine:

- current `HEAD`: `ace33e7` on `codex/per-group-adamw-variants`
- recovered combo: `ba60543` on `codex/wte-plain-adamw`

## Commands

- `uv run train.py --smoke-test`
- `uv run train.py`

Logs:

- `results/rerun_compare/head_smoke.log`
- `results/rerun_compare/ba60543_smoke.log`
- `results/rerun_compare/head_full.log`
- `results/rerun_compare/ba60543_full.log`

## Observed Divergence

- The only behavioral code delta between `ba60543` and current `HEAD` is the WTE optimizer assignment:
  - `ba60543`: `ADAMW_VARIANT_WTE = "adamw"`
  - current `HEAD`: `ADAMW_VARIANT_WTE = "saliencew"`
- In these reruns, current `HEAD` slightly beat the recovered combo:
  - current `HEAD`: `val_bpb = 0.867286`
  - `ba60543`: `val_bpb = 0.869136`
  - rerun gap: `0.001850` in favor of current `HEAD`
- This reverses the historical ordering recorded in the repo:
  - historical current-head-equivalent (`4bc58c2`): `0.857426`
  - historical recovered combo (`ba60543`): `0.852087`
  - historical gap: `0.005339` in favor of `ba60543`

## Interpretation

- The rerun does not reproduce the historical ranking.
- Both reruns are worse than their historical logged values:
  - current `HEAD`: `+0.009860` worse than historical
  - `ba60543`: `+0.017049` worse than historical
- Runtime and memory are nearly identical. The meaningful live difference is the optimizer telemetry:
  - current `HEAD` had a higher AdamW gate mean (`1.0320` vs `0.9873`)
  - current `HEAD` also showed higher promotive and conflict means, while `ba60543` was more strongly aversive

## Practical Read

- We can reconstruct the intended recovered combo exactly.
- We cannot claim from today's reruns that it is still the best state.
- As of these reruns, WTE-on-saliencew slightly outperformed WTE-on-plain-AdamW on this machine, but only by `0.001850`, which is small enough that I would not treat it as settled without repeat runs.
