# Salience Replacement Matrix

This document is the working attack surface for replacing or salience-conditioning the major parts of the stack in this repo.

The rule is simple:

- one part at a time unless explicitly testing interactions
- keep eval `val_bpb` untouched unless the experiment is explicitly about metrics
- log the result against the current best reference
- do not merge a loser into the main line

Current stable best reference at time of writing:

- branch: `codex/no-checkpoint-5060`
- best run: `val_bpb = 0.861187`
- stack:
  - `saliencew v2` on AdamW-handled groups
  - salience gate on Muon matrix path
  - no checkpointing
  - `WARMDOWN_RATIO = 0.1`
  - `MATRIX_LR = 0.05`

## Experiment Protocol

For every part below:

1. fork from the current best branch
2. change only the named part unless the row explicitly says to pair it with another
3. run smoke first
4. if smoke is viable, run a full benchmark
5. score against `current_best_run.json`
6. keep only if it wins under the layered policy

## Part Classes

- `Replace`: remove the current math and substitute salience math directly
- `Gate`: keep the current math but let salience modulate it
- `Aux`: add a salience-driven side objective / side head
- `Partition`: move some params from one regime to another

## Attack Matrix

| Priority | Part | Current Role | Salience Insertion | Class | Risk | What To Measure |
|---|---|---|---|---|---|---|
| 1 | AdamW momentum rules | `exp_avg` on non-matrix groups | salience-dependent persistence / forgetting | Replace/Gate | Medium | `val_bpb`, per-group `P/Q/Xi`, update norm |
| 2 | AdamW second-moment rules | `exp_avg_sq` suppression | reinterpret as aversive / fatigue channel | Replace/Gate | High | `val_bpb`, stability, per-group variance stats |
| 3 | Muon LR gate | matrix-step scale | novelty/retention/fatigue gate on matrix updates | Gate | Medium | `val_bpb`, tokens/sec, Muon group stats |
| 4 | Group partitioning | params assigned to AdamW vs Muon | move selected groups to salience optimizer families | Partition | High | `val_bpb`, failure mode by group |
| 5 | Value embedding optimizer | same family as embeddings | its own salience optimizer or gate | Replace/Gate | High | `val_bpb`, VE behavior, memory/steps |
| 6 | LM head optimizer | current non-matrix path | pure salience or gated salience optimizer | Replace/Gate | High | `val_bpb`, early smoke learning curve |
| 7 | Token embedding optimizer | current non-matrix path | salience-conditioned embedding update | Replace/Gate | High | `val_bpb`, stability, drift |
| 8 | Scalar optimizer | `resid_lambdas`, `x0_lambdas` | salience-conditioned scalar dynamics | Replace/Gate | Medium | branch weighting behavior |
| 9 | Training objective weighting | plain CE on tokens | salience-weighted CE | Aux | Medium | `val_bpb`, runtime, token-weight spread |
| 10 | Auxiliary salience head | none in baseline | predict promotive/aversive token labels | Aux | Low/Medium | whether it regularizes or hurts |
| 11 | Window pattern | `SSSL` | salience decides layer window regime | Gate/Replace | High | `val_bpb`, runtime |
| 12 | Per-layer window schedule | fixed pattern repetition | salience-conditioned schedule | Gate | High | `val_bpb`, cost |
| 13 | VE gating logic | current explicit gate | promotive/aversive/conflict decomposition | Replace | High | smoke viability, `val_bpb` |
| 14 | Value embedding path | additive explicit value path | salience-routed / ablated / redesigned path | Replace | High | capacity vs simplicity tradeoff |
| 15 | Residual pathway structure | `resid_lambdas` / `x0_lambdas` | salience-controlled branch routing | Replace/Gate | High | `val_bpb`, step count |
| 16 | Q/K/V projections | current linear projections | salience-aware parametrization or routing | Replace | Very High | smoke first, runtime |
| 17 | Attention output projection | standard linear output | salience-conditioned output scale | Gate | Medium | `val_bpb`, runtime |
| 18 | Head specialization / redundancy | emergent only | explicit salience allocation across heads | Gate/Aux | High | head stats, `val_bpb` |
| 19 | Token embedding init | large std | salience-derived init geometry | Replace | High | early learning curve |
| 20 | LM head init | very small std | salience-derived init geometry | Replace | High | early learning curve |
| 21 | Value embedding init | current uniform init | salience-aligned init | Replace | High | early learning curve |
| 22 | Tied vs untied embeddings/head | untied | salience-aware tying or partial tying | Replace | High | smoke viability, `val_bpb` |
| 23 | Depth | current fixed 8 | salience-guided depth curriculum / route | Replace/Gate | Very High | params vs performance |
| 24 | Width / aspect ratio | current fixed | salience-aware capacity allocation | Replace | Very High | params vs performance |
| 25 | Head dim / number of heads / KV heads | fixed architecture | salience-aware head budget | Replace | Very High | runtime vs `val_bpb` |
| 26 | Total parameter budget | static | salience-aware budget allocation | Replace | Very High | same budget comparisons |
| 27 | Weight decay rules | mostly off on AdamW groups | aversive/fatigue-driven decay | Replace/Gate | Medium | `val_bpb`, norm drift |
| 28 | Per-group betas | hand-tuned constants | salience-conditioned timescales | Gate | Medium | stability, learning speed |
| 29 | Per-group epsilon | static numerics | salience-conditioned denominator floor | Gate | Medium | stability edge cases |
| 30 | Update clipping / normalization | implicit via optimizer math | salience-dependent clip / norm rule | Gate | Medium | spike prevention vs underupdate |
| 31 | SDPA backend behavior | fixed PyTorch path | likely not a salience target directly | Skip unless forced | Low | only change if backend changes |
| 32 | QK normalization | current `norm(q), norm(k)` | salience-conditioned normalization | Replace | High | attention stability |
| 33 | RoPE settings | default current base | salience likely orthogonal here | Low priority | Low | only if evidence appears |
| 34 | Causal mask construction | fixed | only salience-relevant if sparse routing emerges | Low priority | Medium | runtime + correctness |
| 35 | Sequence packing / target encoding | plain shift | salience-coded training target / packing | Replace/Aux | High | objective path only |

## Immediate Ordered Plan

These are the parts that currently deserve attention first.

### Wave 1: Optimizer Core

1. AdamW second-moment rules
2. AdamW group partitioning
3. Value embedding optimizer
4. LM head optimizer
5. Scalar optimizer
6. Muon salience gate tuning

Reason:
- we already have proof that salience-conditioned optimizer math can help
- this is where the strongest empirical traction has been

### Wave 2: Objective / Encoding

1. salience-weighted CE
2. auxiliary salience head
3. token target construction
4. document / sequence weighting

Reason:
- the objective path is not winning yet, but it is conceptually aligned and worth refining

### Wave 3: Architecture

1. VE gating logic
2. value embedding path
3. residual pathway structure
4. head specialization / redundancy

Reason:
- these are expensive and easy to make worse
- do them only after optimizer/objective leverage slows down

## Per-Part Current Notes

### AdamW / SalienceW

Known:
- pure AdamW best was worse than current saliencew stack
- `saliencew v2` is a validated winner on non-matrix groups

Open:
- which non-matrix groups actually contribute the gain
- whether second-moment math is the real locus rather than final-step gating

### Muon

Known:
- a mild salience gate on Muon improved the best run
- pushing the novelty term harder regressed slightly
- Muon groups showed:
  - high novelty
  - negative retention
  - tiny second moment on large matrix groups

Open:
- whether salience should gate Muon LR, momentum, or variance normalization
- whether negative retention is signal or artifact

### Objective Path

Known:
- first objective versions hurt
- weakening the objective moved it closer to viability but still lost to the optimizer winner

Open:
- whether salience should supervise token weights, token classes, spans, or documents
- whether the current `P/Q` target construction is simply wrong

## Logging Requirements

Every experiment touching a part in this matrix should record:

- part name
- replacement class (`Replace`, `Gate`, `Aux`, `Partition`)
- current reference commit
- smoke result
- full result
- whether the part was isolated or combined

Recommended `results.tsv` description format:

```text
human: part=<name> class=<class> note=<short note>
```

Examples:

```text
human: part=adamw_second_moment class=Replace note=salience fatigue denominator
human: part=muon_lr_gate class=Gate note=novelty-biased gate
human: part=salience_objective class=Aux note=weak PQ token regularizer
```

## Stop Conditions

Stop pushing a part family when:

- three isolated variants all lose cleanly
- smoke repeatedly fails before full runs
- regressions are large and structural rather than tunable

At that point, move down the matrix instead of forcing it.
