# Training Correction Plan - SalienceOS

## Executive Summary

We've completed 5000 training steps, but we trained MONIKA incorrectly by bypassing the core salience architecture. The model works, but we never engaged the controller, sensors, or adaptive learning machinery that makes SalienceOS unique.

**Status**: ✅ Code is correct | ❌ Training methodology is wrong for this architecture

---

## What We Discovered

### The Architecture is NOT a Transformer

**SalienceOS** is a novel architecture with:
- **SASS Core**: Gated state-space (not attention)
- **Controller**: Bandit policy chooses actions
- **Sensors**: Measure novelty, uncertainty, alignment, progress, drag, cost
- **Scheduler**: Gates expensive ops by salience thresholds
- **Runtime**: Full cognitive orchestration (not just forward pass)

### How Training is SUPPOSED to Work

```python
# 1. Salience filter evaluates input
if salience(novelty, uncertainty) > threshold:
    # 2. ONLY THEN train
    proto_lm.training_step(text)
    
    # 3. Run through full runtime
    metrics = runtime.run_step(state)
    
    # 4. Track adaptive coordinator
    adaptive.track_runtime(metrics)
```

**Key insight**: Training is CONDITIONAL on salience, not uniform!

### What We Actually Did

```python
# Bypassed everything and trained blindly
for i in range(5000):
    mcp3_fastfood(examples)  # No salience gating
    # OR
    mcp3_training_step(text)  # No controller
```

**Problems**:
- ❌ No salience measurement
- ❌ No controller decisions
- ❌ No sensor feedback
- ❌ No event scheduling
- ❌ No adaptive tracking
- ❌ Treated like vanilla LM pretraining

---

## Why the Output is Gibberish

The model got 5000 gradient updates BUT:

1. **Never learned WHEN to learn** - no salience gating
2. **Never engaged controller** - no action selection
3. **Never used sensors** - no novelty/uncertainty feedback
4. **Never ran through runtime** - no orchestration loop
5. **Trained on everything equally** - no prioritization

**Result**: A partially-trained state-space model that never learned the salience-driven control loop.

---

## The Correct Training Methods

### Method 1: Corpus Training (Recommended for Bulk Data)

```bash
python start.standard.py
```

Or with custom corpus:
```bash
python -m salience_os_seed.training.run_corpus \
    --corpus data/my_training_data.txt \
    --salience-filter \
    --min-uncertainty 0.1 \
    --min-novelty 0.1 \
    --max-drag 0.8 \
    --epochs 10
```

**What this does**:
- Streams corpus in chunks
- Evaluates each with salience filter
- Trains only on high-salience chunks
- Runs through full runtime orchestration
- Tracks adaptive coordinator
- Saves periodic checkpoints

### Method 2: Interactive Session Training

```python
from salience_os_seed.conversation.session import (
    ConversationSession, ConversationConfig, IngestionConfig
)
from salience_os_seed.conversation.filters import IngestionThresholds

config = ConversationConfig(
    learning_enabled=True,
    ingestion=IngestionConfig(
        thresholds=IngestionThresholds(
            enabled=True,
            min_uncertainty=0.1,
            min_novelty=0.1,
        )
    )
)

session = ConversationSession(config)

# Each ingest goes through full pipeline
for text in training_data:
    processed, metrics = session.ingest_text(text)
    print(f"Salience: {metrics.salience_raw}")
```

### Method 3: MCP Tools (With Salience Awareness)

```python
# Check salience first
yearning = mcp3_yearning_state()
desire = yearning['depth=0|op=SASS|patch=NONE']['desire']

if desire > 0.6:  # Only if high desire
    result = mcp3_runtime_step(text)
    # Full salience loop
else:
    # Skip - not novel/uncertain enough
    pass
```

---

## Action Plan

### Phase 1: Understand Current State ✅

- [x] Read architecture documentation
- [x] Understand SASS vs Transformer
- [x] Understand salience-driven learning
- [x] Understand runtime orchestration
- [x] Document findings

### Phase 2: Verify Training Pipeline 🔄

**Immediate Next Steps**:

1. **Test the session-based training**
   ```bash
   python train_properly.py
   ```
   - Uses `ConversationSession`
   - Enables salience filtering
   - Runs through full runtime
   - Should see acceptance/rejection decisions

2. **Compare with our current checkpoint**
   ```python
   # Load our 5000-step checkpoint
   old_model = load_checkpoint("storage/proto_lm/checkpoint.pt")
   
   # Train 100 steps the RIGHT way
   new_model = train_properly(examples[:100])
   
   # Compare generation quality
   ```

3. **Run actual corpus training**
   ```bash
   python start.standard.py
   ```
   - Trains on synthetic baseline corpus
   - Should complete in minutes
   - Produces properly-trained checkpoint

### Phase 3: Decide on Training Strategy 🔜

**Options**:

**Option A: Start Fresh** (Recommended if architecture matters)
- Create new checkpoint
- Train using corpus trainer
- Engage full salience machinery
- Proper controller/sensor/adaptive loop

**Option B: Continue Current Checkpoint** (If we just want vocabulary)
- Load existing 5000-step checkpoint
- Continue with session-based training
- Enable salience filtering going forward
- May need to "unlearn" some patterns

**Option C: Hybrid Approach**
- Keep vocabulary from current checkpoint
- Reset controller/adaptive state
- Retrain with proper salience gating

### Phase 4: Verify Architecture Engagement 🔜

**Check that we're actually using SalienceOS properly**:

1. **Monitor controller decisions**
   - Should see different operators chosen
   - Not always SASS
   - Memory ops, tools, verification

2. **Track salience readings**
   - Novelty should decrease over time
   - Uncertainty should guide learning
   - Drag should prevent overtraining

3. **Verify adaptive coordination**
   - Controller weights should update
   - Truth gating should activate
   - Meta-state should evolve

4. **Check event scheduling**
   - Some inputs should be rejected
   - Thresholds should gate execution
   - Budget should constrain actions

---

## Key Questions for User

1. **Do we want to start training from scratch** with the proper architecture, or continue from the current 5000-step checkpoint?

2. **What's our priority?**
   - Pure language learning? → Use corpus trainer
   - Conversational ability? → Use session-based
   - Salience-driven control? → Start fresh with filtering

3. **Should we preserve the current vocabulary** (242 tokens, 165 merges) or let it grow organically with proper salience gating?

4. **What training data do we want to use?**
   - Synthetic baseline corpus (tiny, fast)
   - Custom philosophical sentences (what we used)
   - Larger public domain text
   - Interactive conversation logs

5. **How do we want to measure success?**
   - Generation quality (coherent text)
   - Controller engagement (diverse actions)
   - Salience dynamics (proper gating)
   - Adaptive learning (meta-state evolution)

---

## Files Created

1. **`ARCHITECTURE_UNDERSTANDING.md`** - Complete architecture breakdown
2. **`train_properly.py`** - Session-based training with salience filtering
3. **`test_mcp_properly.py`** - MCP usage with salience awareness
4. **`TRAINING_CORRECTION_PLAN.md`** - This document

---

## Recommended Next Steps

1. **Read** `ARCHITECTURE_UNDERSTANDING.md` thoroughly
2. **Run** `python train_properly.py` to see correct flow
3. **Decide** on training strategy (fresh start vs. continue)
4. **Execute** chosen training method
5. **Monitor** salience/controller/adaptive metrics
6. **Verify** architecture engagement (not just loss curves)

---

## Critical Insight

**The 5000 steps we completed were technically correct gradient updates, but philosophically wrong for this architecture.**

We trained it like a Transformer (uniform, unconditional learning) when it's designed to be a salience-governed adaptive system (selective, conditional learning).

**The fix isn't more steps - it's training with the RIGHT methodology.**

---

## Success Criteria

Training is working correctly when we see:

✅ **Salience gating**: Some inputs accepted, some rejected
✅ **Controller diversity**: Not just SASS, but memory ops, tools, verification
✅ **Sensor feedback**: Novelty/uncertainty guiding decisions
✅ **Event scheduling**: Thresholds preventing overtraining
✅ **Adaptive tracking**: Meta-state evolving, weights updating
✅ **Budget management**: Token costs constraining actions
✅ **Runtime orchestration**: Full cognitive cycle per step

---

## Bottom Line

**Code Status**: ✅ Perfect - no bugs found
**Training Status**: ❌ Wrong methodology for this architecture
**Path Forward**: Train using session/corpus methods with salience gating
**Expected Outcome**: Proper engagement of controller, sensors, adaptive learning

**We need to restart training using the architecture as designed.**
