# MONIKA Fix Plan - Complete Solution

## Root Cause Summary

**The model learned to repeat the last character with 100% certainty instead of learning language.**

- `"Hello"` → predicts `'o'` with p=1.0000
- Deterministic generation: `"Hello" → "Helloooooo"`
- Even overfitting on 3 words fails: `"Hello" → "Helloooooo"`

**Why**: Character-level training + small model capacity = learns simplest pattern (repetition)

---

## The Fix: Multi-Pronged Approach

### Phase 1: Architectural Improvements ⚡ (IMMEDIATE)

**1.1 Increase Model Capacity**
- Current: embed_dim=128, 2.6M params
- Target: embed_dim=256, ~10M params
- Rationale: Need capacity to memorize word patterns

**1.2 Increase Sequence Length**
- Current: 64 tokens
- Target: 128 tokens
- Rationale: Need longer context for word-level patterns

**1.3 Add Word-Level Loss Component**
- Current: Pure next-token loss
- Target: Hybrid loss (token + word boundary detection)
- Rationale: Give explicit signal for word boundaries

### Phase 2: Training Data Improvements 📚 (IMMEDIATE)

**2.1 Pre-tokenize with Word Boundaries**
- Add explicit `<WORD_START>` and `<WORD_END>` tokens
- Train model to recognize word structure
- Example: `<WS>Hello<WE> <WS>there<WE>`

**2.2 Start with Word-Level Training**
- Phase 1: Train on single words (1000 steps)
- Phase 2: Train on word pairs (1000 steps)
- Phase 3: Train on simple sentences (5000 steps)
- Phase 4: Full corpus

**2.3 Balance Training Examples**
- Ensure NO over-representation of character repetition
- Filter out data with excessive doubled letters
- Add counter-examples (diverse character sequences)

### Phase 3: Vocabulary Strategy 🔤 (IMMEDIATE)

**3.1 Pre-seed Common Words**
- Add top 100 common words directly to vocabulary
- Don't rely on BPE to discover them
- Example: "the", "is", "a", "to", "Hello", "thank", "you"

**3.2 Force Whole-Word Merges**
- Bias merge scoring toward complete words
- Penalize mid-word merges
- Encourage word-boundary-respecting merges

**3.3 Increase Vocabulary Growth Rate**
- Current: Every 100 steps
- Target: Every 20 steps initially, then decrease
- Rationale: Need faster vocabulary discovery

### Phase 4: Loss Function Improvements 🎯 (MODERATE PRIORITY)

**4.1 Anti-Repetition Loss Term**
```python
# Penalize predicting same token repeatedly
repetition_penalty = -log(P(token[i] == token[i-1]))
total_loss = prediction_loss + 0.1 * repetition_penalty
```

**4.2 Word-Completion Reward**
```python
# Bonus for completing words correctly
if predicted_sequence in vocabulary_words:
    total_loss -= 0.5  # Reward
```

**4.3 Entropy Regularization**
```python
# Penalize zero-entropy predictions
entropy_penalty = -entropy(prediction_distribution)
total_loss += 0.05 * entropy_penalty
```

### Phase 5: Generation Improvements 🎲 (IMMEDIATE)

**5.1 Add Temperature Annealing**
- Start with temp=1.0
- Decrease to 0.7 after first few tokens
- Prevent early collapse to repetition

**5.2 Stronger Repetition Penalty**
- Current: 1.8-2.0
- Target: 3.0-5.0 for same token
- Add position-aware penalty (stronger for recent tokens)

**5.3 Look-ahead Filtering**
```python
# Reject candidates that would cause immediate repetition
if next_token == last_token and last_token == second_last_token:
    mask_out_token(next_token)
```

---

## Implementation Priority

### 🔴 CRITICAL (Do First)
1. Increase embed_dim to 256
2. Pre-seed vocabulary with common words
3. Implement anti-repetition in generation (look-ahead filtering)
4. Train on word-level data (single words → pairs → sentences)

### 🟡 HIGH PRIORITY (Do Next)
5. Add word boundary tokens
6. Increase vocabulary growth rate
7. Implement anti-repetition loss term
8. Increase sequence length to 128

### 🟢 MEDIUM PRIORITY (Nice to Have)
9. Word-completion reward
10. Entropy regularization
11. Hybrid training schedule

---

## Quick Win Test: Minimal Viable Fix

**Goal**: Get SOME coherent generation in 1 hour

**Approach**:
1. Create fresh model with embed_dim=256
2. Pre-seed vocab with 50 common words
3. Train ONLY on those 50 words (10K repetitions)
4. Use aggressive anti-repetition in generation
5. Test if it can generate ANY of the 50 words correctly

**Success Criteria**:
- Can generate at least 5/50 words correctly
- No infinite character repetition
- Some word boundaries respected

---

## Full Solution Test: Production Ready

**Goal**: Conversational ability

**Approach**:
1. All Phase 1-3 improvements
2. Train on synthetic baseline (10K steps)
3. Add word boundary training
4. Progressive curriculum (words → pairs → sentences)

**Success Criteria**:
- Can complete common phrases: "Hello" → " there"
- Understands word boundaries
- No repetition collapse
- 80%+ accuracy on validation set

---

## Files to Create/Modify

### New Files
1. `train_with_fixes.py` - Implement all improvements
2. `improved_config.py` - New TrainingConfig with fixes
3. `word_level_curriculum.py` - Progressive training data
4. `anti_repetition_loss.py` - Custom loss function

### Modified Files
1. `_torch_impl.py` - Add anti-repetition to sample()
2. `vocab.py` - Pre-seed common words
3. `trainer.py` - Hybrid loss function

---

## Estimated Timeline

- **Quick Win**: 1 hour (minimal viable fix)
- **Full Solution**: 4-6 hours (all improvements)
- **Testing & Iteration**: 2-4 hours
- **Total**: 7-11 hours to production-ready

---

## Risk Mitigation

**Risk 1**: Increased model size causes OOM
- **Mitigation**: Start with embed_dim=192, not 256
- **Fallback**: Reduce batch size

**Risk 2**: Anti-repetition breaks valid patterns
- **Mitigation**: Make penalty position-aware
- **Fallback**: Lower penalty coefficient

**Risk 3**: Word-seeding doesn't help
- **Mitigation**: Test with just 10 words first
- **Fallback**: Focus on architectural fixes instead

---

## Next Actions

1. **Implement Quick Win** (recommended to validate approach)
2. **Run comprehensive fix** (if quick win succeeds)
3. **Iterate based on results**

Which would you like me to start with?
