# MK3 Critical Issues - Action Items

## Status: ⛔ NOT PRODUCTION READY - Critical Fixes Required

---

## 🔴 CRITICAL (Must Fix Immediately)

### Issue #1: Catastrophic Performance Bottleneck in Salience Attention

**File:** `core/salience_layers.py`
**Lines:** 107-168
**Function:** `SalienceAttentionLayer.forward()`

**Problem:**
```python
# Lines 107-144: Nested loops instead of batched operations
for head_idx in range(self.num_heads):  # Loop over heads
    for q_idx in range(batch_size * seq_len_q):  # Loop over positions
        # Compute salience for each position individually
```

**Impact:**
- For batch=4, seq_len=128, heads=8: **4,096 sequential iterations**
- Makes attention 100-1000x slower than necessary
- **Completely unusable** for any realistic sequence length

**Fix:**
```python
# Replace lines 107-168 with batched computation:
# 1. Expand Q and K to [batch, heads, seq_q, seq_k, head_dim] using broadcasting
# 2. Compute all salience scores in parallel
# 3. Use torch.matmul and broadcasting instead of loops
```

**Estimated Effort:** 1-2 days
**Priority:** ⛔ **IMMEDIATE** (blocks all usage)

---

## 🟠 HIGH PRIORITY (Fix Soon)

### Issue #2: Inefficient Token-Vector Conversion

**File:** `calm/continuous_model.py`
**Lines:** 206-223 (tokenize_to_vectors), 242-257 (vectors_to_tokens)

**Problem:**
```python
# Lines 206-212: Sequential encoding
for i in range(num_chunks):
    chunk = token_chunks[:, i, :]  # [batch, chunk_size]
    vector = self.autoencoder.encode(chunk)  # Sequential!
    continuous_vectors.append(vector)
```

**Impact:**
- 64 sequential operations for 512-token sequence
- Misses GPU parallelization
- 10-50x slower than batched approach

**Fix:**
```python
# Batch all chunks at once:
# 1. Reshape token_chunks to [batch * num_chunks, chunk_size]
# 2. Single forward pass: all_vectors = self.autoencoder.encode(all_chunks)
# 3. Reshape back to [batch, num_chunks, vector_dim]
```

**Estimated Effort:** 4-8 hours
**Priority:** 🟠 **HIGH** (major performance impact)

---

### Issue #3: Autoencoder Training Data Iteration Bug

**File:** `calm/likelihood_free.py`
**Lines:** 311-344
**Function:** `LikelihoodFreeTrainer.train_autoencoder()`

**Problem:**
```python
# Line 299: Method signature
def train_autoencoder(self, token_ids: torch.Tensor, num_steps: int = 1000, ...):
    # Lines 311-344: Samples from same tensor repeatedly
    for step in range(num_steps):
        chunk = token_ids[:, start_idx:start_idx + chunk_size]
        # Always uses the same token_ids tensor!
```

**Impact:**
- Overfits to single batch if token_ids is small
- Can't iterate over full dataset
- Training quality degraded

**Fix:**
```python
# Change signature to accept data loader:
def train_autoencoder(self, data_loader, num_steps: int = 1000, ...):
    for step, batch in enumerate(data_loader):
        if step >= num_steps:
            break
        # Use batch from data_loader
```

**Estimated Effort:** 1-2 hours
**Priority:** 🟠 **HIGH** (affects training correctness)

---

## 🟡 MEDIUM PRIORITY (Fix Before Production)

### Issue #4: Generation Sampling Not Implemented

**File:** `calm/continuous_model.py`
**Lines:** 334-380
**Function:** `ContinuousAutoregressiveModel.generate()`

**Problem:**
```python
# Lines 370-372: Temperature and top_p not actually used
if temperature != 1.0:
    next_vector = next_vector / temperature
# But no actual sampling! Just uses the predicted vector directly
```

**Impact:**
- No diversity in generation
- Can't control generation behavior
- Not true autoregressive generation

**Fix:**
```python
# After temperature scaling:
# 1. Decode vector to token logits
# 2. Apply top_p / top_k sampling
# 3. Re-encode sampled tokens to vector
# OR: Implement sampling in continuous space using noise injection
```

**Estimated Effort:** 1 day
**Priority:** 🟡 **MEDIUM** (for text generation quality)

---

### Issue #5: Memory Buffer Management Suboptimal

**File:** `calm/continuous_model.py`
**Lines:** 147-169
**Function:** `ContinuousAutoregressiveModel.update_memory_buffer()`

**Problem:**
```python
# Lines 153-154: Random sampling
num_samples = min(embeddings.shape[0], self.memory_buffer_size // 4)
indices = torch.randint(..., (num_samples,), ...)
sampled = embeddings[indices]
# Random sampling might miss important items
```

**Impact:**
- Fatigue computation might not work optimally
- Better strategies exist (importance sampling, recency weighting)

**Fix:**
```python
# Options:
# 1. Use reservoir sampling for fair sampling
# 2. Use importance scores to keep most salient items
# 3. Use recency weighting (keep recent + important old items)
```

**Estimated Effort:** 4-8 hours
**Priority:** 🟡 **MEDIUM** (performance optimization)

---

## 🔵 LOW PRIORITY (Nice to Have)

### Issue #6: No KV Caching for Generation

**File:** `calm/continuous_model.py`
**Function:** `generate()`

**Problem:**
- Each generation step recomputes attention for all previous positions
- Standard transformer optimization (KV caching) not implemented

**Impact:**
- Generation slower than necessary
- Still faster than token-by-token due to vector prediction

**Fix:**
- Implement KV caching similar to standard transformers

**Estimated Effort:** 2-3 days
**Priority:** 🔵 **LOW** (optimization)

---

### Issue #7: Missing Unit Tests

**Status:** Tests created in `tests/test_mk3_comprehensive.py` but not run yet

**Impact:**
- Can't verify correctness
- Can't catch regressions

**Fix:**
- Run tests (requires PyTorch installation)
- Add CI/CD pipeline

**Priority:** 🔵 **LOW** (but recommended)

---

## Summary of Fixes Needed

| Issue | Severity | Effort | Status |
|-------|----------|--------|--------|
| #1: Attention bottleneck | ⛔ CRITICAL | 1-2 days | ❌ Not fixed |
| #2: Token conversion | 🟠 HIGH | 4-8 hours | ❌ Not fixed |
| #3: Training loop | 🟠 HIGH | 1-2 hours | ❌ Not fixed |
| #4: Generation sampling | 🟡 MEDIUM | 1 day | ❌ Not fixed |
| #5: Memory buffer | 🟡 MEDIUM | 4-8 hours | ❌ Not fixed |
| #6: KV caching | 🔵 LOW | 2-3 days | ❌ Not fixed |
| #7: Tests | 🔵 LOW | - | ⚠️ Created, not run |

**Total Estimated Effort to Make Viable:** 1-2 weeks full-time

---

## Recommendation

**DO NOT USE FOR:**
- ❌ Production systems
- ❌ Research experiments (too slow to train)
- ❌ Benchmarking (invalid due to bugs)

**CAN USE FOR:**
- ✅ Code review and learning
- ✅ Understanding the architecture
- ✅ As a starting point for fixes

**Next Steps:**
1. Fix Issue #1 (attention) - IMMEDIATE
2. Fix Issue #2 (conversion) - HIGH
3. Fix Issue #3 (training) - HIGH
4. Validate with tests
5. Then consider production use
