# Checkpoint Loading Bug Fix

## Problem
Tensor size mismatch errors (e.g., "tensor a (320) must match tensor b (384)") occurred after loading checkpoints that had undergone vocabulary expansions. The error persisted even when loading earlier checkpoints.

## Root Cause
The `load_checkpoint` method in `_torch_impl.py` loaded model state_dict but did NOT reinitialize the SASS core or ModuleManager. These components retained stale tensor references from previous vocabulary capacity expansions, causing dimension mismatches during forward/backward passes.

## 4D Debugging Process
1. **Spatial Dimension**: Traced tensor sizes through embed → SASS → output layers
2. **Temporal Dimension**: Identified corruption persisted across checkpoint loads
3. **Causal Dimension**: Found `_ensure_capacity` correctly reinitializes but `load_checkpoint` doesn't
4. **State Dimension**: Discovered in-memory MCP server state diverged from on-disk checkpoints

## Solution
Added explicit reinitialization of SASS core and ModuleManager in `load_checkpoint` method after loading state_dict (lines 711-725):

```python
# Force reinitialization of SASS core and ModuleManager to prevent stale tensor references
# This is critical after loading checkpoints that may have undergone vocab expansions
sass_cfg = SASSConfig(
    d_model=self.config.embed_dim,
    state_channels=self.config.embed_dim * 2,
    num_layers=4,
    dropout=0.05,
)
self.core = SASSCore(sass_cfg).to(self.device)

self.module_manager = ModuleManager(
    embed_dim=self.config.embed_dim,
    blueprints=self._build_module_blueprints(self.config.module_configs),
    device=self.device,
)
```

## Verification
- ✅ Direct model loading from checkpoint works
- ✅ MCP `load_checkpoint` tool works without `restart_server`
- ✅ Training and generation succeed after checkpoint load
- ✅ No tensor size mismatch errors

## Files Modified
- `salience_os_seed/proto_lm/_torch_impl.py` (lines 710-728)

## Date
2025-10-21
