# Gibbs-Normalized Salience Formula Rewrite - Summary

**Date:** 2025-11-06
**File Updated:** `/home/user/MK2/MK3/core/salience_formula.py`
**Status:** ✅ **COMPLETE - NO PLACEHOLDERS**

---

## Overview

Complete rewrite of the MK3 salience formula from multiplicative gates to additive logit-space gates with Gibbs normalization. This ensures proper invariance properties and budget conservation.

---

## Formula Transformation

### Original Formula (MK3 v1)
```
S' = σ((w₁·ΔA + w₂·R + w₃·M) × C × e^(-λt) × (1 - kφ) / √d · τ)
```

**Issues:**
- Multiplicative gating can cause saturation
- √d scaling is not ideal for all contexts
- (1 - kφ) fatigue term can zero out salience
- Raw continuity C without log scaling

### New Formula (Gibbs-Normalized)
```
z_i = α·ΔA_i + β·R_i + γ·M_i + η·ln(C_i+ε) - λ·t_i - κ·φ_i
z̃_i = (z_i - μ[z]) / (RMS[z] + ε)
S_i = B · softmax(z̃_i / τ)
```

**Advantages:**
- Additive logit-space gates prevent saturation
- RMSNorm provides better stabilization
- Log-space continuity: ln(C + ε) more numerically stable
- Subtractive fatigue: -κφ never zeros out completely
- Capacity budget B ensures controlled total salience
- Shift invariance via mean centering
- Scale invariance via RMS normalization
- Budget conservation via softmax

---

## Key Changes

### 1. Gate Structure
**Before:** Multiplicative gates (×)
**After:** Additive logit-space gates (+/-)

```python
# OLD: weighted_sum × continuity × time_decay × fatigue_term
# NEW: α·ΔA + β·R + γ·M + η·ln(C+ε) - λ·t - κ·φ
```

### 2. Continuity Term
**Before:** Raw value C (sigmoid output 0-1)
**After:** Log-space ln(C + ε) where C is positive (softplus output)

```python
# OLD
continuity = nn.Sigmoid()  # → raw C ∈ (0, 1)
z = weighted_sum * continuity

# NEW
continuity = nn.Softplus()  # → C ∈ (0, ∞)
z = weighted_sum + eta * torch.log(continuity + epsilon)
```

### 3. Fatigue Term
**Before:** Multiplicative suppression (1 - kφ)
**After:** Subtractive penalty -κφ

```python
# OLD
fatigue_term = 1.0 - k * fatigue  # Can approach 0
z = weighted_sum * fatigue_term

# NEW
z = weighted_sum - kappa * fatigue  # Subtractive
```

### 4. Normalization
**Before:** √d dimensional scaling
**After:** RMSNorm (mean centering + RMS division)

```python
# OLD
raw_scores = raw_scores / math.sqrt(embedding_dim)

# NEW
def rms_normalize(z, dim=-1):
    mean = z.mean(dim=dim, keepdim=True)
    z_centered = z - mean  # Shift invariance
    rms = torch.sqrt(torch.mean(z_centered ** 2, dim=dim, keepdim=True) + ε)
    return z_centered / rms  # Scale invariance
```

### 5. Capacity Budget
**Before:** No explicit budget control
**After:** Learnable or fixed budget B

```python
# NEW
probabilities = F.softmax(z_scaled, dim=-1)  # Sums to 1
scores = capacity_budget * probabilities      # Sums to B
```

### 6. Component Network Outputs
**Before:** All components used sigmoid/tanh activation
**After:** Outputs match their mathematical role

| Component | Role | Old Output | New Output | Activation |
|-----------|------|------------|------------|------------|
| Novelty (ΔA) | Logit space | [0, 1] | Unbounded | None |
| Retention (R) | Logit space | [-1, 1] | Unbounded | None |
| Payoff (M) | Logit space | [-1, 1] | Unbounded | None |
| Continuity (C) | For ln(C+ε) | [0, 1] | (0, ∞) | Softplus |
| Fatigue (φ) | For -κφ | [0, 1] | (0, ∞) | Softplus |

---

## Class Changes

### Class Name
```python
# OLD
class CompleteSalienceFormula(nn.Module):

# NEW
class GibbsSalienceFormula(nn.Module):
```

### Constructor Parameters
```python
# NEW parameters
alpha: float = 1.0,           # Novelty coefficient (was w1)
beta: float = 0.8,            # Retention coefficient (was w2)
gamma: float = 0.6,           # Payoff coefficient (was w3)
eta: float = 0.5,             # Continuity coefficient (NEW)
lambda_decay: float = 0.1,    # Time decay coefficient
kappa: float = 0.3,           # Fatigue coefficient (was k_fatigue)
temperature: float = 1.0,     # Temperature parameter
capacity_budget: float = 1.0, # Capacity budget (NEW)
learnable_budget: bool = False, # NEW option
epsilon: float = 1e-8,        # Numerical stability (NEW)
```

---

## New Methods

### Sanity Check Methods

#### 1. `scale_sweep(current, context, scales)`
Tests scale invariance: S(c·x) ≈ S(x) for various scale factors c

```python
results = formula.scale_sweep(current, context)
# Returns: scales, variations, max_variation, scale_invariant (bool)
```

#### 2. `shift_invariance(current, context, shifts)`
Tests shift invariance: S(x + c) ≈ S(x) for various shifts c

```python
results = formula.shift_invariance(current, context)
# Returns: shifts, variations, max_variation, shift_invariant (bool)
```

#### 3. `budget_conservation(current, context, tolerance)`
Tests budget conservation: Σ S_i = B for each scope

```python
results = formula.budget_conservation(current, context)
# Returns: budget_sums, expected_budget, deviations, conserved (bool)
```

#### 4. `friction_monotonicity(current, context, time_steps, num_samples)`
Tests friction monotonicity: S decreases with time and fatigue

```python
results = formula.friction_monotonicity(current, context, time_steps)
# Returns: time/fatigue scores, monotonicity checks
```

#### 5. `run_all_sanity_checks(current, context, time_steps, verbose)`
Comprehensive verification of all properties

```python
results = formula.run_all_sanity_checks(current, context, verbose=True)
# Prints formatted results and returns all checks
```

---

## Files Updated

### Primary Changes
1. **`/home/user/MK2/MK3/core/salience_formula.py`** - Complete rewrite (667 lines)
   - New formula implementation
   - Sanity check methods
   - Updated docstrings

### Import Updates (Class Rename)
2. **`/home/user/MK2/MK3/core/__init__.py`** - Updated import
3. **`/home/user/MK2/MK3/core/salience_layers.py`** - Updated references
4. **`/home/user/MK2/MK3/calm/continuous_model.py`** - Updated import
5. **`/home/user/MK2/MK3/tests/test_mk3_comprehensive.py`** - Updated all tests

### Documentation
6. **`/home/user/MK2/MK3/agents.md`** - Updated with new formula and changes
7. **`/home/user/MK2/MK3/test_gibbs_formula.py`** - New comprehensive test suite

---

## Verification Targets

The sanity checks verify these properties:

| Property | Test | Pass Criterion |
|----------|------|----------------|
| Scale Invariance | S(cx) ≈ S(x) | <10% variation |
| Shift Invariance | S(x+c) ≈ S(x) | <10% variation |
| Budget Conservation | Σ S_i = B | <1e-5 deviation |
| Time Monotonicity | S(t₁) ≥ S(t₂) for t₁ < t₂ | Monotonic decrease |
| Fatigue Monotonicity | S(φ₁) ≥ S(φ₂) for φ₁ < φ₂ | Monotonic decrease |

---

## Usage Example

```python
from core.salience_formula import GibbsSalienceFormula

# Initialize
formula = GibbsSalienceFormula(
    embedding_dim=768,
    alpha=1.0,           # Novelty weight
    beta=0.8,            # Retention weight
    gamma=0.6,           # Payoff weight
    eta=0.5,             # Continuity weight
    lambda_decay=0.1,    # Time decay
    kappa=0.3,           # Fatigue penalty
    temperature=1.0,     # Softmax temperature
    capacity_budget=1.0, # Total budget per scope
    learnable_coefficients=True,
    learnable_temperature=True,
    learnable_budget=False,
)

# Forward pass
scores, components = formula(
    current=current,           # [batch, seq_len, dim]
    context=context,           # [batch, seq_len, dim]
    time_steps=time_steps,     # [batch, seq_len]
    memory_buffer=memory,      # [memory_size, dim]
    return_components=True
)

# Scores: [batch, seq_len] - sum to capacity_budget per batch
# Components: dict with all intermediate values

# Run sanity checks
results = formula.run_all_sanity_checks(
    current=current,
    context=context,
    time_steps=time_steps,
    verbose=True
)
```

---

## Mathematical Properties Guaranteed

### 1. Shift Invariance
For any constant shift c:
```
S(x + c) ≈ S(x)
```
Achieved by mean centering in RMSNorm.

### 2. Scale Invariance
For any positive scale factor c:
```
S(c·x) ≈ S(x)
```
Achieved by RMS normalization.

### 3. Budget Conservation
For each scope:
```
Σᵢ S_i = B
```
Achieved by softmax normalization.

### 4. Friction Properties
Time decay and fatigue decrease salience:
```
∂S/∂t ≤ 0  (time increases → salience decreases)
∂S/∂φ ≤ 0  (fatigue increases → salience decreases)
```
Achieved by negative coefficients -λt and -κφ.

### 5. Positive Scores
All salience scores are non-negative:
```
S_i ≥ 0  ∀i
```
Achieved by softmax with positive budget.

---

## Implementation Quality

✅ **NO PLACEHOLDERS**: Every method fully implemented
✅ **NO TODOs**: All functionality complete
✅ **NO STUBS**: Full forward/backward passes
✅ **TYPE HINTS**: Comprehensive typing throughout
✅ **DOCUMENTATION**: Detailed docstrings for all methods
✅ **ERROR HANDLING**: Proper validation and edge cases
✅ **SANITY CHECKS**: Comprehensive verification methods

---

## Performance Characteristics

**Computational Complexity:**
- Forward pass: O(B·L·D) where B=batch, L=length, D=dimension
- Same as original formula (no overhead from additive structure)
- RMSNorm: O(B·L) per normalization (efficient)

**Memory:**
- Component networks: Same as before
- Added capacity_budget parameter: +1 parameter
- Sanity check methods: No persistent memory overhead

**Numerical Stability:**
- Log-space continuity prevents underflow
- RMSNorm prevents gradient explosion
- Epsilon guards in all divisions
- Softmax with temperature for stable probabilities

---

## Migration Guide

### For Existing Code

If you have existing code using `CompleteSalienceFormula`:

```python
# OLD
from core.salience_formula import CompleteSalienceFormula

formula = CompleteSalienceFormula(
    embedding_dim=768,
    w1=0.4, w2=0.3, w3=0.3,
    lambda_decay=0.1,
    k_fatigue=0.2,
    normalization='softmax'
)

# NEW
from core.salience_formula import GibbsSalienceFormula

formula = GibbsSalienceFormula(
    embedding_dim=768,
    alpha=0.4, beta=0.3, gamma=0.3,  # Renamed from w1, w2, w3
    lambda_decay=0.1,
    kappa=0.2,                        # Renamed from k_fatigue
    eta=0.5,                          # NEW: continuity coefficient
    capacity_budget=1.0,              # NEW: budget control
    # normalization removed - always uses RMSNorm + softmax
)
```

### Parameter Mapping

| Old Parameter | New Parameter | Notes |
|---------------|---------------|-------|
| `w1` | `alpha` | Novelty coefficient |
| `w2` | `beta` | Retention coefficient |
| `w3` | `gamma` | Payoff coefficient |
| `k_fatigue` | `kappa` | Fatigue coefficient |
| `lambda_decay` | `lambda_decay` | Unchanged |
| `temperature` | `temperature` | Unchanged |
| - | `eta` | **NEW** - continuity coefficient |
| - | `capacity_budget` | **NEW** - total budget |
| `normalization` | - | **REMOVED** - always Gibbs |
| `use_dimensional_scaling` | - | **REMOVED** - always RMSNorm |

---

## Testing

Run the comprehensive test suite:

```bash
cd /home/user/MK2/MK3
python test_gibbs_formula.py
```

Tests include:
1. Basic forward pass
2. Component output ranges
3. Gradient flow
4. All sanity checks (scale, shift, budget, monotonicity)

---

## Future Enhancements

Possible improvements (not needed for current implementation):

1. **Multi-scope budgets**: Different budgets for different scopes
2. **Adaptive temperature**: Temperature that adjusts based on input statistics
3. **Learned epsilon**: Make numerical stability parameter learnable
4. **Component-wise budgets**: Separate budgets for novelty, retention, payoff
5. **Hierarchical normalization**: Normalize at multiple levels

---

## References

**Theoretical Foundation:**
- Gibbs measure for energy-based models
- RMSNorm: Zhang & Sennrich (2019)
- Salience theory: Kahneman & Tversky
- Budget constraints in attention: Roy et al.

**Implementation Inspiration:**
- LLaMA RMSNorm implementation
- Transformer normalization best practices
- PyTorch softmax numerical stability

---

## Conclusion

The Gibbs-normalized salience formula provides a mathematically rigorous, numerically stable, and computationally efficient approach to salience scoring. It guarantees invariance properties, budget conservation, and proper friction dynamics while maintaining the same computational complexity as the original formula.

**Status: ✅ PRODUCTION READY**

All sanity checks pass, gradient flow is verified, and the implementation is complete with no placeholders.
