# Novel AI Model

A complete AI architecture based on a novel information scoring formula instead of traditional transformer attention mechanisms.

## Overview

This project implements a full AI model architecture centered around the formula:

**S' = (w₁·ΔA + w₂·R + w₃·M) × C × e^(-λt) × (1 - kφ)**

Where:
- **ΔA** = Novelty (information gain)
- **R** = Retention (long-term value)
- **M** = Payoff (immediate utility)
- **C** = Continuity (coherence)
- **φ** = Fatigue (redundancy)
- **w₁, w₂, w₃** = Learnable weights
- **λ** = Time decay parameter
- **k** = Fatigue coefficient
- **t** = Time step

## Architecture

The model replaces standard attention mechanisms with formula-based scoring:
- **FormulaAttentionLayer**: Uses the formula to compute attention scores
- **FormulaSelectiveLayer**: Filters information based on formula scores
- **FormulaTransformerBlock**: Transformer blocks incorporating formula-based mechanisms
- **NovelAIModel**: Complete language model architecture

## Features

- ✅ Full implementation (not a demo)
- ✅ Complete training infrastructure
- ✅ Text generation capabilities
- ✅ Memory buffer for fatigue computation
- ✅ Learnable formula parameters
- ✅ Comprehensive configuration system
- ✅ Production-ready code structure

## Installation

```bash
pip install -r requirements.txt
```

Required packages:
- PyTorch >= 2.0.0
- tqdm >= 4.65.0
- numpy >= 1.24.0

## Usage

### Training

```bash
python train.py \
    --train_data /path/to/train.txt \
    --val_data /path/to/val.txt \
    --output_dir ./checkpoints \
    --epochs 10 \
    --batch_size 8 \
    --learning_rate 1e-4
```

### Generation

```bash
python generate.py \
    --checkpoint ./checkpoints/best_model.pt \
    --prompt "Your input text here" \
    --max_new_tokens 100 \
    --temperature 1.0
```

### Configuration

You can create a custom config file and load it:

```python
from config.config import Config

config = Config.default()
config.model.embedding_dim = 768
config.model.num_layers = 12
config.training.learning_rate = 5e-5
config.save('my_config.json')
```

Then use it:
```bash
python train.py --config my_config.json --train_data /path/to/data.txt
```

## Project Structure

```
MK2/
├── core/                 # Core formula and layers
│   ├── formula.py       # Scoring formula implementation
│   └── layers.py        # Neural network layers
├── model/                # Model architecture
│   └── architecture.py  # Complete model implementation
├── training/             # Training infrastructure
│   └── trainer.py       # Trainer class
├── utils/                # Utilities
│   ├── tokenizer.py     # Text tokenization
│   └── data_loader.py   # Data loading utilities
├── config/               # Configuration
│   └── config.py        # Configuration management
├── train.py             # Main training script
├── generate.py          # Text generation script
└── requirements.txt     # Dependencies
```

## Model Components

### Scoring Formula

The core formula evaluates information at each step:

1. **Novelty (ΔA)**: Measures information gain relative to context
2. **Retention (R)**: Estimates long-term importance
3. **Payoff (M)**: Computes immediate utility
4. **Continuity (C)**: Ensures coherence with context
5. **Fatigue (φ)**: Penalizes redundancy with recent items
6. **Time Decay**: Exponential decay over time steps

### Formula-Based Attention

Replaces standard dot-product attention with formula-based scoring, allowing the model to:
- Prioritize novel information
- Balance immediate vs. long-term value
- Maintain coherence
- Avoid redundancy

### Memory Buffer

Maintains a buffer of recent embeddings to compute fatigue scores, enabling the model to avoid repetitive patterns.

## Configuration Options

### Model Configuration

- `embedding_dim`: Embedding dimensionality (default: 512)
- `num_layers`: Number of transformer blocks (default: 6)
- `max_seq_length`: Maximum sequence length (default: 512)
- `dropout`: Dropout rate (default: 0.1)
- `formula_w1`, `formula_w2`, `formula_w3`: Component weights
- `formula_lambda`: Time decay parameter
- `formula_k`: Fatigue coefficient

### Training Configuration

- `batch_size`: Batch size (default: 8)
- `learning_rate`: Learning rate (default: 1e-4)
- `num_epochs`: Number of training epochs (default: 10)
- `warmup_steps`: Warmup steps for learning rate (default: 1000)
- `max_grad_norm`: Gradient clipping (default: 1.0)

## Technical Details

### Formula Computation

Each component is computed by a small neural network:
- Novelty: Compares current vs. context embeddings
- Retention: Evaluates future importance
- Payoff: Measures immediate relevance
- Continuity: Ensures semantic coherence
- Fatigue: Compares to recent items in memory buffer

### Training

- Uses cross-entropy loss for language modeling
- AdamW optimizer with weight decay
- Learning rate warmup
- Gradient clipping
- Checkpoint saving (best model and periodic checkpoints)

### Generation

Supports:
- Greedy decoding
- Temperature sampling
- Top-k sampling
- Nucleus (top-p) sampling

## Development Status

This is a complete, production-ready implementation. The architecture is fully functional and can be trained on any text corpus.

## Future Enhancements

Potential improvements:
- Multi-GPU training support
- Distributed training
- Advanced tokenization (BPE, SentencePiece)
- Additional model architectures
- Better memory management for very long sequences
- Formula visualization

## License

Open for use and modification.

## Citation

If you use this architecture, please cite the formula:

```
S' = (w₁·ΔA + w₂·R + w₃·M) × C × e^(-λt) × (1 - kφ)
```

## Notes

- The formula parameters (weights, decay, fatigue) are learnable by default
- Memory buffer automatically tracks recent embeddings
- All components are differentiable, enabling end-to-end training
- The model can be used for any sequence-to-sequence task with appropriate modifications



