# Benchmark Training and Evaluation

## Overview

The training system is now set up to train the novel AI model and evaluate it on standard benchmarks, allowing comparison with other open-source models.

## What's Been Set Up

### 1. Benchmark Evaluation Infrastructure (`evaluation/benchmarks.py`)
- **PerplexityEvaluator**: Computes perplexity (standard LM metric)
- **AccuracyEvaluator**: Computes next-token prediction accuracy
- **BenchmarkSuite**: Comprehensive evaluation framework

### 2. Training with Evaluation (`training/train_with_eval.py`)
- Trains model on your dataset
- Evaluates on test sets during training
- Saves best model based on validation loss
- Generates evaluation reports

### 3. Benchmark Scripts
- `scripts/run_benchmarks.py`: Run comprehensive evaluations
- `scripts/create_larger_sample.py`: Generate sample training data
- `scripts/download_wikitext.py`: Download WikiText-2 benchmark dataset

### 4. Comparison Framework
- Compares against baseline models (GPT-2 Small/Medium/Large)
- Generates detailed comparison reports
- Tracks perplexity and accuracy metrics

## Current Training Status

Training has been started with the command:
```bash
python training/train_with_eval.py \
    --train_data ./data/sample_train.txt \
    --val_data ./data/sample_valid.txt \
    --test_data ./data/sample_test.txt \
    --test_names sample_test \
    --epochs 5 \
    --batch_size 4 \
    --learning_rate 1e-3 \
    --eval_interval 1 \
    --output_dir ./checkpoints/benchmark_run
```

**Training Data**: 5,440 texts for training, 680 for validation, 680 for testing

## Monitoring Training

Check training progress:
```bash
# View checkpoint directory
ls ./checkpoints/benchmark_run/

# Check for best model
ls ./checkpoints/benchmark_run/best_model.pt

# View training logs (if available)
tail -f ./checkpoints/benchmark_run/training.log
```

## After Training Completes

### 1. Run Comprehensive Benchmarks
```bash
python scripts/run_benchmarks.py \
    --checkpoint ./checkpoints/benchmark_run/best_model.pt \
    --datasets ./data/sample_test.txt \
    --dataset_names sample_test \
    --output_dir ./benchmark_results \
    --model_name "Novel AI Model (Formula-Based)"
```

### 2. View Results
Results will be saved in:
- `./benchmark_results/benchmark_results.json` - Detailed metrics
- `./benchmark_results/benchmark_comparison_report.txt` - Comparison report

### 3. Compare Metrics

The report will show:
- **Perplexity**: Lower is better (target: <35 for GPT-2 Small equivalent)
- **Accuracy**: Higher is better (target: >45% for GPT-2 Small equivalent)
- **Top-5 Accuracy**: Token in top 5 predictions (target: >65%)

## Benchmark Baselines (Reference)

For comparison, approximate baseline results:

| Model | Parameters | Perplexity | Accuracy |
|-------|-----------|------------|----------|
| Transformer Base | ~65M | ~40 | ~42% |
| GPT-2 Small | 124M | ~35 | ~45% |
| GPT-2 Medium | 355M | ~26 | ~50% |
| GPT-2 Large | 774M | ~22 | ~52% |

*Note: These are approximate values and depend on the dataset.*

## Expected Performance

For a model with similar architecture to GPT-2 Small (124M params):

**Good Performance:**
- Perplexity: 30-40
- Accuracy: 40-50%
- Top-5 Accuracy: 60-70%

**Excellent Performance:**
- Perplexity: <30
- Accuracy: >50%
- Top-5 Accuracy: >70%

## Real-World Benchmarks

For serious evaluation, use standard datasets:

### WikiText-2
```bash
# Download
python scripts/download_wikitext.py

# Train
python training/train_with_eval.py \
    --train_data ./data/wikitext_train.txt \
    --val_data ./data/wikitext_valid.txt \
    --test_data ./data/wikitext_test.txt \
    --test_names wikitext_test \
    --epochs 10 \
    --batch_size 8 \
    --output_dir ./checkpoints/wikitext_run
```

### Penn Treebank (PTB)
- Small, widely-used benchmark
- Download from official sources
- Similar workflow

## Key Metrics Explained

### Perplexity
- **Definition**: exp(cross_entropy_loss)
- **Meaning**: Average number of choices model considers for next token
- **Lower is better**: 20 is excellent, 30-40 is good, >50 needs improvement

### Accuracy
- **Definition**: Percentage of correct next-token predictions
- **Meaning**: How often model predicts the correct token
- **Higher is better**: >50% is excellent, 40-50% is good

### Top-5 Accuracy
- **Definition**: Percentage where correct token is in top 5 predictions
- **Meaning**: Model is "close" even if not exactly right
- **Higher is better**: >70% is excellent

## Troubleshooting

**Training is slow:**
- Reduce batch size
- Reduce max_seq_length
- Use smaller model (fewer layers)

**Out of memory:**
- Reduce batch_size
- Reduce max_seq_length
- Reduce embedding_dim

**Poor performance:**
- More training data
- More training epochs
- Tune learning rate
- Increase model size

## Next Steps

1. **Wait for training to complete** (check `./checkpoints/benchmark_run/`)
2. **Run benchmark evaluation** when model is ready
3. **Compare results** against baselines
4. **Iterate**: Adjust hyperparameters and retrain if needed

## Model Architecture Details

The novel AI model uses:
- Formula-based attention instead of standard dot-product
- Learnable weights (w1, w2, w3) for novelty, retention, payoff
- Memory buffer for fatigue computation
- Standard transformer infrastructure

This allows the model to:
- Prioritize novel information
- Balance immediate vs. long-term value
- Maintain coherence
- Avoid redundancy

## Questions?

Check:
- `BENCHMARK_TRAINING.md` - Detailed training guide
- `README.md` - General model documentation
- `evaluation/benchmarks.py` - Evaluation code



