# Benchmark Training Guide

This guide explains how to train the novel AI model and evaluate it on standard benchmarks.

## Quick Start

### 1. Create Training Data

```bash
# Create sample data for testing
python scripts/create_larger_sample.py

# Or download real datasets (like WikiText)
python scripts/download_wikitext.py
```

### 2. Train with Evaluation

```bash
python training/train_with_eval.py \
    --train_data ./data/sample_train.txt \
    --val_data ./data/sample_valid.txt \
    --test_data ./data/sample_test.txt \
    --test_names sample_test \
    --epochs 10 \
    --batch_size 8 \
    --learning_rate 1e-4 \
    --eval_interval 1 \
    --output_dir ./checkpoints/benchmark_run \
    --save_best
```

### 3. Run Comprehensive Benchmarks

After training, run detailed benchmarks:

```bash
python scripts/run_benchmarks.py \
    --checkpoint ./checkpoints/benchmark_run/best_model.pt \
    --datasets ./data/sample_test.txt \
    --dataset_names sample_test \
    --output_dir ./benchmark_results \
    --model_name "Novel AI Model (Formula-Based)"
```

## Benchmark Metrics

The evaluation includes:

1. **Perplexity**: Standard language modeling metric (lower is better)
   - Measures how well the model predicts the next token
   - Formula: exp(cross_entropy_loss)

2. **Accuracy**: Next token prediction accuracy (higher is better)
   - Top-1 accuracy: Exact token match
   - Top-5 accuracy: Token in top 5 predictions

3. **Loss**: Average cross-entropy loss (lower is better)

## Comparison with Other Models

The benchmark script includes approximate baseline results for:
- GPT-2 Small (124M parameters): PPL ~35, Accuracy ~45%
- GPT-2 Medium (355M parameters): PPL ~26, Accuracy ~50%
- GPT-2 Large (774M parameters): PPL ~22, Accuracy ~52%
- Transformer Base: PPL ~40, Accuracy ~42%

## Real-World Datasets

For serious benchmarking, use standard datasets:

- **WikiText-2**: General domain, 4M tokens
  - Download: `python scripts/download_wikitext.py`
  
- **Penn Treebank (PTB)**: Small corpus for quick benchmarks

- **OpenWebText**: Large-scale web text (requires significant resources)

## Training Tips

1. **Start Small**: Use small models and datasets for initial testing
2. **Monitor Metrics**: Watch validation loss and perplexity
3. **Adjust Learning Rate**: If loss isn't decreasing, try lower LR
4. **Use Validation Set**: Always validate to avoid overfitting
5. **Save Checkpoints**: Enable `--save_best` to keep best model

## Expected Performance

For a model with similar parameters to GPT-2 Small:
- **Perplexity**: Should aim for 30-40 on WikiText-2
- **Accuracy**: Should achieve 40-50% on next token prediction
- **Training Time**: ~few hours on GPU for small models

## Advanced Usage

### Custom Datasets

Prepare your dataset as plain text files, one paragraph per line (separated by blank lines).

### Multi-GPU Training

Current implementation uses single GPU. For multi-GPU, wrap model in `nn.DataParallel` or use `DistributedDataParallel`.

### Hyperparameter Tuning

Key parameters to tune:
- `embedding_dim`: 256-1024 (larger = better but slower)
- `num_layers`: 4-12 (more layers = deeper model)
- `learning_rate`: 1e-5 to 1e-3
- `batch_size`: As large as memory allows
- Formula weights (`w1`, `w2`, `w3`): Start with defaults, let them learn

## Results Location

After training and evaluation:
- Model checkpoints: `./checkpoints/benchmark_run/`
- Benchmark results: `./benchmark_results/benchmark_results.json`
- Comparison report: `./benchmark_results/benchmark_comparison_report.txt`

## Troubleshooting

**Issue**: "num_samples should be a positive integer"
- **Solution**: Dataset is too small. Use `create_larger_sample.py` or a real dataset.

**Issue**: Out of memory
- **Solution**: Reduce `batch_size` or `max_seq_length`

**Issue**: Loss not decreasing
- **Solution**: Lower learning rate, check data quality, ensure sufficient training data



