# MK6

MK6 is a bottom-up rebuild centered on the Gibbs-normalized salience math, without CALM.

Core direction:
- token-native language model
- salience-guided sparse causal attention
- local window attention plus globally salient prefix tokens
- byte-level tokenizer with no external dependency

Presets:
- `smoke`
- `small_5060`
- `h100_base`

Smoke test:

```powershell
python train.py --preset smoke --test_mode --checkpoint_dir checkpoints\smoke_run
python generate.py --checkpoint checkpoints\smoke_run\best_model.pt --prompt "Salience is"
```

5060-scale run:

```powershell
python train.py --preset small_5060 --train_file path\to\train.txt --val_file path\to\val.txt
```
