# autoresearch

This is an experiment to have the LLM do its own research.

## Setup

To set up a new experiment, work with the user to:

1. **Agree on a run tag**: propose a tag based on today's date (e.g. `mar5`). The branch `autoresearch/<tag>` must not already exist - this is a fresh run.
2. **Create the branch**: `git checkout -b autoresearch/<tag>` from current master.
3. **Read the in-scope files**: The repo is small. Read these files for full context:
   - `README.md` - repository context.
   - `prepare.py` - fixed constants, data prep, tokenizer, dataloader, evaluation. Do not modify.
   - `train.py` - the file you modify. Model architecture, optimizer, training loop.
   - `human_ideas.md` - operator idea inbox. If it exists, read it. If it does not exist, create it from the local template and leave it untracked.
4. **Verify data exists**: Check that the local autoresearch cache contains data and a tokenizer. If not, tell the human to run `uv run prepare.py`.
5. **Initialize results.tsv**: Create `results.tsv` with just the header row. The baseline will be recorded after the first run.
6. **Confirm and go**: Confirm setup looks good.

Once you get confirmation, kick off the experimentation.

## Experimentation

Each experiment runs on a single GPU. The training script runs for a fixed 5-minute training budget (wall clock training time, excluding startup/compilation). You launch it simply as: `uv run train.py`.

**What you CAN do:**
- Modify `train.py` - this is the only file you edit. Everything is fair game: model architecture, optimizer, hyperparameters, training loop, batch size, model size, etc.

**What you CANNOT do:**
- Modify `prepare.py`. It is read-only. It contains the fixed evaluation, data loading, tokenizer, and training constants.
- Install new packages or add dependencies. You can only use what's already in `pyproject.toml`.
- Modify the evaluation harness. The `evaluate_bpb` function in `prepare.py` is the ground truth metric.

**Use a layered metric stack.** In this repo, `val_bpb` is the primary metric for the byte-level trunk. It is not the whole-system king metric.

- **Primary trunk metric:** lowest held-out `val_bpb`
- **Hard gates:** no crash, no timeout, and no unreasonable resource blowup
- **Secondary runtime proxies:** `total_seconds`, `num_steps`, and `peak_vram_mb`
- **Future system metric:** if this repo ever gains real selector / controller / task evals, those downstream utility metrics outrank pure trunk proxies for those layers

This means:
- For architecture and optimizer changes to the trunk, optimize `val_bpb` first.
- For runtime and controller changes, treat `val_bpb` as necessary but not sufficient.
- Do not confuse the compression gauge with the steering wheel.

**VRAM** is a soft constraint. Some increase is acceptable for meaningful `val_bpb` gains, but it should not blow up dramatically.

**Simplicity criterion**: All else being equal, simpler is better. A small improvement that adds ugly complexity is not worth it. Conversely, removing something and getting equal or better results is a great outcome - that is a simplification win. When evaluating whether to keep a change, weigh the complexity cost against the improvement magnitude. A 0.001 val_bpb improvement that adds 20 lines of hacky code is probably not worth it. A 0.001 val_bpb improvement from deleting code definitely is. An improvement of about 0 with much simpler code is still a win.

**The first run**: Your very first run should always be to establish the baseline, so run the training script as-is.

## Keep / Discard Policy

Use the following decision rule when deciding whether to advance the branch:

1. Reject crashes, OOMs, invalid metrics, and timeout runs.
2. If `val_bpb` improves materially, keep the run unless resource use blows up enough to make the result impractical.
3. If `val_bpb` is effectively tied (for this repo, think on the order of `~0.002` or less), use runtime proxies to break the tie:
   - prefer lower `total_seconds`
   - prefer higher `num_steps`
   - prefer lower `peak_vram_mb`
4. Do **not** keep a clearly worse `val_bpb` just because a proxy improved, unless the human explicitly says this phase is about runtime/efficiency rather than trunk quality.
5. When in doubt, treat `val_bpb` as the trunk truth signal and the other metrics as budget-aware tie-breakers.

## Human Idea Injection

The human is allowed to inject new ideas at any time through `human_ideas.md`.

- Treat `human_ideas.md` as a raw inbox, not a polished spec. The human may paste rough math, paper references, pseudocode, constraints, or direct commands.
- Do not require a strict format. Your job is to translate rough notes into concrete `train.py` experiments.
- Before **every single experiment**, reread `human_ideas.md` fresh from disk. Do not rely on the version you saw at setup time.
- If the human added something new, prioritize it over self-generated ideas unless it is clearly invalid, impossible under the repo constraints, or would obviously waste a run.
- If the human writes a hard constraint such as `constraint: do not change X`, honor it until the file changes.
- When an experiment is derived from the inbox, make that explicit in `results.tsv` with a short prefix like `human:` so the provenance is visible.
- Do not rewrite or clean up the human's notes unless explicitly asked. The inbox belongs to the human.

## Output format

Once the script finishes it prints a summary like this:

```
---
val_bpb:          0.997900
training_seconds: 300.1
total_seconds:    325.9
peak_vram_mb:     45060.2
mfu_percent:      39.80
total_tokens_M:   499.6
num_steps:        953
num_params_M:     50.3
depth:            8
```

You can extract the key metrics from the log file:

```
grep "^val_bpb:\|^peak_vram_mb:\|^total_seconds:\|^num_steps:" run.log
```

You can score a run against the current best reference with:

```
uv run score_run.py run.log --reference current_best_run.json --mode trunk
```

## Logging results

When an experiment is done, log it to `results.tsv` (tab-separated, not comma-separated).

The TSV has a header row and 5 columns:

```
commit	val_bpb	memory_gb	status	description
```

1. git commit hash (short, 7 chars)
2. val_bpb achieved (e.g. `1.234567`) - use `0.000000` for crashes
3. peak memory in GB, round to `.1f` (e.g. `12.3` - divide `peak_vram_mb` by `1024`) - use `0.0` for crashes
4. status: `keep`, `discard`, or `crash`
5. short text description of what this experiment tried

In the description, include the experiment intent where useful, for example:
- `trunk:` for model-quality changes
- `runtime:` for throughput / memory / control changes
- `human:` when the idea came from `human_ideas.md`

If the run is near a tie on `val_bpb`, include the supporting runtime facts in the description, e.g. `sec=449.9 steps=36`.

Example:

```
commit	val_bpb	memory_gb	status	description
a1b2c3d	0.997900	44.0	keep	baseline
b2c3d4e	0.993200	44.2	keep	increase LR to 0.04
c3d4e5f	1.005000	44.0	discard	switch to GeLU activation
d4e5f6g	0.000000	0.0	crash	double model width (OOM)
```

## The experiment loop

The experiment runs on a dedicated branch (e.g. `autoresearch/mar5` or `autoresearch/mar5-gpu0`).

LOOP FOREVER:

1. Refresh operator context: look at the git state and reread `human_ideas.md` from disk.
2. Choose the next experiment. Prefer a fresh human-supplied idea when available; otherwise generate your own.
3. Tune `train.py` with that experimental idea by directly hacking the code.
4. git commit
5. Run the experiment: `uv run train.py > run.log 2>&1` (redirect everything - do not use tee or let output flood your context)
6. Read out the results:
   - raw metrics: `grep "^val_bpb:\|^peak_vram_mb:\|^total_seconds:\|^num_steps:" run.log`
   - layered decision: `uv run score_run.py run.log --reference current_best_run.json --mode trunk`
7. If the grep output is empty, the run crashed. Run `tail -n 50 run.log` to read the Python stack trace and attempt a fix. If you cannot get things to work after more than a few attempts, give up.
8. Record the results in the TSV (do not commit `results.tsv`; leave it untracked)
9. Decide using the layered metric stack:
   - materially better `val_bpb` wins
   - near-ties are broken by `total_seconds`, `num_steps`, and `peak_vram_mb`
   - clearly worse `val_bpb` usually loses
10. If the run wins under that policy, advance the branch and keep the git commit.
11. If the run loses under that policy, reset back to where you started.

The idea is that you are a completely autonomous researcher trying things out. If they work, keep. If they do not, discard. Keep advancing the branch only when the result justifies it. If you feel like you are getting stuck, rewind very sparingly.

**Timeout**: Each experiment should take about 5 minutes of training plus startup/eval overhead. If a run exceeds 10 minutes without a good reason, kill it and treat it as a failure.

**Crashes**: If a run crashes (OOM, bug, etc.), use judgment. If it is something dumb and easy to fix (typo, missing import), fix it and rerun. If the idea itself is fundamentally broken, skip it, log `crash` in the TSV, and move on.

**NEVER STOP**: Once the experiment loop has begun, do not pause to ask the human if you should continue. Do not ask "should I keep going?" or "is this a good stopping point?". The human may be asleep or away from the machine and expects you to continue indefinitely until manually stopped. You are autonomous. If you run out of ideas, think harder - reread the in-scope files, combine near-misses, and try more radical changes.
