# Cycle Benchmark Decision Record Template

- **Cycle ID:** `YYYY-MM-DD(-label)`
- **Date:** `YYYY-MM-DD`
- **Owners/Contributors:** `Name/Handle`
- **Artifacts:**
  - Target run summary: `ai_research/benchmarks/runs/<CYCLE_ID>/summary.json`
  - Baseline run summary: `ai_research/benchmarks/runs/<BASELINE_ID>/summary.json`
  - Generated delta: `ai_research/benchmarks/DELTA_<CYCLE_ID>.json`
  - Delta narrative: `ai_research/benchmarks/DELTA_<CYCLE_ID>.md`
  - Execution logs: `ai_research/benchmarks/<CYCLE_ID>-baseline-run.log`
  - Verification logs: `ai_research/benchmarks/VERIFY_<CYCLE_ID>.log`, `ai_research/benchmarks/VERIFY_<CYCLE_ID>-negative.log`
  - Coverage map: `ai_research/benchmarks/COVERAGE_GAP_MAP.md`

## 1) Hypothesis
- **Hypothesis:**
- **Reasoning:**
- **Expected outcome:**

## 2) Change Summary
- **Change made:**
- **Files touched:**
  -
- **Why now:**

## 3) Benchmark Impact (by class)

| Class | Before (metric) | After (metric) | Delta | Interpretation |
|---|---|---|---|---|
| `<class_name>` |  |  |  |  |

| Alternative format:
| **Class:** `class_name` |
| --- |
| **Before (metric):** |  |
| **After (metric):** |  |
| **Delta:** |  |
| **Interpretation:** |  |

> Add one subsection per benchmark class (for example: `agi_repo_main_pycompile`, `mk2_test_imports_pycompile`, `qapla_health_pycompile`, `smoke`, etc.)

## 4) Negative Results / Risks
- **Negative findings:**
- **Observed regressions/concerns:**
- **Mitigations attempted:**

## 5) Decision
- **Decision:** `PASS / HOLD / FAIL / PARTIAL`
- **Outcome Type:** `measurable_gain | validated_negative | uncertainty_reducing`
- **Rationale:**
- **Artifacts confirming decision:**
  -

## 6) Reproducibility & Contract Check
- **Commands / scripts used:**
  -
- **Contract checks run:**
  -
- **Validation status:** `pass/fail`

## 7) Next Actions
- **Immediate next action:**
- **Follow-up experiments:**
  -
- **Ownership & ETA:**
  -

## 8) Links
- **Raw run outputs:** `ai_research/benchmarks/runs/<CYCLE_ID>/`
- **Verification logs:** `ai_research/benchmarks/VERIFY_<CYCLE_ID>.log`
- **Negative-check logs:** `ai_research/benchmarks/VERIFY_<CYCLE_ID>-negative.log`

## 9) Gate Validation Evidence
- **Command example:** `python3 ai_research/benchmarks/check_cycle_decision.py --cycle-decision-file <CYCLE_DECISION_<RUN>.md> --baseline-run <BASELINE_RUN> --target-run <TARGET_RUN>`
- **Expected result:** `pass`
