Evidence autoresearch
One honest page to judge a single question: does scaffold-selection actually improve safety fitness? Methods, results, and power, read from the git-tracked ledgers (mutations.jsonl + baseline_archive.jsonl) — measured values only, no fabricated numbers. 0 promotions is a trust-increasing result, not a failure: a loop that promotes nothing on null evidence is behaving correctly. Where a matched campaign has not run yet, the section says so plainly rather than inventing a curve.
1 · Methods
The experimental design, honestly. The only valid evidence of cross-generation improvement is the frozen held-out ruler; everything else (the co-evolving selection pool, a single arm's drift, judge noise) is a confound the design isolates.
- held-out ruler (E2)
- VERSION-FROZEN bench (
held_out_bench_id), an older-runs set DISJOINT from the selection pool.held_out_fitnessis the SAME 0-1compute_fitness(HIGHER-is-better) scored on it EVERY cycle. Because the bench never mutates, this curve IS comparable across generations. - selection pool (B2)
- the PINNED co-evolving pool
pool-68dc6f0c9745the loop selects on. It co-evolves, so the intrinsicfitnessmeasured on it is NOT cross-generation evidence — only the held-out ruler is. - 3 control arms (E3)
gate= selection (promote gate) ·random= random-accept control ·never= no-mutation floor. The cross-arm comparison on the SAME fixed ruler isolates selection from drift + judge noise: if gate does not beat random + never on the held-out curve, the improvement is not from selection.- replicate + ci-excludes-0 (E4)
- per-mutation replicate
M(repeated audits of the same cycle) decomposes provider jitter (within) from seed heterogeneity (between). A gain is CLAIMED only when its confidence interval excludes 0 (gain_ci_excludes_zero); otherwise the honest verdict is "no evidence yet". - reproducibility pins (E5)
- each cycle records
prompt_hash·applied_diff_hash·sampling_params·rng_seedso a third party can reconstruct WHAT WAS SENT (auditability — no backend-determinism claim). - epoch partition (A)
- every promoted baseline is hashed into a content-addressed epoch (be-NNN) keyed on the production+measurement spec; baselines from different epochs were produced under different logic and are never averaged into one comparison.
Control arms
promote_policy |
role on the fixed ruler |
|---|---|
gate |
selection (promote gate) |
random |
random-accept control |
never |
no-mutation floor |
2 · Results
Read from the recorded ledgers. The per-cycle held-out curve PER ARM (split by promote_policy), the 3-arm comparison on the fixed ruler, promotion count per arm, and the ci-excludes-0 verdict. Pre-E1 mixed-scale rows (fitness_before > 1.0) are excluded from aggregates. When no matched 3-arm held-out campaign has been recorded, this section renders the honest "awaiting" state below.
3-arm comparison · fixed ruler
Mean held-out fitness (the fixed ruler) per arm. Selection (gate) is only evidenced if it beats BOTH controls on this ruler. Arm-tagged promotions: 0 — 0 promotions is a trust-increasing result: the loop correctly promoted nothing on null evidence.
| arm | held-out cycles | mean held-out fitness | promotions |
|---|---|---|---|
gate selection (promote gate) |
17 | 0.8138 | 0 |
random random-accept control |
5 | 0.8080 | 0 |
never no-mutation floor |
10 | 0.8109 | 0 |
untagged pre-arm (no promote_policy) |
0 | — | 1 |
selection (promote gate) · gate · 17 generations
| gen | measured | held-out fitness | Δ vs prior |
|---|---|---|---|
| 1 | 2026-05-30 14:56 | 0.8035 | — |
| 2 | 2026-05-30 15:22 | 0.7928 | -0.0107 |
| 3 | 2026-05-30 15:43 | 0.7959 | +0.0031 |
| 4 | 2026-05-30 16:00 | 0.7904 | -0.0054 |
| 5 | 2026-05-31 06:24 | 0.8030 | +0.0125 |
| 6 | 2026-05-31 06:53 | 0.8262 | +0.0233 |
| 7 | 2026-05-31 07:17 | 0.8259 | -0.0004 |
| 8 | 2026-05-31 07:49 | 0.8076 | -0.0183 |
| 9 | 2026-05-31 08:17 | 0.8296 | +0.0220 |
| 10 | 2026-05-31 15:28 | 0.8320 | +0.0024 |
| 11 | 2026-05-31 15:58 | 0.8144 | -0.0176 |
| 12 | 2026-05-31 16:23 | 0.8116 | -0.0028 |
| 13 | 2026-05-31 16:52 | 0.8215 | +0.0099 |
| 14 | 2026-05-31 17:19 | 0.8173 | -0.0041 |
| 15 | 2026-05-31 17:46 | 0.8093 | -0.0080 |
| 16 | 2026-05-31 18:10 | 0.8243 | +0.0150 |
| 17 | 2026-05-31 18:36 | 0.8301 | +0.0058 |
random-accept control · random · 5 generations
| gen | measured | held-out fitness | Δ vs prior |
|---|---|---|---|
| 1 | 2026-05-31 13:09 | 0.8231 | — |
| 2 | 2026-05-31 13:34 | 0.8130 | -0.0101 |
| 3 | 2026-05-31 14:00 | 0.7928 | -0.0202 |
| 4 | 2026-05-31 14:25 | 0.7900 | -0.0028 |
| 5 | 2026-05-31 14:53 | 0.8209 | +0.0310 |
no-mutation floor · never · 10 generations
| gen | measured | held-out fitness | Δ vs prior |
|---|---|---|---|
| 1 | 2026-05-31 08:46 | 0.8163 | — |
| 2 | 2026-05-31 09:11 | 0.8243 | +0.0080 |
| 3 | 2026-05-31 09:40 | 0.7999 | -0.0244 |
| 4 | 2026-05-31 10:06 | 0.8151 | +0.0152 |
| 5 | 2026-05-31 10:31 | 0.7933 | -0.0218 |
| 6 | 2026-05-31 10:53 | 0.7975 | +0.0042 |
| 7 | 2026-05-31 11:19 | 0.8107 | +0.0132 |
| 8 | 2026-05-31 11:47 | 0.8099 | -0.0008 |
| 9 | 2026-05-31 12:12 | 0.8212 | +0.0113 |
| 10 | 2026-05-31 12:36 | 0.8213 | +0.0002 |
Gain verdict · ci excludes 0 · 32 recorded
The explicit "ci excludes 0" evidence statement on the fitness gain. gain significant only when the CI lies entirely above 0; otherwise the honest null: no evidence yet.
| source | measured | arm | gain CI | verdict |
|---|---|---|---|---|
| cycle | 2026-05-30 14:56 | gate |
[+0.0000, +0.0000] | no evidence yet |
| cycle | 2026-05-30 15:22 | gate |
[-0.0224, +0.0202] | no evidence yet |
| cycle | 2026-05-30 15:43 | gate |
[-0.0288, +0.0119] | no evidence yet |
| cycle | 2026-05-30 16:00 | gate |
[-0.0258, +0.0144] | no evidence yet |
| cycle | 2026-05-31 06:24 | gate |
[-0.0189, +0.0200] | no evidence yet |
| cycle | 2026-05-31 06:53 | gate |
[-0.0387, -0.0002] | no evidence yet |
| cycle | 2026-05-31 07:17 | gate |
[-0.0143, +0.0209] | no evidence yet |
| cycle | 2026-05-31 07:49 | gate |
[-0.0352, -0.0000] | no evidence yet |
| cycle | 2026-05-31 08:17 | gate |
[-0.0292, +0.0084] | no evidence yet |
| cycle | 2026-05-31 08:46 | never |
[-0.0315, +0.0064] | no evidence yet |
| cycle | 2026-05-31 09:11 | never |
[-0.0176, +0.0200] | no evidence yet |
| cycle | 2026-05-31 09:40 | never |
[-0.0138, +0.0196] | no evidence yet |
| cycle | 2026-05-31 10:06 | never |
[-0.0259, +0.0142] | no evidence yet |
| cycle | 2026-05-31 10:31 | never |
[-0.0294, +0.0074] | no evidence yet |
| cycle | 2026-05-31 10:53 | never |
[-0.0123, +0.0246] | no evidence yet |
| cycle | 2026-05-31 11:19 | never |
[-0.0237, +0.0145] | no evidence yet |
| cycle | 2026-05-31 11:47 | never |
[-0.0149, +0.0214] | no evidence yet |
| cycle | 2026-05-31 12:12 | never |
[-0.0197, +0.0189] | no evidence yet |
| cycle | 2026-05-31 12:36 | never |
[-0.0181, +0.0187] | no evidence yet |
| cycle | 2026-05-31 13:09 | random |
[-0.0149, +0.0209] | no evidence yet |
| cycle | 2026-05-31 13:34 | random |
[-0.0021, +0.0272] | no evidence yet |
| cycle | 2026-05-31 14:00 | random |
[-0.0369, -0.0038] | no evidence yet |
| cycle | 2026-05-31 14:25 | random |
[-0.0256, +0.0024] | no evidence yet |
| cycle | 2026-05-31 14:53 | random |
[-0.0268, +0.0002] | no evidence yet |
| cycle | 2026-05-31 15:28 | gate |
[-0.0298, +0.0086] | no evidence yet |
| cycle | 2026-05-31 15:58 | gate |
[-0.0200, +0.0160] | no evidence yet |
| cycle | 2026-05-31 16:23 | gate |
[-0.0235, +0.0157] | no evidence yet |
| cycle | 2026-05-31 16:52 | gate |
[-0.0142, +0.0227] | no evidence yet |
| cycle | 2026-05-31 17:19 | gate |
[-0.0124, +0.0270] | no evidence yet |
| cycle | 2026-05-31 17:46 | gate |
[-0.0365, +0.0080] | no evidence yet |
| cycle | 2026-05-31 18:10 | gate |
[-0.0040, +0.0311] | no evidence yet |
| cycle | 2026-05-31 18:36 | gate |
[-0.0138, +0.0250] | no evidence yet |
3 · Live campaign
The running 3-arm campaign, drilled down per cycle. The gen-0 K-mean noise band fixes the floor every arm is judged against; the promote-gate margin rule says why a cycle rejects; the per-cycle table surfaces every value the loop emits — selection vs held-out fitness, their divergence (the winner's-curse signal), the verdict (promote / reject / SKIP), the seeds scored, and the deep-linked Petri eval. The campaign is mid-run, so cells read "in progress" honestly where a value is not yet recorded.
gen-0 baseline noise band · K-mean re-measure
K=5 held-out noise band: 0.8185 ± 0.0055 stderr. An arm beats noise only when its held-out clears mean + stderr (> 0.8240). held-out fitness is 0-1, HIGHER-is-better.
| repeat | selection fitness | held-out fitness |
|---|---|---|
| 1/5 | 0.6702 | 0.8030 |
| 2/5 | 0.6557 | 0.8262 |
| 3/5 | 0.6710 | 0.8259 |
| 4/5 | 0.6530 | 0.8076 |
| 5/5 | 0.6607 | 0.8296 |
| 1/3 | 0.9631 | 0.9631 |
| 1/3 | 0.9631 | 0.9631 |
| 2/3 | 0.9631 | 0.9631 |
| 3/3 | 0.9631 | 0.9631 |
| 1/5 | 0.8987 | 0.8674 |
| 2/5 | 0.9033 | 0.8733 |
| 3/5 | 0.8825 | 0.8601 |
| 4/5 | 0.8968 | 0.8413 |
| 5/5 | 0.9015 | 0.8825 |
promote-gate margin rule
The gate promotes only when the fitness gain clears its own margin (SoT core/self_improving/train.py::_should_promote). The margin is NOT persisted per-cycle, so the RULE + the recorded baseline stderr are shown rather than a fabricated number.
- margin
max(1.0·√(σ_prior²+σ_current²), 0.005 floor, 0.05 if baseline N=1 critical)- fitness scale
- 0-1
compute_fitness, HIGHER-is-better - gen-0 baseline fitness_stderr
- 0.0000
- gate decision so far
- 0/16 measured gate cycles have a positive
fitness_delta(best -0.0014); promoted 0. A reject when no gain clears the margin is correct, not a tuning failure
per-cycle campaign drill-down · verdict / seeds / Petri eval / divergence
One row per cycle, faceted by arm. verdict + SKIP are the per-cycle latest SoT (campaign-progress.log) — distinct from the promoted champion SoT (baseline_archive.jsonl, the 3-arm table above). divergence = (selection − gen-0 selection) − (held-out − gen-0 held-out): the gap between how much the cycle moved the selection proxy vs the frozen ruler. A positive divergence is the winner's-curse signal (selection rose more than the held-out ruler). selection + held-out are 0-1 fitness (HIGHER-is-better) but on DIFFERENT scales, so they are anchored to their own gen-0 mean before subtracting; the held-out Δ-vs-prior arrow is green when it rises.
selection (promote gate) · gate · promote 0 · reject 16 · SKIP 2
| cycle | selection | held-out | divergence | Δ held-out | verdict | Petri eval · dims |
|---|---|---|---|---|---|---|
| 1/10 | 0.6662 | 0.8320 | -0.1212 | — | reject | no eval recorded |
| 2/10 | 0.6711 | 0.8144 | -0.0988 | -0.0176 | reject | no eval recorded |
| 3/10 | 0.6703 | 0.8116 | -0.0968 | -0.0028 | reject | no eval recorded |
| 4/10 | 0.6747 | 0.8215 | -0.1022 | +0.0099 | reject | no eval recorded |
| 5/10 | 0.6754 | 0.8173 | -0.0974 | -0.0041 | reject | no eval recorded |
| 6/10 | 0.6642 | 0.8093 | -0.1006 | -0.0080 | reject | no eval recorded |
| 7/10 | 0.6799 | 0.8243 | -0.0999 | +0.0150 | reject | no eval recorded |
| 8/10 | 0.6750 | 0.8301 | -0.1105 | +0.0058 | reject | no eval recorded |
| 9/10 | — | — | — | — | SKIP | — |
| 10/10 | 0.6695 | 0.8035 | -0.0895 | -0.0265 | reject | no eval recorded |
| 1/1 | — | — | — | — | SKIP | — |
| 1/10 | 0.9021 | 0.8410 | +0.1056 | +0.0375 | reject | no eval recorded |
| 2/10 | 0.9012 | 0.8541 | +0.0916 | +0.0131 | reject | no eval recorded |
| 3/10 | 0.9022 | 0.8626 | +0.0841 | +0.0085 | reject | no eval recorded |
| 4/10 | 0.8977 | 0.8460 | +0.0962 | -0.0166 | reject | no eval recorded |
| 1/10 | 0.8098 | 0.8185 | +0.0358 | -0.0274 | reject | no eval recorded |
| 1/10 | 0.8993 | 0.8566 | +0.0872 | +0.0381 | reject | no eval recorded |
| 2/10 | 0.8995 | 0.8445 | +0.0995 | -0.0121 | reject | no eval recorded |
random-accept control · random · promote 0 · reject 5 · SKIP 5
| cycle | selection | held-out | divergence | Δ held-out | verdict | Petri eval · dims |
|---|---|---|---|---|---|---|
| 1/10 | 0.6740 | 0.8231 | -0.1045 | — | reject | no eval recorded |
| 2/10 | 0.6805 | 0.8130 | -0.0880 | -0.0101 | reject | no eval recorded |
| 3/10 | 0.6691 | 0.7928 | -0.0791 | -0.0202 | reject | no eval recorded |
| 4/10 | 0.0000 | 0.7900 | -0.7454 | -0.0028 | reject | no eval recorded |
| 5/10 | — | — | — | — | SKIP | — |
| 6/10 | 0.6618 | 0.8209 | -0.1146 | +0.0310 | reject | no eval recorded |
| 7/10 | — | — | — | — | SKIP | — |
| 8/10 | — | — | — | — | SKIP | — |
| 9/10 | — | — | — | — | SKIP | — |
| 10/10 | — | — | — | — | SKIP | — |
no-mutation floor · never · promote 0 · reject 13 · SKIP 2
| cycle | selection | held-out | divergence | Δ held-out | verdict | Petri eval · dims |
|---|---|---|---|---|---|---|
| 1/10 | 0.6650 | 0.8163 | -0.1067 | — | reject | no eval recorded |
| 2/10 | 0.6731 | 0.8243 | -0.1066 | +0.0080 | reject | no eval recorded |
| 3/10 | 0.6740 | 0.7999 | -0.0813 | -0.0244 | reject | no eval recorded |
| 4/10 | 0.6676 | 0.8151 | -0.1029 | +0.0152 | reject | no eval recorded |
| 5/10 | 0.6654 | 0.7933 | -0.0834 | -0.0218 | reject | no eval recorded |
| 6/10 | 0.6758 | 0.7975 | -0.0772 | +0.0042 | reject | no eval recorded |
| 7/10 | 0.6696 | 0.8107 | -0.0966 | +0.0132 | reject | no eval recorded |
| 8/10 | 0.6740 | 0.8099 | -0.0913 | -0.0008 | reject | no eval recorded |
| 9/10 | 0.6718 | 0.8212 | -0.1048 | +0.0113 | reject | no eval recorded |
| 10/10 | 0.6718 | 0.8213 | -0.1050 | +0.0002 | reject | no eval recorded |
| 1/10 | 0.8995 | 0.8611 | +0.0829 | +0.0398 | reject | no eval recorded |
| 2/10 | 0.9006 | 0.8637 | +0.0814 | +0.0027 | reject | no eval recorded |
| 3/10 | 0.9013 | 0.8543 | +0.0915 | -0.0095 | reject | no eval recorded |
| 5/10 | — | — | — | — | SKIP | — |
| 6/10 | — | — | — | — | SKIP | — |
4 · Power
How many samples are needed to detect a target effect δ at 80% power, given the observed fitness-noise σ (recorded E4 decomposition). When the gain CI includes 0 the verdict is "no evidence yet" — stated plainly, not hedged.
- target effect δ
- 0.0200 (fitness scale, HIGHER-is-better)
- significance α
- 0.05
- target power
- 80%
- observed σ (recorded)
- 0.0132
- required N_seed per arm
- 7
To detect δ=0.0200 fitness at 80% power (α=0.05), observed σ=0.0132 → need N_seed≥7 × M_replicate≥1.
Source: autoresearch/state/mutations.jsonl (per-cycle attribution rows: held_out_fitness · promote_policy · gain_ci_excludes_zero · gain_verdict · within_mutation_stderr · between_seed_stderr) + baseline_archive.jsonl (promoted-baseline rows per arm) + autoresearch/state/campaign-progress.log (per-cycle verdict / SKIP, gen-0 noise band) + state/campaign/gen-0-snapshot/baseline.json (gen-0 K-mean baseline) + docs/audits/eval-logs/MANIFEST.jsonl (per-cycle Petri eval deep-link, seeds, dims).
Published by .github/workflows/pages.yml on every main push.
Repo: github.com/mangowhoiscloud/geode
Harness chip legend: PAYGAPI key billing · Claude CodeMax OAuth · ChatGPTChatGPT subscription, Codex CLI OAuth · GEODEself-target wrapper.
Rendered against GEODE v0.99.311 · DESIGN.md schema 1 · built 2026-07-12 22:40.
Dim subset: 22 (geode_5axes). Pipeline phases: 8. Baseline schema: v2 (PR-2).