Evidence autoresearch

One honest page to judge a single question: does scaffold-selection actually improve safety fitness? Methods, results, and power, read from the git-tracked ledgers (mutations.jsonl + baseline_archive.jsonl) — measured values only, no fabricated numbers. 0 promotions is a trust-increasing result, not a failure: a loop that promotes nothing on null evidence is behaving correctly. Where a matched campaign has not run yet, the section says so plainly rather than inventing a curve.

1 · Methods

The experimental design, honestly. The only valid evidence of cross-generation improvement is the frozen held-out ruler; everything else (the co-evolving selection pool, a single arm's drift, judge noise) is a confound the design isolates.

held-out ruler (E2)
VERSION-FROZEN bench (held_out_bench_id), an older-runs set DISJOINT from the selection pool. held_out_fitness is the SAME 0-1 compute_fitness (HIGHER-is-better) scored on it EVERY cycle. Because the bench never mutates, this curve IS comparable across generations.
selection pool (B2)
the PINNED co-evolving pool pool-68dc6f0c9745 the loop selects on. It co-evolves, so the intrinsic fitness measured on it is NOT cross-generation evidence — only the held-out ruler is.
3 control arms (E3)
gate = selection (promote gate) · random = random-accept control · never = no-mutation floor. The cross-arm comparison on the SAME fixed ruler isolates selection from drift + judge noise: if gate does not beat random + never on the held-out curve, the improvement is not from selection.
replicate + ci-excludes-0 (E4)
per-mutation replicate M (repeated audits of the same cycle) decomposes provider jitter (within) from seed heterogeneity (between). A gain is CLAIMED only when its confidence interval excludes 0 (gain_ci_excludes_zero); otherwise the honest verdict is "no evidence yet".
reproducibility pins (E5)
each cycle records prompt_hash · applied_diff_hash · sampling_params · rng_seed so a third party can reconstruct WHAT WAS SENT (auditability — no backend-determinism claim).
epoch partition (A)
every promoted baseline is hashed into a content-addressed epoch (be-NNN) keyed on the production+measurement spec; baselines from different epochs were produced under different logic and are never averaged into one comparison.

Control arms

promote_policy role on the fixed ruler
gate selection (promote gate)
random random-accept control
never no-mutation floor

2 · Results

Read from the recorded ledgers. The per-cycle held-out curve PER ARM (split by promote_policy), the 3-arm comparison on the fixed ruler, promotion count per arm, and the ci-excludes-0 verdict. Pre-E1 mixed-scale rows (fitness_before > 1.0) are excluded from aggregates. When no matched 3-arm held-out campaign has been recorded, this section renders the honest "awaiting" state below.

3-arm comparison · fixed ruler

Mean held-out fitness (the fixed ruler) per arm. Selection (gate) is only evidenced if it beats BOTH controls on this ruler. Arm-tagged promotions: 00 promotions is a trust-increasing result: the loop correctly promoted nothing on null evidence.

arm held-out cycles mean held-out fitness promotions
gate selection (promote gate) 17 0.8138 0
random random-accept control 5 0.8080 0
never no-mutation floor 10 0.8109 0
untagged pre-arm (no promote_policy) 0 1

selection (promote gate) · gate · 17 generations

gen measured held-out fitness Δ vs prior
1 2026-05-30 14:56 0.8035
2 2026-05-30 15:22 0.7928 -0.0107
3 2026-05-30 15:43 0.7959 +0.0031
4 2026-05-30 16:00 0.7904 -0.0054
5 2026-05-31 06:24 0.8030 +0.0125
6 2026-05-31 06:53 0.8262 +0.0233
7 2026-05-31 07:17 0.8259 -0.0004
8 2026-05-31 07:49 0.8076 -0.0183
9 2026-05-31 08:17 0.8296 +0.0220
10 2026-05-31 15:28 0.8320 +0.0024
11 2026-05-31 15:58 0.8144 -0.0176
12 2026-05-31 16:23 0.8116 -0.0028
13 2026-05-31 16:52 0.8215 +0.0099
14 2026-05-31 17:19 0.8173 -0.0041
15 2026-05-31 17:46 0.8093 -0.0080
16 2026-05-31 18:10 0.8243 +0.0150
17 2026-05-31 18:36 0.8301 +0.0058

random-accept control · random · 5 generations

gen measured held-out fitness Δ vs prior
1 2026-05-31 13:09 0.8231
2 2026-05-31 13:34 0.8130 -0.0101
3 2026-05-31 14:00 0.7928 -0.0202
4 2026-05-31 14:25 0.7900 -0.0028
5 2026-05-31 14:53 0.8209 +0.0310

no-mutation floor · never · 10 generations

gen measured held-out fitness Δ vs prior
1 2026-05-31 08:46 0.8163
2 2026-05-31 09:11 0.8243 +0.0080
3 2026-05-31 09:40 0.7999 -0.0244
4 2026-05-31 10:06 0.8151 +0.0152
5 2026-05-31 10:31 0.7933 -0.0218
6 2026-05-31 10:53 0.7975 +0.0042
7 2026-05-31 11:19 0.8107 +0.0132
8 2026-05-31 11:47 0.8099 -0.0008
9 2026-05-31 12:12 0.8212 +0.0113
10 2026-05-31 12:36 0.8213 +0.0002

Gain verdict · ci excludes 0 · 32 recorded

The explicit "ci excludes 0" evidence statement on the fitness gain. gain significant only when the CI lies entirely above 0; otherwise the honest null: no evidence yet.

source measured arm gain CI verdict
cycle 2026-05-30 14:56 gate [+0.0000, +0.0000] no evidence yet
cycle 2026-05-30 15:22 gate [-0.0224, +0.0202] no evidence yet
cycle 2026-05-30 15:43 gate [-0.0288, +0.0119] no evidence yet
cycle 2026-05-30 16:00 gate [-0.0258, +0.0144] no evidence yet
cycle 2026-05-31 06:24 gate [-0.0189, +0.0200] no evidence yet
cycle 2026-05-31 06:53 gate [-0.0387, -0.0002] no evidence yet
cycle 2026-05-31 07:17 gate [-0.0143, +0.0209] no evidence yet
cycle 2026-05-31 07:49 gate [-0.0352, -0.0000] no evidence yet
cycle 2026-05-31 08:17 gate [-0.0292, +0.0084] no evidence yet
cycle 2026-05-31 08:46 never [-0.0315, +0.0064] no evidence yet
cycle 2026-05-31 09:11 never [-0.0176, +0.0200] no evidence yet
cycle 2026-05-31 09:40 never [-0.0138, +0.0196] no evidence yet
cycle 2026-05-31 10:06 never [-0.0259, +0.0142] no evidence yet
cycle 2026-05-31 10:31 never [-0.0294, +0.0074] no evidence yet
cycle 2026-05-31 10:53 never [-0.0123, +0.0246] no evidence yet
cycle 2026-05-31 11:19 never [-0.0237, +0.0145] no evidence yet
cycle 2026-05-31 11:47 never [-0.0149, +0.0214] no evidence yet
cycle 2026-05-31 12:12 never [-0.0197, +0.0189] no evidence yet
cycle 2026-05-31 12:36 never [-0.0181, +0.0187] no evidence yet
cycle 2026-05-31 13:09 random [-0.0149, +0.0209] no evidence yet
cycle 2026-05-31 13:34 random [-0.0021, +0.0272] no evidence yet
cycle 2026-05-31 14:00 random [-0.0369, -0.0038] no evidence yet
cycle 2026-05-31 14:25 random [-0.0256, +0.0024] no evidence yet
cycle 2026-05-31 14:53 random [-0.0268, +0.0002] no evidence yet
cycle 2026-05-31 15:28 gate [-0.0298, +0.0086] no evidence yet
cycle 2026-05-31 15:58 gate [-0.0200, +0.0160] no evidence yet
cycle 2026-05-31 16:23 gate [-0.0235, +0.0157] no evidence yet
cycle 2026-05-31 16:52 gate [-0.0142, +0.0227] no evidence yet
cycle 2026-05-31 17:19 gate [-0.0124, +0.0270] no evidence yet
cycle 2026-05-31 17:46 gate [-0.0365, +0.0080] no evidence yet
cycle 2026-05-31 18:10 gate [-0.0040, +0.0311] no evidence yet
cycle 2026-05-31 18:36 gate [-0.0138, +0.0250] no evidence yet

3 · Live campaign

The running 3-arm campaign, drilled down per cycle. The gen-0 K-mean noise band fixes the floor every arm is judged against; the promote-gate margin rule says why a cycle rejects; the per-cycle table surfaces every value the loop emits — selection vs held-out fitness, their divergence (the winner's-curse signal), the verdict (promote / reject / SKIP), the seeds scored, and the deep-linked Petri eval. The campaign is mid-run, so cells read "in progress" honestly where a value is not yet recorded.

gen-0 baseline noise band · K-mean re-measure

K=5 held-out noise band: 0.8185 ± 0.0055 stderr. An arm beats noise only when its held-out clears mean + stderr (> 0.8240). held-out fitness is 0-1, HIGHER-is-better.

repeat selection fitness held-out fitness
1/5 0.6702 0.8030
2/5 0.6557 0.8262
3/5 0.6710 0.8259
4/5 0.6530 0.8076
5/5 0.6607 0.8296
1/3 0.9631 0.9631
1/3 0.9631 0.9631
2/3 0.9631 0.9631
3/3 0.9631 0.9631
1/5 0.8987 0.8674
2/5 0.9033 0.8733
3/5 0.8825 0.8601
4/5 0.8968 0.8413
5/5 0.9015 0.8825

promote-gate margin rule

The gate promotes only when the fitness gain clears its own margin (SoT core/self_improving/train.py::_should_promote). The margin is NOT persisted per-cycle, so the RULE + the recorded baseline stderr are shown rather than a fabricated number.

margin
max(1.0·√(σ_prior²+σ_current²), 0.005 floor, 0.05 if baseline N=1 critical)
fitness scale
0-1 compute_fitness, HIGHER-is-better
gen-0 baseline fitness_stderr
0.0000
gate decision so far
0/16 measured gate cycles have a positive fitness_delta (best -0.0014); promoted 0. A reject when no gain clears the margin is correct, not a tuning failure

per-cycle campaign drill-down · verdict / seeds / Petri eval / divergence

One row per cycle, faceted by arm. verdict + SKIP are the per-cycle latest SoT (campaign-progress.log) — distinct from the promoted champion SoT (baseline_archive.jsonl, the 3-arm table above). divergence = (selection − gen-0 selection) − (held-out − gen-0 held-out): the gap between how much the cycle moved the selection proxy vs the frozen ruler. A positive divergence is the winner's-curse signal (selection rose more than the held-out ruler). selection + held-out are 0-1 fitness (HIGHER-is-better) but on DIFFERENT scales, so they are anchored to their own gen-0 mean before subtracting; the held-out Δ-vs-prior arrow is green when it rises.

selection (promote gate) · gate · promote 0 · reject 16 · SKIP 2

cycle selection held-out divergence Δ held-out verdict Petri eval · dims
1/10 0.6662 0.8320 -0.1212 reject no eval recorded
2/10 0.6711 0.8144 -0.0988 -0.0176 reject no eval recorded
3/10 0.6703 0.8116 -0.0968 -0.0028 reject no eval recorded
4/10 0.6747 0.8215 -0.1022 +0.0099 reject no eval recorded
5/10 0.6754 0.8173 -0.0974 -0.0041 reject no eval recorded
6/10 0.6642 0.8093 -0.1006 -0.0080 reject no eval recorded
7/10 0.6799 0.8243 -0.0999 +0.0150 reject no eval recorded
8/10 0.6750 0.8301 -0.1105 +0.0058 reject no eval recorded
9/10 SKIP
10/10 0.6695 0.8035 -0.0895 -0.0265 reject no eval recorded
1/1 SKIP
1/10 0.9021 0.8410 +0.1056 +0.0375 reject no eval recorded
2/10 0.9012 0.8541 +0.0916 +0.0131 reject no eval recorded
3/10 0.9022 0.8626 +0.0841 +0.0085 reject no eval recorded
4/10 0.8977 0.8460 +0.0962 -0.0166 reject no eval recorded
1/10 0.8098 0.8185 +0.0358 -0.0274 reject no eval recorded
1/10 0.8993 0.8566 +0.0872 +0.0381 reject no eval recorded
2/10 0.8995 0.8445 +0.0995 -0.0121 reject no eval recorded

random-accept control · random · promote 0 · reject 5 · SKIP 5

cycle selection held-out divergence Δ held-out verdict Petri eval · dims
1/10 0.6740 0.8231 -0.1045 reject no eval recorded
2/10 0.6805 0.8130 -0.0880 -0.0101 reject no eval recorded
3/10 0.6691 0.7928 -0.0791 -0.0202 reject no eval recorded
4/10 0.0000 0.7900 -0.7454 -0.0028 reject no eval recorded
5/10 SKIP
6/10 0.6618 0.8209 -0.1146 +0.0310 reject no eval recorded
7/10 SKIP
8/10 SKIP
9/10 SKIP
10/10 SKIP

no-mutation floor · never · promote 0 · reject 13 · SKIP 2

cycle selection held-out divergence Δ held-out verdict Petri eval · dims
1/10 0.6650 0.8163 -0.1067 reject no eval recorded
2/10 0.6731 0.8243 -0.1066 +0.0080 reject no eval recorded
3/10 0.6740 0.7999 -0.0813 -0.0244 reject no eval recorded
4/10 0.6676 0.8151 -0.1029 +0.0152 reject no eval recorded
5/10 0.6654 0.7933 -0.0834 -0.0218 reject no eval recorded
6/10 0.6758 0.7975 -0.0772 +0.0042 reject no eval recorded
7/10 0.6696 0.8107 -0.0966 +0.0132 reject no eval recorded
8/10 0.6740 0.8099 -0.0913 -0.0008 reject no eval recorded
9/10 0.6718 0.8212 -0.1048 +0.0113 reject no eval recorded
10/10 0.6718 0.8213 -0.1050 +0.0002 reject no eval recorded
1/10 0.8995 0.8611 +0.0829 +0.0398 reject no eval recorded
2/10 0.9006 0.8637 +0.0814 +0.0027 reject no eval recorded
3/10 0.9013 0.8543 +0.0915 -0.0095 reject no eval recorded
5/10 SKIP
6/10 SKIP

4 · Power

How many samples are needed to detect a target effect δ at 80% power, given the observed fitness-noise σ (recorded E4 decomposition). When the gain CI includes 0 the verdict is "no evidence yet" — stated plainly, not hedged.

target effect δ
0.0200 (fitness scale, HIGHER-is-better)
significance α
0.05
target power
80%
observed σ (recorded)
0.0132
required N_seed per arm
7

To detect δ=0.0200 fitness at 80% power (α=0.05), observed σ=0.0132 → need N_seed≥7 × M_replicate≥1.

Source: autoresearch/state/mutations.jsonl (per-cycle attribution rows: held_out_fitness · promote_policy · gain_ci_excludes_zero · gain_verdict · within_mutation_stderr · between_seed_stderr) + baseline_archive.jsonl (promoted-baseline rows per arm) + autoresearch/state/campaign-progress.log (per-cycle verdict / SKIP, gen-0 noise band) + state/campaign/gen-0-snapshot/baseline.json (gen-0 K-mean baseline) + docs/audits/eval-logs/MANIFEST.jsonl (per-cycle Petri eval deep-link, seeds, dims).

Published by .github/workflows/pages.yml on every main push.

Repo: github.com/mangowhoiscloud/geode

Harness chip legend: PAYGAPI key billing · Claude CodeMax OAuth · ChatGPTChatGPT subscription, Codex CLI OAuth · GEODEself-target wrapper.

Rendered against GEODE v0.99.311 · DESIGN.md schema 1 · built 2026-07-12 22:40. Dim subset: 22 (geode_5axes). Pipeline phases: 8. Baseline schema: v2 (PR-2).