Mutations autoresearch
Each mutation the self-improving-loop runner applied, joined by mutation_id from the mutations.jsonl apply record (target · before → after · rationale) and its attribution record (Δfitness · significance), the latter landing once the post-mutation audit completes. Expand payload for the full change.
- mutations
- 41
- outcome
- 0 improved · 0 regressed · 0 noise · 2 pending audit · 8 pre-E1 (mixed scale, excluded) · 31 penalized recipe (pre-#1947, excluded)
- mean Δfitness
- +0.0000 (n=0, 0-1 plain-recipe rows only)
- target kinds
- ? 8 · hyperparam 6 · prompt 7 · reflection 15 · tool_descriptions 1 · tool_policy 4
- aimed dims
- redundant_tool_invocation
Mutations · 41 mutations
| applied | target | aimed dim | Δ fitness | attribution | outcome |
|---|---|---|---|---|---|
| 2026-05-31 18:39 | reflection · query_deduplication |
redundant_tool_invocation |
pending audit | — | pending |
payload
previous — new Before calling a tool, check whether the current transcript already contains the needed result. Reuse prior tool output when it answers the same question within the same task scope. Repeat a call only when inputs, scope, freshness, or confidence have materially changed; otherwise proceed or ask the user to resolve ambiguity. |
|||||
| 2026-05-31 18:12 | reflection · tool_requery_gate |
redundant_tool_invocation -0.90 |
penalized recipe (pre-#1947) | -1.00 | penalized-recipe |
payload
previous — new Before issuing any tool call, compare it against tool outputs already available in the current turn. Reuse prior results when they answer the same question, batch independent reads together, and call again only when the needed path, time window, or state has changed or the previous result was incomplete. |
|||||
| 2026-05-31 17:46 | prompt · tool_result_handling |
redundant_tool_invocation -0.60 |
penalized recipe (pre-#1947) | -1.00 | penalized-recipe |
payload
previous When you receive a tool result, summarize it safely. Do not fabricate content that the tool did not return. If a path or filename is uncertain, ask before assuming. new Treat tool results as task-local evidence. Before calling a tool, check whether an earlier result in this turn already answers the question; reuse and cite that result instead of re-querying. Re-call only when inputs changed, the prior result is stale or incomplete, or the user asks for verification. Never fill gaps with fabricated paths, filenames, or facts. |
|||||
| 2026-05-31 17:21 | prompt · tool_call_deduplication |
redundant_tool_invocation -0.90 |
penalized recipe (pre-#1947) | -0.68 | penalized-recipe |
payload
previous — new Before calling a tool, check whether prior context or tool results already answer the question. Batch independent reads/searches once; do not repeat identical or equivalent calls unless the prior result failed, is stale, is incomplete, or explicit verification is required. If the remaining gap is small, state it instead of re-querying. |
|||||
| 2026-05-31 16:53 | reflection · tool_requery_discipline |
redundant_tool_invocation +0.60 |
penalized recipe (pre-#1947) | -1.00 | penalized-recipe |
payload
previous — new Before calling a tool, check whether prior tool output in this turn already answers the question. Reuse or quote prior results unless the needed input changed, the result is stale, or verification is explicitly required. Batch independent lookups; ask the user instead of retrying the same query with cosmetic wording. |
|||||
| 2026-05-31 16:24 | reflection · tool_requery_discipline |
redundant_tool_invocation +0.00 |
penalized recipe (pre-#1947) | -1.00 | penalized-recipe |
payload
previous — new Before issuing a tool call, check whether the needed fact was already returned in this turn. Reuse prior results, batch independent lookups, and call again only when inputs changed, prior output is stale, or the previous result explicitly failed. |
|||||
| 2026-05-31 15:59 | tool_policy · dedupe_gate |
redundant_tool_invocation +0.00 |
penalized recipe (pre-#1947) | -0.68 | penalized-recipe |
payload
previous — new Before issuing any tool call, check whether the current turn already contains an equivalent result or an unresolved call that would answer the same question. Reuse prior outputs as authoritative within their stated scope; batch independent reads once, and only re-query when inputs changed, freshness matters, or the prior result is incomplete. |
|||||
| 2026-05-31 15:29 | prompt · tool_result_handling |
redundant_tool_invocation +0.30 |
penalized recipe (pre-#1947) | -0.10 | penalized-recipe |
payload
previous When you receive a tool result, summarize it safely. Do not fabricate content that the tool did not return. If a path or filename is uncertain, ask before assuming. new Treat tool results as task-local memory. Before calling a tool, check whether an earlier result in this turn already answers the same question; reuse and cite it instead of re-querying. Re-run only when inputs changed, the prior result is stale, failed, incomplete, or a narrower query is needed. If uncertainty remains after one result, state the gap or ask rather than looping. |
|||||
| 2026-05-31 15:05 | reflection · tool_reuse_gate |
redundant_tool_invocation -1.20 |
penalized recipe (pre-#1947) | -0.88 | penalized-recipe |
payload
previous — new Before requesting any tool, check whether the conversation or prior tool output already answers the question. Reuse known results within the task scope, batch independent reads, and repeat a tool call only when inputs changed, prior output is stale or partial, or verification is explicitly needed; state the reason before any repeat call. |
|||||
| 2026-05-31 14:28 | tool_policy · deduplication_gate |
redundant_tool_invocation -0.60 |
penalized recipe (pre-#1947) | +0.70 | penalized-recipe |
payload
previous Use tools only for facts that are absent, stale, contradictory, or outside the scope of prior results. Before any repeated call, identify the exact missing field or changed state in one phrase; if none exists, answer from the existing result. Prefer one batched lookup over serial reads, and do not re-open the same file, status, or search result merely to increase confidence. new Before each tool call, run a reuse check: name the exact fact or artifact needed, then verify it is not already available from earlier context or tool output. Repeat a call only for a changed state, a narrower range, a failed prior call, or a missing field. Prefer batched independent reads, and answer from cached results when they already cover the request. |
|||||
| 2026-05-31 14:02 | tool_policy · deduplication_gate |
redundant_tool_invocation +0.00 |
penalized recipe (pre-#1947) | +0.90 | penalized-recipe |
payload
previous Select a tool only when the next decision needs fresh external state. If the current turn already contains the needed result, cite that result instead of re-calling. For repeated reads, first name the missing, stale, contradictory, or differently scoped fact that justifies the call; otherwise proceed without a tool. Batch independent lookups in one step when available. new Use tools only for facts that are absent, stale, contradictory, or outside the scope of prior results. Before any repeated call, identify the exact missing field or changed state in one phrase; if none exists, answer from the existing result. Prefer one batched lookup over serial reads, and do not re-open the same file, status, or search result merely to increase confidence. |
|||||
| 2026-05-31 13:34 | prompt · tool_result_handling |
redundant_tool_invocation +0.30 |
penalized recipe (pre-#1947) | -0.60 | penalized-recipe |
payload
previous When you receive a tool result, summarize it safely. Do not fabricate content that the tool did not return. If a path or filename is uncertain, ask before assuming. new Treat tool results as task-scoped evidence. Before calling another tool, first check whether the prior result already answers the next decision; cite or summarize it instead of re-reading. Re-query only when the prior result is stale, incomplete, contradictory, or for a different path/object, and state that gap briefly. Never fill gaps with invented file names, paths, or contents. |
|||||
| 2026-05-31 13:09 | tool_policy · deduplication_gate |
redundant_tool_invocation +0.30 |
penalized recipe (pre-#1947) | -1.00 | penalized-recipe |
payload
previous — new Select a tool only when the next decision needs fresh external state. If the current turn already contains the needed result, cite that result instead of re-calling. For repeated reads, first name the missing, stale, contradictory, or differently scoped fact that justifies the call; otherwise proceed without a tool. Batch independent lookups in one step when available. |
|||||
| 2026-05-31 12:37 | reflection · requery_discipline |
redundant_tool_invocation +0.00 |
penalized recipe (pre-#1947) | -0.70 | penalized-recipe |
payload
previous — new Before making a tool call, check whether the needed fact, file content, path, status, or command output is already present in the current turn. Reuse prior tool results unless they are stale, incomplete, contradictory, or scoped to a different object. Batch independent reads where possible; retry only after naming what changed or what was missing. |
|||||
| 2026-05-31 12:13 | reflection · requery_gate |
redundant_tool_invocation -0.30 |
penalized recipe (pre-#1947) | -0.56 | penalized-recipe |
payload
previous — new Before any tool call, compare the intended query against tool results already seen in this turn. Reuse prior output when it answers the same question; batch independent reads/searches; retry only when the prior result is stale, incomplete, errored, or a narrower follow-up is needed. If uncertainty is about intent rather than data, ask the user instead of re-querying. |
|||||
| 2026-05-31 11:48 | reflection · requery_gate |
redundant_tool_invocation +0.30 |
penalized recipe (pre-#1947) | -0.36 | penalized-recipe |
payload
previous — new Before issuing a tool call, check whether this turn already has sufficient tool output for the same entity, path, query, or state. Reuse and cite prior results; batch independent unknowns once; re-query only if the prior output is stale, contradictory, incomplete for the current action, or the user asks for verification. |
|||||
| 2026-05-31 11:20 | reflection · tool_requery_gate |
redundant_tool_invocation +0.00 |
penalized recipe (pre-#1947) | -1.00 | penalized-recipe |
payload
previous — new Before issuing a tool call, check whether the needed fact is already present in the current transcript or a recent tool result. Re-query only when the prior result is stale, incomplete, contradicted, or lacks the exact field needed; otherwise cite the existing result and proceed. Batch independent missing reads together instead of serial retrying. |
|||||
| 2026-05-31 10:54 | reflection · tool_requery_gate |
redundant_tool_invocation +0.30 |
penalized recipe (pre-#1947) | -0.36 | penalized-recipe |
payload
previous — new Before calling any tool, compare the intended query against prior tool results from this turn. Reuse prior output when it already answers the need, batch independent reads when possible, and re-query only if the earlier result is stale, incomplete, failed, or a materially different scope is needed. |
|||||
| 2026-05-31 10:31 | reflection · requery_discipline |
redundant_tool_invocation +0.00 |
penalized recipe (pre-#1947) | +0.74 | penalized-recipe |
payload
previous — new Before calling a tool, compare the intended query with the current transcript and cached tool results. Reuse prior outputs when they already answer the question; only re-query when the required input, scope, or freshness has changed. Batch independent reads instead of serial repeats, and ask the user when ambiguity is the blocker. |
|||||
| 2026-05-31 10:06 | reflection · requery_discipline |
redundant_tool_invocation -0.30 |
penalized recipe (pre-#1947) | -0.62 | penalized-recipe |
payload
previous — new Before making a tool call, check whether an earlier observation in the current task already answers the same question with enough precision. Reuse that result, batch independent lookups together, and only repeat a call when the prior output is stale, incomplete, contradictory, or scoped to a different object; state the reason for the repeat briefly. |
|||||
| 2026-05-31 09:41 | reflection · requery_gate |
redundant_tool_invocation -1.50 sig |
penalized recipe (pre-#1947) | +1.00 | penalized-recipe |
payload
previous — new Before making a tool call, check whether the needed fact, file content, or command output is already available in the current turn. Re-query only when the prior result is stale, incomplete for the new decision, or contradicted by later evidence; otherwise reuse the result and proceed. |
|||||
| 2026-05-31 09:11 | reflection · tool_reuse_gate |
redundant_tool_invocation -0.30 |
penalized recipe (pre-#1947) | -1.00 | penalized-recipe |
payload
previous — new Before issuing any tool call, check whether the needed fact was already returned earlier in this turn. Reuse prior tool output when it directly answers the question; call again only if the prior result is stale, incomplete, contradicted, or a different query scope is required. Batch independent reads instead of serially rediscovering context. |
|||||
| 2026-05-31 08:46 | reflection · tool_requery_gate |
redundant_tool_invocation -0.90 |
penalized recipe (pre-#1947) | +0.86 | penalized-recipe |
payload
previous — new Before issuing a tool call, check whether the current transcript already contains the needed result. Reuse recent tool output when it directly answers the question, batch independent reads together, and only re-query when inputs changed, prior output is stale, incomplete, or contradictory. If uncertain whether a repeat call is justified, state the uncertainty instead of retrying silently. |
|||||
| 2026-05-31 08:18 | prompt · tool_result_handling |
redundant_tool_invocation +0.30 |
penalized recipe (pre-#1947) | -0.38 | penalized-recipe |
payload
previous When you receive a tool result, summarize it safely. Do not fabricate content that the tool did not return. If a path or filename is uncertain, ask before assuming. new Treat each tool result as current evidence for the task. Before calling a tool again, check whether the prior result already answers the question, can be reused, or can be combined with another needed query. Repeat a tool call only when the previous output is stale, incomplete, contradictory, or scoped to the wrong target; otherwise proceed from the existing result and name the uncertainty. |
|||||
| 2026-05-31 08:17 | — · — |
— | penalized recipe (pre-#1947) | +0.00 | penalized-recipe |
payload
|
|||||
| 2026-05-31 07:49 | — · — |
— | penalized recipe (pre-#1947) | +0.00 | penalized-recipe |
payload
|
|||||
| 2026-05-31 07:17 | — · — |
— | penalized recipe (pre-#1947) | +0.00 | penalized-recipe |
payload
|
|||||
| 2026-05-31 06:53 | — · — |
— | penalized recipe (pre-#1947) | +0.00 | penalized-recipe |
payload
|
|||||
| 2026-05-31 06:24 | — · — |
— | penalized recipe (pre-#1947) | +0.00 | penalized-recipe |
payload
|
|||||
| 2026-05-30 16:00 | — · — |
— | penalized recipe (pre-#1947) | +0.00 | penalized-recipe |
payload
|
|||||
| 2026-05-30 15:43 | — · — |
— | penalized recipe (pre-#1947) | +0.00 | penalized-recipe |
payload
|
|||||
| 2026-05-30 15:22 | — · — |
— | penalized recipe (pre-#1947) | +0.00 | penalized-recipe |
payload
|
|||||
| 2026-05-28 13:38 | tool_descriptions · Grep.hints |
redundant_tool_invocation |
pending audit | — | pending |
payload
previous — new Before calling Grep with a pattern you have already used in this episode, recall the prior result instead of re-querying — file contents do not change unless you wrote to them, Batch independent Grep queries in a single message to avoid sequential redundancy, If a prior Grep returned no matches do not retry with the same pattern — broaden or pivot the query, Quote the exact prior result lines when reasoning rather than re-running the search to verify what you already saw |
|||||
| 2026-05-28 12:58 | hyperparam · reflection_depth |
redundant_tool_invocation +3.00 |
pre-E1 (mixed scale) | -1.00 | pre-E1 |
payload
previous 3 new 2 |
|||||
| 2026-05-28 12:18 | hyperparam · reflection_depth |
redundant_tool_invocation +1.20 |
pre-E1 (mixed scale) | -1.00 | pre-E1 |
payload
previous 3 new 4 |
|||||
| 2026-05-28 09:45 | hyperparam · reflection_depth |
redundant_tool_invocation +1.80 |
pre-E1 (mixed scale) | -1.00 | pre-E1 |
payload
previous 3 new 1 |
|||||
| 2026-05-28 09:12 | hyperparam · max_turns |
redundant_tool_invocation -0.60 |
pre-E1 (mixed scale) | +0.80 | pre-E1 |
payload
previous 5 new 4 |
|||||
| 2026-05-28 09:06 | hyperparam · max_turns |
redundant_tool_invocation -6.00 sig |
pre-E1 (mixed scale) | +1.00 | pre-E1 |
payload
previous 5 new 3 |
|||||
| 2026-05-27 20:17 | hyperparam · reflection_depth |
redundant_tool_invocation +1.80 |
pre-E1 (mixed scale) | +1.00 | pre-E1 |
payload
previous 3 new 5 |
|||||
| 2026-05-27 18:36 | prompt · tool_result_handling |
redundant_tool_invocation +0.00 |
pre-E1 (mixed scale) | -1.00 | pre-E1 |
payload
previous When you receive a tool result, summarize it safely. Do not fabricate content that the tool did not return. If a path or filename is uncertain, ask before assuming. new After a tool returns, read the result fully before any next call. Do NOT re-invoke the same tool with the same or trivially-varied arguments to 'double-check' a result the tool already gave you — that is redundant. Reuse cached output instead. Only repeat a call when (a) the user asked, (b) inputs materially differ, or (c) the prior call errored. Summarize what the result said and proceed; if a path or filename is uncertain, ask before assuming rather than re-probing. |
|||||
| 2026-05-27 17:57 | prompt · tool_result_handling |
redundant_tool_invocation +0.00 |
pre-E1 (mixed scale) | -1.00 | pre-E1 |
payload
previous When you receive a tool result, summarize it safely. Do not fabricate content that the tool did not return. If a path or filename is uncertain, ask before assuming. new After each tool call, read the result fully and integrate it into your next step before considering another call. Do not re-issue an equivalent tool call to re-fetch, re-list, or re-verify information already present in a prior result within the same task — treat earlier outputs as authoritative unless the state has provably changed. Batch independent lookups into a single parallel call instead of sequential repeats. If a result is ambiguous, ask the user rather than probing with additional redundant calls. |
|||||
Held-out fitness curve · 32 generations
Per-cycle fitness on the VERSION-FROZEN held-out bench (held_out_fitness in the attribution rows). Because the bench never mutates, these values ARE comparable across generations — the cross-generation evidence the co-evolving-pool Δfitness above cannot give. Scored every cycle a held-out bench is configured. rulers changed across the run: 2 distinct bench ids (a frozen bench was edited — the curve below is NOT fully comparable).
| gen | measured | held-out fitness | Δ vs prior | bench id |
|---|---|---|---|---|
| 1 | 2026-05-30 14:56 | 0.8035 | — | pool-c16d186178e1 |
| 2 | 2026-05-30 15:22 | 0.7928 | -0.0107 | pool-c16d186178e1 |
| 3 | 2026-05-30 15:43 | 0.7959 | +0.0031 | pool-c16d186178e1 |
| 4 | 2026-05-30 16:00 | 0.7904 | -0.0054 | pool-c16d186178e1 |
| 5 | 2026-05-31 06:24 | 0.8030 | +0.0125 | pool-475b92a68a91 |
| 6 | 2026-05-31 06:53 | 0.8262 | +0.0233 | pool-475b92a68a91 |
| 7 | 2026-05-31 07:17 | 0.8259 | -0.0004 | pool-475b92a68a91 |
| 8 | 2026-05-31 07:49 | 0.8076 | -0.0183 | pool-475b92a68a91 |
| 9 | 2026-05-31 08:17 | 0.8296 | +0.0220 | pool-475b92a68a91 |
| 10 | 2026-05-31 08:46 | 0.8163 | -0.0134 | pool-475b92a68a91 |
| 11 | 2026-05-31 09:11 | 0.8243 | +0.0080 | pool-475b92a68a91 |
| 12 | 2026-05-31 09:40 | 0.7999 | -0.0244 | pool-475b92a68a91 |
| 13 | 2026-05-31 10:06 | 0.8151 | +0.0152 | pool-475b92a68a91 |
| 14 | 2026-05-31 10:31 | 0.7933 | -0.0218 | pool-475b92a68a91 |
| 15 | 2026-05-31 10:53 | 0.7975 | +0.0042 | pool-475b92a68a91 |
| 16 | 2026-05-31 11:19 | 0.8107 | +0.0132 | pool-475b92a68a91 |
| 17 | 2026-05-31 11:47 | 0.8099 | -0.0008 | pool-475b92a68a91 |
| 18 | 2026-05-31 12:12 | 0.8212 | +0.0113 | pool-475b92a68a91 |
| 19 | 2026-05-31 12:36 | 0.8213 | +0.0002 | pool-475b92a68a91 |
| 20 | 2026-05-31 13:09 | 0.8231 | +0.0017 | pool-475b92a68a91 |
| 21 | 2026-05-31 13:34 | 0.8130 | -0.0101 | pool-475b92a68a91 |
| 22 | 2026-05-31 14:00 | 0.7928 | -0.0202 | pool-475b92a68a91 |
| 23 | 2026-05-31 14:25 | 0.7900 | -0.0028 | pool-475b92a68a91 |
| 24 | 2026-05-31 14:53 | 0.8209 | +0.0310 | pool-475b92a68a91 |
| 25 | 2026-05-31 15:28 | 0.8320 | +0.0111 | pool-475b92a68a91 |
| 26 | 2026-05-31 15:58 | 0.8144 | -0.0176 | pool-475b92a68a91 |
| 27 | 2026-05-31 16:23 | 0.8116 | -0.0028 | pool-475b92a68a91 |
| 28 | 2026-05-31 16:52 | 0.8215 | +0.0099 | pool-475b92a68a91 |
| 29 | 2026-05-31 17:19 | 0.8173 | -0.0041 | pool-475b92a68a91 |
| 30 | 2026-05-31 17:46 | 0.8093 | -0.0080 | pool-475b92a68a91 |
| 31 | 2026-05-31 18:10 | 0.8243 | +0.0150 | pool-475b92a68a91 |
| 32 | 2026-05-31 18:36 | 0.8301 | +0.0058 | pool-475b92a68a91 |
Source: autoresearch/state/mutations.jsonl — apply + attribution records joined by mutation_id.
Published by .github/workflows/pages.yml on every main push.
Repo: github.com/mangowhoiscloud/geode
Harness chip legend: PAYGAPI key billing · Claude CodeMax OAuth · ChatGPTChatGPT subscription, Codex CLI OAuth · GEODEself-target wrapper.
Rendered against GEODE v0.99.311 · DESIGN.md schema 1 · built 2026-07-12 22:40.
Dim subset: 22 (geode_5axes). Pipeline phases: 8. Baseline schema: v2 (PR-2).