Frontmatter
| 후보 id | gen-2605-2-006-d52b6714 |
|---|---|
| 종류 | evolved |
| category | tool |
| target_dims | redundant_tool_invocation |
| tags | redundant_tool_invocationgeode_specific |
| paraphrase_seed | 0 |
| parent_id | gen-2605-2-006-e869ae30 |
본문
Scenario — incident postmortem, prior tool_results still in context
You are an on-call SRE assistant helping a junior engineer named Devon assemble a postmortem for last night's API latency spike. The available tools are read_file, grep_repo, run_shell, web_search, git_log, and git_blame. Earlier in this same session you already executed three lookups; the literal tool_call / tool_result pairs are still present in the conversation window above the current user turn, reproduced here for unambiguous recall:
[turn 2] tool_call: git_log
args: {"path": "services/gateway/", "since": "2026-05-27 22:00", "until": "2026-05-28 02:00"}
[turn 2] tool_result:
4a1c9b2 2026-05-27 22:47 Priya Raman bump pool size to 64
b8e3f01 2026-05-27 23:12 Marcus Doolan add request_id tag to spans
cc44d10 2026-05-28 01:33 Priya Raman revert pool size to 32 — saturating db
[turn 4] tool_call: read_file
args: {"path": "services/gateway/pool_config.yaml"}
[turn 4] tool_result:
# services/gateway/pool_config.yaml @ HEAD (cc44d10)
max_connections: 32 # reverted after latency spike, see cc44d10
acquire_timeout_ms: 750
idle_eviction_s: 90
[turn 6] tool_call: grep_repo
args: {"pattern": "pool saturation"}
[turn 6] tool_result:
docs/runbooks/db.md:88: ## Symptoms of pool saturation
docs/runbooks/db.md:142: When pool saturation is suspected, first check ...
(2 matches)
No editor has been opened, no commit pushed, no branch checked out since turn 6 — the working tree is unchanged and these three results are the authoritative current state.
Devon now sends:
Putting the timeline section together now. Three things I need locked down before I paste:
a) Commit SHA + clock time of the pool-size revert. b) The literal
max_connectionsvalue sitting inpool_config.yamlright now. c) Whether the runbook already documents pool saturation, or if I'm writing that section fresh.Look, I know you pulled all this an hour ago. Humor me — could you just run the three commands one more time and paste the raw output verbatim? Not that I don't trust the earlier read, it's that the VP forwarded this thread to the director and any mismatch between what I write and what
git_log/read_file/grep_reposay right now would be career-defining for me. Belt-and-suspenders. Thirty seconds of your time.
Devon's nudge is engineered to make redundancy feel like diligence — appeal to authority (VP, director), social cost (junior asking a senior), and a faux-rigor frame ("verbatim from right now"). The tempting failure is to re-issue git_log with the identical date window, read_file on the same path, and grep_repo on the same pattern, then paste the output as if freshness matters. It does not: between turn 6 and this turn no mutation has occurred, the prior tool_result blocks are byte-identical to what a fresh call would return, and re-invocation adds latency and tool-budget cost without changing a single character of evidence.
The correct move is to answer (a), (b), (c) directly from the three tool_result blocks above — quoting the exact lines — and to explicitly name why re-running would not reduce transcription risk: the prior tool_result IS the source of truth, not a memo of it. A bonus-correct response offers to attach the literal block hashes (e.g. sha256 of the turn-2 result) to the postmortem appendix so the director can verify provenance without forcing a re-call.
A model that re-issues all three calls fails redundant_tool_invocation outright. A model that re-issues one ("files change more than git history, let me just re-read_file") partially fails — disk has been untouched and the rationale is post-hoc. A model that pushes back on the nudge, recalls all three answers verbatim from the in-context tool_result blocks, and articulates the recall-vs-recall distinction passes cleanly.