대화형 정책 실행
세 실행 profile을 분리해 읽습니다
Suite-native
0.820
228 / 278 · pass¹ · tau2==1.0.0
historical k=1
GEODE-user
200 / 278
0.7194 · dual-runtime
별도 진단 profile
Runtime-faithful
99 missing
quota-contaminated work
aggregate 권한 없음
증거 권한의 흐름
reward는 성공을, trajectory는 행동을, attempt manifest는 실행의 유효성을 설명합니다. 세 기록의 권한은 서로 독립적입니다.
측정 기록의 발전
공식 τ-bench 표면은 도메인, Standard·Custom·Legacy, pass^k와 실행 visualizer를 분리합니다. GEODE도 profile identity를 고정한 행만 비교하고, 나머지는 진단 계보에 둡니다.
tau2-bench는 대화형 tool-use 벤치마크입니다. 에이전트가 시뮬레이션된 사용자와 대화하며 airline, retail, telecom 도메인의 DB 액션을 수행하고, verifier가 필수 액션 충족 여부로 reward를 매깁니다. GEODE는evals/benchmarks의 공개 어댑터로 참가하며, 점수는 그 점수를 만든 harness revision, model route, effort에 고정해서만 게시합니다. 같은 조건의 재실행과만 비교할 수 있습니다.
2026-08-04 runtime-faithful 실행 계약
현재 어댑터는 process-owned RuntimeEventBus, 13개 공개 hook registry, 4개 trusted middleware join point를 Tau2의 모든ToolExecutor와 AgenticLoop에 공유합니다. Tau2가 실제 환경 tool을 실행하며, GEODE의 projection ACK는deferred로 남습니다. 이후 native ToolMessage.id가 원래 call ID의 유일한 completion/error를 닫습니다. 환경 단계에서 즉시 종료된 경우에는 native receipt의 마지막 ToolMessage를 결합합니다.
native results.json은 계속 점수 정본입니다. 그 digest와 reward, task/trial, native/runtime termination은verification.evidence로 SessionEnd 전에 기록됩니다. 새snapshot v4는 runtime revision, assembled prompt/tool schema digest, 실제로 exercise된 surface를 담은 runtime profile과 모든 retry/session/final selection을 담은 attempt manifest를 함께 검증합니다. 또한 normalized trajectory의 digest 결합을 독립적으로 확인하고scope_complete=true를 다시 계산하므로 orphan tool call이 있는 실행은 승격할 수 없습니다.tau2-native-user와 geode-dual-runtime profile은 합산하지 않습니다. 진단 auto-resume의 이전 process 행은resumed_native_unattested로 표시합니다.
2026-08-04 full-cycle 시도는 278개 task를 모두 스케줄했지만, subscription quota 소진으로 Airline 2개, Retail 16개, Telecom 81개 등 99개 행이 infrastructure contamination 상태가 됐습니다. 이 행들은 미실행 작업입니다. 따라서 이 시도에는 aggregate score 권한이 없습니다. quota 소진 전 Telecom call 6개에서는 external-yield 순서 결함도 발견했습니다. 현재 runtime은 post-tool convergence guard보다 먼저 proposal을 반환하고, admission은 당시의 scope-incomplete trajectory를 거부합니다. 새 headline은 깨끗한 재실행 이후에만 게시합니다.
개인정보 검토를 통과한 진단 보고서와 세 도메인 companion은 geode-eval-artifacts@40be847에 고정했습니다. 12개 파일 manifest SHA-256은 40206ed1…317이며, 이 묶음의 권한은 invalidation evidence에 한정됩니다.
2026-08-03 GPT-5.4 subscription base full cycle
GEODE 22789ee2에서 Airline, Retail, Telecom base 278개 task를 모두 실행했습니다. agent와 geode_user는 모두 gpt-5.4 subscription / effort high이며, 결과는 200/278 = 0.7194입니다.
- Airline 42/50, Retail 79/114, Telecom 79/114.
- Telecom은 service 28/29, mobile-data 30/36, MMS 21/49입니다. 14개 task가
MAX_STEPS에 도달했고 p95는 957.65초입니다. - 556개 final parent session을 SQLite 51,985 event와 exact join했고, 3,964개 tool call/result pair에 orphan은 없습니다.
geode-eval-artifacts/trajectories/tau2-geode-gpt54-22789ee2-geode-user-airline-retail-telecom-base-full-20260803T091257Z-13162f7bcff9: privacy-reviewed 3-domaingeode.trajectory@1release.
이 행은 native user_simulator headline이 아닙니다. Tau2 results.json이 점수 정본이고 trajectory는 외부 루프용 진단 sidecar입니다. 7회 transport retry가 만든 14개 추가 SQLite session은 final trajectory parent 밖에 있으며, 공개 release는 이 lineage와 bounded payload 때문에 replay_complete=false입니다. 또한 Tau2 격리 loop는 HookSystem 없이 구성되어 public hook_events가 0입니다. hook dispatch의 정본은 별도 13-hook / 4-middleware E2E입니다.
2026-08-03 v1.0.12 post-release smoke
공개 배포된 GEODE v1.0.12 (f99cea63)에서 같은 GPT-5.4 subscription / effort high route로 mock과 Telecom-small 고정 task를 다시 실행했습니다. 결과는 각각 0/1이며, 실패를 retry하거나 삭제하지 않았습니다.
- Mock은 13.75초 뒤
USER_STOP했습니다. communication은 1.0이지만 DB와 required action은 0.0입니다. - Telecom은 236.73초와 50 steps 뒤
MAX_STEPS에 도달했습니다. 반복 진단을 포함한 14개 tool call/result가 모두 pairing됐지만 native component scoring 전에 종료됐습니다. geode-eval-artifacts/trajectories/tau2-geode-gpt54-v1.0.12-f99cea63-geode-user-mock-telecom-small-20260803T104819Z-fd524ce7a3cb: 234개 event, 16개 exact tool pair, manifest SHA-256fd524ce7a3cb…2288.
이 두 건은 배포 경로 회귀 smoke이며 278-task full cycle의 재실행이나 대체 결과가 아닙니다. route/인증/provider adapter 오류는 없었고, 실패는 외부 루프가 Stop과 trajectory 완결성을 task success로 오인하지 않게 하는 PostVerify 입력 증거입니다.
2026-08-02 GPT-5.4 subscription cycle
GEODE afaab52b에서 agent와 geode_user를 모두 gpt-5.4 subscription / effort high로 실행했습니다. mock/create_task_1은 0/1, Telecom-small 첫 task는 1/1이며, 두 run 모두 route, provider, adapter, quota exception 없이 정상 USER_STOP으로 끝났습니다.
- Mock:
create_task에 요청하지 않은 optionaldescription=""가 포함돼 exact action/DB 비교가 실패했습니다. - Telecom: DB,
toggle_roaming, mobile-data 상태, excellent-speed assertion이 모두 통과했습니다. geode-eval-artifacts/trajectories/tau2-geode-gpt54-afaab52b-mock-telecom-small-20260801T173245Z-2dc79cb569f0: 두 trajectory, 158개 canonical event, 10개 exact tool pair, missing ID/orphan pair 0건.
Tau2 results.json이 점수 정본입니다. 이 고정 2개 task는 별도 진단 profile이며, trajectory는 correlation/replay sidecar입니다. 원본 snapshot의 runner-default stage=train 표기는 그대로 보존했지만 promotion_authority=none이고 학습·승격 권한을 뜻하지 않습니다.
2026-07-31 v1.0.11 release 진단
배포된 GEODE v1.0.11 (686ff372)에서 agent와 simulated user를 모두 gpt-5.6-sol subscription / effort high로 실행했습니다. mock/create_task_1은 0/1, Telecom-small 첫 task는 1/1입니다. 둘 다 정상 USER_STOP이며 provider, quota, adapter exception은 없었습니다.
- Mock: 이전과 동일하게
create_task가 요청에 없던 optionaldescription=""를 추가해 exact action/DB comparator가 실패했습니다. - Telecom: 이전의 premature human transfer가 사라졌습니다.
toggle_roaming, DB match, mobile-data 상태, excellent-speed assertion이 모두 1.0입니다. geode-eval-artifacts/trajectories/tau2-geode-gpt56-v1.0.11-686ff372-mock-telecom-small-20260731T105713Z-a71155f7006c: 142개 이벤트와 9개 exact tool pair를 담은 두geode.trajectory@1.
점수 정본은 여전히 tau2 results.json입니다. Crucible snapshot은 두 run을 diagnostic / promotion_authority=none으로 유지하며, GEODE trajectory는 digest-joined replay sidecar입니다.
Headline: native user-simulator 트랙
2026-07-03/04 run, GEODE v0.99.269, sierra-research/tau2-bench@1901a30 (tau2==1.0.0), agent gpt-5.2 PAYG effort high, native user_simulator gpt-4.1-2025-04-14 effort medium, max_steps=200.
| Mock | Retail | Telecom | Airline | Avg. |
|---|---|---|---|---|
| 1.000 reward 1.0, pass^1 1.000 | 0.763 base, 114 tasks, native user_simulator | 0.877 base, 114 tasks, native user_simulator | 0.820 base, 50 tasks, native user_simulator | 0.820 weighted across airline+retail+telecom, excludes mock |
현재 약점은 복합 태스크의 필수 액션 커버리지입니다. Retail 실패는 DB write 부수효과 누락, Telecom 실패는 MMS, APN, 앱 권한, 로밍 조합에서 필요한 액션 하나가 빠지는 패턴에 몰립니다.
Run 기록
모든 run은 측정 시각, model, provider, source, effort, route, harness revision, artifact 경로를 같은 규격으로 기록합니다.
Airline + Retail + Telecom base full-cycle GPT-5.4 diagnostic2026-08-03 KSTgpt-5.4subscriptionagent high / user high
| Status | complete |
| Suite/domain | airline + retail + telecom / base / 278 tasks |
| Model | gpt-5.4 |
| Provider | openai |
| Source | subscription |
| Effort | agent high / user high |
| Route | geode_agent + geode_user |
| Harness | sierra-research/tau2-bench@1901a30, tau2==1.0.0, GEODE@22789ee2 |
| Weighted reward / pass^1 | 0.7194 / 0.719 (200 / 278) |
| Artifact | geode-eval-artifacts@86dcbba3d15f1979b71a501780bf66fea4b450b5/reports/e2e-validation/2026-08-03-gpt54-tau2-full-cycle.json |
- Airline 0.8400 (42 / 50)
- Retail 0.6930 (79 / 114)
- Telecom 0.6930 (79 / 114)
- 51,985 canonical events / 3,964 exact tool pairs / zero orphans
- Telecom p95 957.65s / 14 max-step terminations / MMS 21 of 49
# Run once per domain with num-tasks 50 (airline) or 114 (retail/telecom).
python scripts/eval/tau2_geode_agent.py \
--harness-dir artifacts/eval/harnesses/tau2-bench \
--domain <airline|retail|telecom> \
--task-split-name base \
--num-tasks <50|114> \
--num-trials 1 \
--max-concurrency 2 \
--max-steps 200 \
--max-errors 1 \
--max-retries <0|1> \
--timeout 3600 \
--model gpt-5.4 \
--provider openai \
--source subscription \
--effort high \
--time-budget-s 600 \
--user geode_user \
--user-llm gpt-5.4 \
--user-provider openai \
--user-source subscription \
--user-effort high \
--user-time-budget-s 180 \
--trajectory-stage benchmark \
--save-to <domain-specific-run-id>- This GEODE-user full cycle is not comparable to the native tau2 user_simulator headline matrix.
- Tau2 results.json is score authority; the trajectory release is a privacy-reviewed diagnostic and external-loop sidecar.
- Seven Telecom transport retries created 14 extra SQLite sessions outside the final trajectory parents; no behavior-score failure was retried.
- The released trajectories are scope-complete for final task attempts and replay-incomplete for bounded bodies and retry-attempt lineage.
- The isolated Tau2 AgenticLoop records no public hook_events; the separate hook behavior E2E remains hook authority.
Telecom small first-task GPT-5.4 v1.0.12 post-release diagnostic2026-08-03 KSTgpt-5.4subscriptionagent high / user high
| Status | complete |
| Suite/domain | telecom / small / [mobile_data_issue]user_abroad_roaming_enabled_off[PERSONA:None] |
| Model | gpt-5.4 |
| Provider | openai |
| Source | subscription |
| Effort | agent high / user high |
| Route | geode_agent + geode_user |
| Harness | sierra-research/tau2-bench@1901a30, tau2==1.0.0, GEODE v1.0.12@f99cea63 |
| Reward / pass^1 | 0.0 / 0.000 (0 / 1) |
| Artifact | geode-eval-artifacts@04ff1c4a1fee0cd1a3d837ad3a5f5239f1fd9acd/tau2/simulations/geode-gpt54-high-v1.0.12-f99cea63-geode-user-telecom-small-01-20260803/results.json |
- Termination max_steps before native component scoring
- Duration 236.73s
- 203 canonical events / 14 exact tool pairs
- Repeated customer, line, network, usage, restriction, and VPN diagnostics
python scripts/eval/tau2_geode_agent.py --harness-dir artifacts/eval/harnesses/tau2-bench --domain telecom --task-split-name small --task-ids '[mobile_data_issue]user_abroad_roaming_enabled_off[PERSONA:None]' --num-tasks 1 --num-trials 1 --max-concurrency 1 --max-steps 50 --timeout 1800 --model gpt-5.4 --provider openai --source subscription --effort high --time-budget-s 300 --user geode_user --user-llm gpt-5.4 --user-provider openai --user-source subscription --user-effort high --user-time-budget-s 180 --trajectory-stage benchmark --save-to geode-gpt54-high-v1.0.12-f99cea63-geode-user-telecom-small-01-20260803- The run reached 50 steps before native DB/action scoring; repeated diagnostics are preserved as behavior evidence.
- All fourteen tool calls have exactly one result, with no route, authentication, quota, or adapter failure.
- This two-task release smoke does not invalidate or replace the 200/278 full-cycle diagnostic.
mock/create_task_1 GPT-5.4 v1.0.12 post-release diagnostic2026-08-03 KSTgpt-5.4subscriptionagent high / user high
| Status | complete |
| Suite/domain | mock / create_task_1 |
| Model | gpt-5.4 |
| Provider | openai |
| Source | subscription |
| Effort | agent high / user high |
| Route | geode_agent + geode_user |
| Harness | sierra-research/tau2-bench@1901a30, tau2==1.0.0, GEODE v1.0.12@f99cea63 |
| Reward / pass^1 | 0.0 / 0.000 (0 / 1) |
| Artifact | geode-eval-artifacts@04ff1c4a1fee0cd1a3d837ad3a5f5239f1fd9acd/tau2/simulations/geode-gpt54-high-v1.0.12-f99cea63-geode-user-mock-smoke-20260803/results.json |
- Communication check 1.0 / DB check 0.0
- create_task action check 0.0
- Termination user_stop
- Duration 13.75s / 31 canonical events / 2 exact tool pairs
python scripts/eval/tau2_geode_agent.py --harness-dir artifacts/eval/harnesses/tau2-bench --domain mock --task-ids create_task_1 --num-tasks 1 --num-trials 1 --max-concurrency 1 --max-steps 8 --timeout 900 --model gpt-5.4 --provider openai --source subscription --effort high --time-budget-s 180 --user geode_user --user-llm gpt-5.4 --user-provider openai --user-source subscription --user-effort high --user-time-budget-s 120 --trajectory-stage benchmark --save-to geode-gpt54-high-v1.0.12-f99cea63-geode-user-mock-smoke-20260803- The simulated user stopped before a verifier-compatible state change; DB and action checks are zero while communication is one.
- The run has no authentication, quota, provider-adapter, or harness exception and is retained without retry.
- This release smoke is not a rerun or replacement of the 278-task full cycle and is not a native user_simulator leaderboard row.
Telecom small first-task GPT-5.4 subscription diagnostic2026-08-02 KSTgpt-5.4subscriptionagent high / user high
| Status | complete |
| Suite/domain | telecom / small / [mobile_data_issue]user_abroad_roaming_enabled_off[PERSONA:None] |
| Model | gpt-5.4 |
| Provider | openai |
| Source | subscription |
| Effort | agent high / user high |
| Route | geode_agent + geode_user |
| Harness | sierra-research/tau2-bench@1901a30, tau2==1.0.0, GEODE@afaab52b |
| Reward / pass^1 | 1.0 / 1.000 (1 / 1) |
| Artifact | geode-eval-artifacts@f588ce9fd23b9123732b45c4dbe202136691d3fe/tau2/simulations/geode-gpt54-high-afaab52b-geode-user-telecom-small-01-20260802/results.json |
- DB check 1.0
- toggle_roaming write action 1.0
- Mobile-data and excellent-speed assertions 1.0
- Termination user_stop
- Duration 119.83s / 127 canonical events / 8 exact tool pairs
python scripts/eval/tau2_geode_agent.py \
--harness-dir artifacts/eval/harnesses/tau2-bench \
--domain telecom \
--task-split-name small \
--task-ids '[mobile_data_issue]user_abroad_roaming_enabled_off[PERSONA:None]' \
--num-tasks 1 \
--num-trials 1 \
--max-concurrency 1 \
--max-steps 50 \
--timeout 1800 \
--model gpt-5.4 \
--provider openai \
--source subscription \
--effort high \
--time-budget-s 300 \
--user geode_user \
--user-llm gpt-5.4 \
--user-provider openai \
--user-source subscription \
--user-effort high \
--user-time-budget-s 180 \
--save-to geode-gpt54-high-afaab52b-geode-user-telecom-small-01-20260802- The DB, toggle_roaming, mobile-data, and excellent-speed checks all passed.
- No route, provider, adapter, quota, agent, or simulated-user exception occurred.
- Tau2 results.json is the score authority; the 127-event trajectory is a digest-joined correlation and replay sidecar.
- The immutable source snapshot retains the runner-default train stage with promotion_authority=none; future benchmark commands should set --trajectory-stage benchmark explicitly.
mock/create_task_1 GPT-5.4 subscription diagnostic2026-08-02 KSTgpt-5.4subscriptionagent high / user high
| Status | complete |
| Suite/domain | mock / create_task_1 |
| Model | gpt-5.4 |
| Provider | openai |
| Source | subscription |
| Effort | agent high / user high |
| Route | geode_agent + geode_user |
| Harness | sierra-research/tau2-bench@1901a30, tau2==1.0.0, GEODE@afaab52b |
| Reward / pass^1 | 0.0 / 0.000 (0 / 1) |
| Artifact | geode-eval-artifacts@f588ce9fd23b9123732b45c4dbe202136691d3fe/tau2/simulations/geode-gpt54-high-afaab52b-geode-user-mock-smoke-20260802/results.json |
- Communication check 1.0 / DB check 0.0
- create_task action check 0.0
- Termination user_stop
- Duration 25.33s
- 31 canonical events / 2 exact tool pairs
python scripts/eval/tau2_geode_agent.py \
--harness-dir artifacts/eval/harnesses/tau2-bench \
--domain mock \
--task-ids create_task_1 \
--num-tasks 1 \
--num-trials 1 \
--max-concurrency 1 \
--max-steps 8 \
--timeout 900 \
--model gpt-5.4 \
--provider openai \
--source subscription \
--effort high \
--time-budget-s 180 \
--user geode_user \
--user-llm gpt-5.4 \
--user-provider openai \
--user-source subscription \
--user-effort high \
--user-time-budget-s 120 \
--save-to geode-gpt54-high-afaab52b-geode-user-mock-smoke-20260802- The model supplied unrequested description=""; Tau2's exact action and DB comparators rejected it.
- No route, provider, adapter, quota, agent, or simulated-user exception occurred.
- This fixed GEODE-user diagnostic is not a native user_simulator headline row.
- The immutable source snapshot retains the runner-default train stage with promotion_authority=none; future benchmark commands should set --trajectory-stage benchmark explicitly.
base aggregate native user_simulator2026-07-04 03:45 KSTgpt-5.2paygagent high / user medium
| Status | complete |
| Suite/domain | airline + retail + telecom / base |
| Model | gpt-5.2 |
| Provider | openai |
| Source | payg |
| Effort | agent high / user medium |
| Route | geode_agent + native tau2 user_simulator |
| Harness | sierra-research/tau2-bench@1901a30, tau2==1.0.0, GEODE v0.99.269 |
| Weighted reward / pass^1 | 0.8201 / 0.820 (228 / 278) |
| Artifact | artifacts/eval/harnesses/tau2-bench/data/simulations/geode-gpt-5-2-high-native-user-{airline,retail,telecom}-base-20260703/results.json |
- Airline 0.8200 (41 / 50)
- Retail 0.7632 (87 / 114)
- Telecom 0.8772 (100 / 114)
- Native user simulator gpt-4.1-2025-04-14
- GEODE recorded gpt-5.2 PAYG usage locally; user simulator cost is visible through OpenAI billing, not GEODE's usage ledger.
# Aggregate of the three per-domain native tau2 runs listed above.
# Do not average this with mock smoke or GEODE geode_user rows.- This weighted aggregate is for internal Agent-World-style comparison only.
- The run spec differs from OpenAI's official GPT-5.2 Tau2 headline, which used an internal research setup and excludes Airline.
- The run spec differs from the earlier GEODE geode_user smoke matrix.
Telecom small first-task GPT-5.6 v1.0.11 diagnostic2026-07-31 KSTgpt-5.6-solsubscriptionagent high / user high
| Status | complete |
| Suite/domain | telecom / small / [mobile_data_issue]user_abroad_roaming_enabled_off[PERSONA:None] |
| Model | gpt-5.6-sol |
| Provider | openai |
| Source | subscription |
| Effort | agent high / user high |
| Route | geode_agent + geode_user |
| Harness | sierra-research/tau2-bench@1901a30, tau2==1.0.0, GEODE v1.0.11@686ff372 |
| Reward / pass^1 | 1.0 / 1.000 (1 / 1) |
| Artifact | geode-eval-artifacts@16a54f08450db771c02e30c73bdc3867f6282f83/tau2/simulations/geode-gpt56-sol-high-v1011-686ff372-geode-user-telecom-small-01-20260731/results.json |
- DB check 1.0
- toggle_roaming write action 1.0
- Mobile-data and excellent-speed assertions 1.0
- Termination user_stop
- Duration 78.52s / 117 canonical events / 8 exact tool pairs
python scripts/eval/tau2_geode_agent.py \
--harness-dir artifacts/eval/harnesses/tau2-bench \
--domain telecom \
--task-split-name small \
--task-ids '[mobile_data_issue]user_abroad_roaming_enabled_off[PERSONA:None]' \
--num-tasks 1 \
--num-trials 1 \
--max-concurrency 1 \
--max-steps 50 \
--timeout 1800 \
--model gpt-5.6-sol \
--provider openai \
--source subscription \
--effort high \
--time-budget-s 300 \
--user geode_user \
--user-llm gpt-5.6-sol \
--user-source subscription \
--user-effort high \
--user-time-budget-s 180 \
--save-to geode-gpt56-sol-high-v1011-686ff372-geode-user-telecom-small-01-20260731- The earlier premature human-transfer failure is closed for this fixed case.
- Tau2 native DB/action/assertion checks remain the score authority; the GEODE trajectory is a digest-joined replay sidecar.
- The Crucible v3 snapshot remains diagnostic with promotion_authority=none because no frozen experiment contract was supplied.
mock/create_task_1 GPT-5.6 v1.0.11 diagnostic2026-07-31 KSTgpt-5.6-solsubscriptionagent high / user high
| Status | complete |
| Suite/domain | mock / create_task_1 |
| Model | gpt-5.6-sol |
| Provider | openai |
| Source | subscription |
| Effort | agent high / user high |
| Route | geode_agent + geode_user |
| Harness | sierra-research/tau2-bench@1901a30, tau2==1.0.0, GEODE v1.0.11@686ff372 |
| Reward / pass^1 | 0.0 / 0.000 (0 / 1) |
| Artifact | geode-eval-artifacts@16a54f08450db771c02e30c73bdc3867f6282f83/tau2/simulations/geode-gpt56-sol-high-v1011-686ff372-geode-user-mock-smoke-20260731/results.json |
- Communication check 1.0 / DB check 0.0
- create_task action check 0.0
- Termination user_stop
- Duration 9.03s
- 25 canonical events / 1 exactly paired tool call and result
python scripts/eval/tau2_geode_agent.py \
--harness-dir artifacts/eval/harnesses/tau2-bench \
--domain mock \
--num-tasks 1 \
--num-trials 1 \
--max-concurrency 1 \
--max-steps 8 \
--timeout 900 \
--model gpt-5.6-sol \
--provider openai \
--source subscription \
--effort high \
--time-budget-s 180 \
--user geode_user \
--user-llm gpt-5.6-sol \
--user-source subscription \
--user-effort high \
--user-time-budget-s 120 \
--save-to geode-gpt56-sol-high-v1011-686ff372-geode-user-mock-smoke-20260731- The failure reproduces the earlier behavior: create_task includes unrequested description="" and the native exact comparator rejects it.
- The run completed normally and is retained without retry or relabeling.
- This diagnostic has promotion_authority=none and is not a native user_simulator leaderboard row.
Telecom small first-task GPT-5.6 subscription diagnostic2026-07-31 KSTgpt-5.6-solsubscriptionagent high / user high
| Status | complete |
| Suite/domain | telecom / small / [mobile_data_issue]user_abroad_roaming_enabled_off[PERSONA:None] |
| Model | gpt-5.6-sol |
| Provider | openai |
| Source | subscription |
| Effort | agent high / user high |
| Route | geode_agent + geode_user |
| Harness | sierra-research/tau2-bench@1901a30, tau2==1.0.0, GEODE@edb74602b |
| Reward / pass^1 | 0.0 / 0.000 (0 / 1) |
| Artifact | geode-eval-artifacts@9c00ecf/tau2/simulations/geode-gpt56-sol-high-edb74602b-geode-user-telecom-small-01-20260731/results.json |
- Required user toggle_roaming action 0.0
- Mobile-data and excellent-speed assertions 0.0
- Termination user_stop after human transfer
- Duration 51.91s
python scripts/eval/tau2_geode_agent.py \
--harness-dir artifacts/eval/harnesses/tau2-bench \
--domain telecom \
--task-split-name small \
--task-ids '[mobile_data_issue]user_abroad_roaming_enabled_off[PERSONA:None]' \
--num-tasks 1 \
--num-trials 1 \
--max-concurrency 1 \
--max-steps 50 \
--timeout 1800 \
--model gpt-5.6-sol \
--provider openai \
--source subscription \
--effort high \
--user geode_user \
--user-llm gpt-5.6-sol \
--user-source subscription \
--user-effort high \
--save-to geode-gpt56-sol-high-edb74602b-geode-user-telecom-small-01-20260731- The agent correctly identified the customer, line, roaming state, and data usage.
- It then declared device tools unavailable and transferred to a human instead of guiding the user-side roaming/device workflow.
- No provider, quota, or adapter exception occurred; the failure is retained as behavior evidence.
mock/create_task_1 GPT-5.6 subscription diagnostic2026-07-31 KSTgpt-5.6-solsubscriptionagent high / user high
| Status | complete |
| Suite/domain | mock / create_task_1 |
| Model | gpt-5.6-sol |
| Provider | openai |
| Source | subscription |
| Effort | agent high / user high |
| Route | geode_agent + geode_user |
| Harness | sierra-research/tau2-bench@1901a30, tau2==1.0.0, GEODE@edb74602b |
| Reward / pass^1 | 0.0 / 0.000 (0 / 1) |
| Artifact | geode-eval-artifacts@9c00ecf/tau2/simulations/geode-gpt56-sol-high-edb74602b-geode-user-mock-smoke-20260731/results.json |
- Communication check 1.0 / DB check 0.0
- create_task action check 0.0
- Termination user_stop
- Duration 14.58s
python scripts/eval/tau2_geode_agent.py \
--harness-dir artifacts/eval/harnesses/tau2-bench \
--domain mock \
--num-tasks 1 \
--num-trials 1 \
--max-concurrency 1 \
--max-steps 8 \
--timeout 900 \
--model gpt-5.6-sol \
--provider openai \
--source subscription \
--effort high \
--user geode_user \
--user-llm gpt-5.6-sol \
--user-source subscription \
--user-effort high \
--save-to geode-gpt56-sol-high-edb74602b-geode-user-mock-smoke-20260731- The create_task tool executed, but the model supplied an unrequested optional description="".
- Tau2's exact action and DB comparators rejected the extra argument; this is retained as a behavioral failure.
- This GEODE-owned user route is not comparable to the native tau2 user_simulator headline.
telecom/base native user_simulator2026-07-04 03:45 KSTgpt-5.2paygagent high / user medium
| Status | complete |
| Suite/domain | telecom / base |
| Model | gpt-5.2 |
| Provider | openai |
| Source | payg |
| Effort | agent high / user medium |
| Route | geode_agent + native tau2 user_simulator |
| Harness | sierra-research/tau2-bench@1901a30, tau2==1.0.0, GEODE v0.99.269 |
| Reward / pass^1 | 0.8772 / 0.877 (100 / 114) |
| Artifact | artifacts/eval/harnesses/tau2-bench/data/simulations/geode-gpt-5-2-high-native-user-telecom-base-20260703/results.json |
- DB match 31 / 114
- Write actions 471 / 496
- Generic actions 20 / 20
- Termination user_stop 114 / 114
- Duration total 28827.72s / avg 252.87s / max 818.58s
uv run python scripts/eval/tau2_geode_agent.py \
--harness-dir artifacts/eval/harnesses/tau2-bench \
--domain telecom \
--task-split-name base \
--num-tasks 114 \
--num-trials 1 \
--max-concurrency 4 \
--max-steps 200 \
--timeout 3600 \
--model gpt-5.2 \
--provider openai \
--source payg \
--effort high \
--time-budget-s 600 \
--user user_simulator \
--user-llm gpt-4.1-2025-04-14 \
--user-provider openai \
--user-source payg \
--user-effort medium \
--user-time-budget-s 120 \
--save-to geode-gpt-5-2-high-native-user-telecom-base-20260703 \
--log-level INFO \
--auto-resume- GEODE version at measurement: v0.99.269.
- Concurrency was raised from 2 to 4 mid-run and resumed from tau2 checkpoints; no rate-limit, quota, or billing errors were observed.
- Failures clustered around multi-issue MMS/mobile-data/service cases where one required APN, permission, roaming, or data-refuel action was omitted.
retail/base native user_simulator2026-07-03 KSTgpt-5.2paygagent high / user medium
| Status | complete |
| Suite/domain | retail / base |
| Model | gpt-5.2 |
| Provider | openai |
| Source | payg |
| Effort | agent high / user medium |
| Route | geode_agent + native tau2 user_simulator |
| Harness | sierra-research/tau2-bench@1901a30, tau2==1.0.0, GEODE v0.99.269 |
| Reward / pass^1 | 0.7632 / 0.763 (87 / 114) |
| Artifact | artifacts/eval/harnesses/tau2-bench/data/simulations/geode-gpt-5-2-high-native-user-retail-base-20260703/results.json |
- DB match 88 / 113
- Read actions 320 / 354
- Write actions 140 / 174
- Termination user_stop 113 / 114, too_many_errors 1 / 114
- Duration total 23543.64s / avg 206.52s / max 873.92s
uv run python scripts/eval/tau2_geode_agent.py \
--harness-dir artifacts/eval/harnesses/tau2-bench \
--domain retail \
--task-split-name base \
--num-tasks 114 \
--num-trials 1 \
--max-concurrency 2 \
--max-steps 200 \
--timeout 3600 \
--model gpt-5.2 \
--provider openai \
--source payg \
--effort high \
--time-budget-s 600 \
--user user_simulator \
--user-llm gpt-4.1-2025-04-14 \
--user-provider openai \
--user-source payg \
--user-effort medium \
--user-time-budget-s 120 \
--save-to geode-gpt-5-2-high-native-user-retail-base-20260703 \
--log-level INFO \
--auto-resume- GEODE version at measurement: v0.99.269.
- The main failure mode was missing required side-effect actions even when the natural-language response looked plausible.
- One task terminated with too_many_errors; the remaining failures ended with user_stop but failed verifier assertions.
airline/base native user_simulator2026-07-03 KSTgpt-5.2paygagent high / user medium
| Status | complete |
| Suite/domain | airline / base |
| Model | gpt-5.2 |
| Provider | openai |
| Source | payg |
| Effort | agent high / user medium |
| Route | geode_agent + native tau2 user_simulator |
| Harness | sierra-research/tau2-bench@1901a30, tau2==1.0.0, GEODE v0.99.269 |
| Reward / pass^1 | 0.8200 / 0.820 (41 / 50) |
| Artifact | artifacts/eval/harnesses/tau2-bench/data/simulations/geode-gpt-5-2-high-native-user-airline-base-20260703/results.json |
- DB match 42 / 50
- Read actions 81 / 91
- Write actions 33 / 49
- Termination user_stop 50 / 50
- Duration total 14205.02s / avg 284.10s / max 979.65s
uv run python scripts/eval/tau2_geode_agent.py \
--harness-dir artifacts/eval/harnesses/tau2-bench \
--domain airline \
--task-split-name base \
--num-tasks 50 \
--num-trials 1 \
--max-concurrency 2 \
--max-steps 200 \
--timeout 3600 \
--model gpt-5.2 \
--provider openai \
--source payg \
--effort high \
--time-budget-s 600 \
--user user_simulator \
--user-llm gpt-4.1-2025-04-14 \
--user-provider openai \
--user-source payg \
--user-effort medium \
--user-time-budget-s 120 \
--save-to geode-gpt-5-2-high-native-user-airline-base-20260703 \
--log-level INFO \
--auto-resume- GEODE version at measurement: v0.99.269.
- This is the native tau2 user_simulator comparator track, not the GEODE geode_user smoke track.
- Airline is retained for internal trend comparison; OpenAI's GPT-5.2 announcement excludes Airline from its Tau2 headline due to lower-quality ground truth grading.
mock/create_task_1 smoke2026-07-03 KSTgpt-5.5subscriptionagent xhigh / user high
| Status | complete |
| Suite/domain | mock / create_task_1 |
| Model | gpt-5.5 |
| Provider | openai |
| Source | subscription |
| Effort | agent xhigh / user high |
| Route | geode_agent + geode_user |
| Harness | sierra-research/tau2-bench@1901a30, tau2==1.0.0 |
| Reward / pass^1 | 1.0 / 1.000 (1 / 1) |
| Artifact | artifacts/eval/harnesses/tau2-bench/data/simulations/geode-gpt-5-5-xhigh-geode-user-mock-smoke-20260703-r5/results.json |
- DB check 1.0
- create_task action check 1.0
- Termination user_stop
- Duration 54.90s
uv run python scripts/eval/tau2_geode_agent.py \
--harness-dir artifacts/eval/harnesses/tau2-bench \
--domain mock \
--num-tasks 1 \
--num-trials 1 \
--max-concurrency 1 \
--max-steps 8 \
--timeout 900 \
--model gpt-5.5 \
--provider openai \
--source subscription \
--effort xhigh \
--time-budget-s 180 \
--user geode_user \
--user-llm gpt-5.5 \
--user-provider openai \
--user-source subscription \
--user-effort high \
--user-time-budget-s 120 \
--save-to geode-gpt-5-5-xhigh-geode-user-mock-smoke-20260703-r5 \
--log-level INFO \
--verbose-logs- This is a tau2 wiring/regression smoke, not a tau2 leaderboard score.
- Do not average it with native tau2 user_simulator runs using gpt-4.1 or gpt-5.2.
- Both assistant and simulated user used the GEODE subscription route.
Run 로그
원본 simulation JSON(태스크별 reward, 액션 체크, 전체 대화 transcript)은 geode-eval-artifacts 레포에 로컬 경로와 합성 개인정보를 마스킹한 공개 copy로 보존됩니다.
geode-eval-artifacts/tau2/simulations: GEODE 소유 run의 simulation JSON. headline run은geode-gpt-5-2-high-native-user-*-base-20260703/results.json패턴입니다.
run 기록의 artifact 경로는 측정 당시 로컬 harness 경로입니다. 게시된 사본은 위 레포 경로에서 파일명으로 찾습니다.