상태 기반 MCP 검증
점수·비용·증거 누락을 함께 봅니다
Available services
64 / 74
86.5% · filesystem + postgres + GitHub
historical slice
Gate 0C
23 / 30
Codex 21 / 30 · common deadline
diagnostic · k=1
Coverage gap
53 tasks
Notion 28 + Playwright 4 + WebArena 21
unmeasured
Gate 0C
59 admitted · 1 withheld
동일 filesystem/standard 30건과 공통 deadline을 사용했습니다. timeout trajectory 한 건은 점수에 남고 공개 admission에서는 보류됐습니다.
Gate 0C 원본 bundle ↗Gate 0B
7 / 15 guard · 10 / 15 unlimited
25K result guard의 직접 ablation입니다. 세 반복 중 네 timeout은 withheld로 남겼고, 단일 treatment 진단에는 승격 권한을 주지 않았습니다.
Gate 0B 원본 bundle ↗측정 기록의 발전
MCPMark Verified는 service별 pass@1과 turns·time·token·cost를 함께 열고, trajectory가 없는 제출을 명시합니다. GEODE 표면도 미측정 service와 withheld trajectory를 점수 옆에 보존합니다.
MCPMark는 실제 MCP 서버(filesystem, Postgres, GitHub, Notion, Playwright 등)를 대상으로 한 tool-use 벤치마크입니다. 태스크마다 독립 검증 스크립트가 결과 상태를 확인합니다. GEODE는 evals/benchmarks의 BaseMCPAgent 어댑터로 참가하고 upstream pipeline.py는 패치하지 않습니다. 점수는 harness commit, 서비스 집합, model route, timeout에 고정해서만 게시합니다.
2026-08-13 GPT-5.4 filesystem/standard 정정 관측
고정된 30개 filesystem/standard task를 GPT-5.4 subscription / effort high로 task별 paired 실행했습니다. GEODE는 21/30 (70.0%), Codex CLI는 20/30 (66.7%)로 GEODE가 1건 앞섰습니다. 60개 trajectory에 3,381 events가 보존됐고, 1,430 tool call/result가 정확히 pairing됐으며 orphan은 없습니다.
사후 source audit에서 원래 사전등록한 equal-hard-deadline 전제가 성립하지 않았음이 확인됐습니다. GEODE는 MCP setup 뒤의loop.arun만, Codex는 내부 MCP startup을 포함하는 process communication을 timed surface로 사용했습니다. 따라서 prospective hypothesis는 invalidated이며, 점수는 retrospective description으로만 남습니다. Native input 총합은 GEODE가 작았지만 cache 제외 입력은 4.20M 대 1.44M으로 더 컸으므로 token-efficiency도 주장하지 않습니다. 공개 bundle에는 정확한 runner가 없으므로 독립 실행 가능한 재현 패키지도 아닙니다.
geode-eval-artifacts/mcpmark/results-paired/mcpmark-filesystem-standard-gpt54-high-geode-codex-k1-boundary-aligned-20260813: 원본 spec·receipt와 이를 supersede하는 정정 analysis·receipt.geode-eval-artifacts/trajectories/mcpmark-geode-gpt54-high-mcpmark-filesystem-standard-gpt54-high-geode-codex--818b13fe1039-20260812T231820Z-ed26f124b9c7: privacy-reviewed GEODE 30-task trajectory release.geode-eval-artifacts/trajectories/mcpmark-codex-gpt54-high-mcpmark-filesystem-standard-gpt54-high-geode-codex--f749317fe281-20260812T231820Z-828560273a4e: privacy-reviewed Codex 30-task trajectory release.
2026-08-12 matched token-efficiency rerun
같은 GPT-5.4 subscription / effort high, 같은 pinned filesystem/easy 10건을 수정 전후로 대조했습니다. 점수는 9/10 (90.0%)로 유지됐고, 입력 토큰은 447,376에서 314,219로 29.8%, 출력 토큰은 25,157에서 20,385로 19.0% 줄었습니다.
10건 중 8건의 입력 토큰이 감소했고 round 수가 같은 4건도 12.5% 감소했습니다. 188개 canonical event와 54/54 exact tool pair에는 orphan이 없습니다. 단, 한 번의 matched trial이므로 MCPMark Verified 점수·신뢰구간·구독 과금 절감으로 일반화하지 않습니다.
geode-eval-artifacts/trajectories/mcpmark-geode-gpt54-high-token-efficiency-rerun-filesystem-easy-20260812T090254Z-35db8b275a36: 원격 read-back과 privacy 검증을 통과한 10개 stable trajectory.- matched rerun report: task별 변화와 promotion 경계를 포함한 판정 근거.
2026-08-03 v1.0.12 post-release regression
공개 배포된 GEODE v1.0.12 (f99cea63)과 gpt-5.4 subscription / effort high로 filesystem/easy 10건을 실행했습니다. 공식 verifier는 9/10 (90.0%), 총 802.2초와 53 turns입니다. 실패한 file_context/uppercase는 다섯 파일을 모두 만들었지만 file_01.txt를 완전히 대문자로 바꾸지 못했습니다.
인증·quota·provider adapter·MCP transport 오류는 없었습니다. 10개 trajectory는 182개 canonical event와 56개 exact tool pair를 보존하며 scope_complete=true, replay_complete=false입니다. v1.0.11의 GPT-5.6 10/10과 비교할 때 모델까지 바뀌었으므로 release 회귀로 단정하지 않습니다.
geode-eval-artifacts/mcpmark/results-geode-agentworld/geode-gpt54-high-v1.0.12-f99cea63-20260803-mcpmark-filesystem-easy: verifier receipt, redacted execution logs, raw/public digest ledger.geode-eval-artifacts/trajectories/mcpmark-geode-gpt54-v1.0.12-f99cea63-filesystem-easy-20260803T104819Z-9636b39c16fb: manifest SHA-2569636b39c16fb…d267로 원격 read-back된 stable release.
2026-07-31 v1.0.11 release regression
배포된 GEODE v1.0.11 (686ff372)과 gpt-5.6-sol subscription / effort high로 filesystem/easy 10건을 재측정했습니다. 공식 verifier는 10/10 (100.0%), 총 596.6초와 56 turns입니다. 이전 edb74602b run의 유일한 실패였던 file_context/uppercase도 통과했습니다.
10개 stable trajectory의 226개 이벤트는 canonical SQLite 행과 ID·session·turn·call·kind까지 일치합니다. 78개 tool call/result가 모두 정확히 pairing됐고 필수 turn ID 누락은 0건입니다.
geode-eval-artifacts/mcpmark/results-geode-agentworld/geode-gpt56-sol-high-v1011-686ff372-20260731-mcpmark-filesystem-easy: 마스킹된 verifier receipt와 ordered MCP execution logs.geode-eval-artifacts/trajectories/mcpmark-geode-gpt56-v1.0.11-686ff372-filesystem-easy-20260731T105713Z-82fe94b01a25: privacy review와 source digest 검증을 통과한 10개geode.trajectory@1.
Headline: Verified available-services 트랙
2026-07-04 run, GEODE v0.99.269 계열, eval-sys/mcpmark@cd45b7f, gpt-5.5 Codex 구독 route, effort xhigh. 이 측정의 범위는 로컬에서 실행 가능했던 standard 슬라이스(filesystem, postgres, github)입니다.
| File | GitHub | Notion | Playwright | Postgres | Avg. |
|---|---|---|---|---|---|
| 83.3% standard, 25 / 30 | 82.6% standard, 19 / 23 | unmeasured unblocked 2026-07-10 (easy smoke 1/1); standard 28 tasks not yet measured | unmeasured live-web subset runnable since 2026-07-10; WebArena subset needs ~100GB images (local disk exceeded) | 95.2% standard, 20 / 21 | 86.5% Measured available services only: filesystem+postgres+github |
Service coverage
| Service | Easy | Standard | Adapter 상태 | Blocker |
|---|---|---|---|---|
filesystem | 10 | 30 | standard run 완료 | historical 25 / 30; paired GPT-5.4 21 / 30 |
postgres | 10 | 21 | standard run 완료 | 20 / 21, postgres-mcp==0.3.0 |
github | 10 | 23 | standard run 완료 | 19 / 23, Docker GitHub MCP server. State Duplication Error 6건의 원인(GITHUB_EVAL_ORG 미영속)은 2026-07-10 제거 |
notion | 10 | 28 | 실측 가능 (easy smoke 1/1, 2026-07-10) | 07-04 스톨 원인은 브라우저 세션 만료로 확정, 재발급 절차 확립. standard 28건 미측정 |
playwright | 0 | 4 | 실행 준비 완료 (2026-07-10) | @playwright/mcp@0.0.68 기동 확인. 4건 미측정 |
playwright_webarena | 10 | 21 | stdio adapter 준비 | WebArena Docker 이미지 실측 119GiB vs 로컬 여유 13GiB. 외장 볼륨 또는 VM 필요 |
insforge | 확인 필요 | 조사 중 | INSFORGE_API_KEY, task manager 인자 호환성 확인 필요 | |
supabase | 확인 필요 | 미지원 | HTTP MCP transport. GEODE MCPServerManager는 현재 stdio 중심 | |
구독 쿼터(429 usage_limit_reached)는 full-suite 연속 실행을 리셋 창 단위로 분할시킵니다. 429 실패는 점수에 포함하지 않고 해당 태스크를 재실행합니다.
Run 기록
Verified available-services aggregate2026-07-04 KSTgpt-5.5subscriptionxhigh
| Status | complete |
| Suite/domain | filesystem + postgres + github / standard |
| Model | gpt-5.5 |
| Provider | openai-codex |
| Source | subscription |
| Effort | xhigh |
| Route | GEODE local MCPMark adapter |
| Harness | eval-sys/mcpmark@cd45b7f, GEODE feature/mcpmark-agentworld-run |
| Accuracy | 86.5% (64 / 74) |
| Artifact | artifacts/eval/harnesses/mcpmark/results-geode-agentworld/geode-gpt55-xhigh-20260704-mcpmark-verified-* |
- Filesystem standard: 25 / 30, 83.3%
- Postgres standard: 20 / 21, 95.2%
- GitHub standard: 19 / 23, 82.6%
- Recorded task execution time: filesystem 13580.6s over 29 recorded tasks, postgres 8765.7s, github 16476.3s
- Notion was not included: no notion_state.json in the local harness environment.
- Playwright/WebArena was not included: required Docker images/service stack were absent.
cd artifacts/eval/harnesses/mcpmark
# Run each available MCP service through the GEODE adapter.
GEODE_REPO_ROOT=<geode-worktree> \
PYTHONPATH=<geode-worktree>:<geode-site-packages> \
GITHUB_EVAL_ORG=mangowhoiscloud \
GEODE_MCPMARK_GITHUB_REPO_VISIBILITY=public \
.venv/bin/python pipeline.py \
--mcp <filesystem|postgres|github> \
--task-suite standard \
--models geode-gpt-5.5 \
--agent geode \
--reasoning-effort xhigh \
--k 1 \
--timeout 1500 \
--exp-name geode-gpt55-xhigh-20260704-mcpmark-verified-<service> \
--output-dir ./results-geode-agentworld- This is not the full MCPMark Verified leaderboard aggregate. It covers only services that were runnable in the local environment: filesystem, postgres, and github.
- The OpenAI model route was the GEODE Codex subscription route, not MCPMark's native LiteLLM OpenAI API route.
- GitHub fixture repositories were made public during execution so the Docker GitHub MCP server could use normal public-repo semantics; all transient repos were deleted by cleanup.
- The filesystem score counts papers/author_folders as a failed no-result transport run after two attempts without meta output.
filesystem/standard Gate 0C common-deadline diagnostic2026-08-14 KSTgpt-5.4subscriptionhigh
| Status | complete |
| Suite/domain | filesystem/standard 30-task paired k=1 |
| Model | gpt-5.4 |
| Provider | openai / codex-cli |
| Source | subscription |
| Effort | high |
| Route | GEODE AgenticLoop and isolated Codex CLI paired by task |
| Harness | eval-sys/mcpmark@cd45b7f, GEODE@f4b37604, Codex CLI 0.145.0@dad1db87 |
| Diagnostic-only verifier pass rate | GEODE 76.7% (23 / 30) · Codex 70.0% (21 / 30) · Δ +6.67 pp |
| Artifact | geode-eval-artifacts@1160fecfe4447f0a3f4cf30a414f29c61776d012/mcpmark/results-paired/mcpmark-gate0c-filesystem30-gpt54-high-20260813t190922z |
- Paired outcomes: 17 both-pass / 3 both-fail / 6 GEODE-only / 4 Codex-only
- Exact token coverage: GEODE 29 / 30 attempts; Codex 30 / 30
- All-arm action wall: GEODE 8,463.252s vs Codex 5,484.322s; runner envelopes 8,536.382s vs 5,548.725s
- Native execution-log calls/errors: GEODE 644 / 51 vs Codex 678 / 17
- Normalized trajectory attempts: GEODE 645 including one recovery projection vs Codex 678
- Read/repeated-read references: GEODE 798 / 81 vs Codex 838 / 213
- One GEODE score-bearing deadline expiration; its token usage is null and its scope-incomplete trajectory is withheld
# Requires the pinned MCPMark checkout and verifier patch, frozen run spec,
# isolated runtime home, and working GPT-5.4 subscription authentication.
.venv/bin/python -m plugins.benchmark_harness.run_mcpmark_pair --profile filesystem30-geode-codex --run-spec <frozen-run-spec.json> --mcpmark-root artifacts/eval/harnesses/mcpmark --output-dir <fresh-output-dir> --python .venv/bin/python- The prospective -10 percentage-point threshold was supported, but promotion_authority remains none.
- This is one direct paired repetition, not k=3 stability, a full MCPMark Verified headline, or an API-key leaderboard claim.
- The common action deadline excludes fixture setup and the post-action verifier; runtime scaffolds and tool-result policies remain different.
- Token totals have unmatched coverage and do not support a billing or token-efficiency claim.
- Fifty-nine scope-complete trajectories are admitted; the one GEODE timeout trajectory remains withheld.
filesystem/standard Gate 0B tool-result-cap diagnostic2026-08-13–14 KSTgpt-5.4subscriptionhigh
| Status | complete |
| Suite/domain | filesystem/standard targeted 5 tasks × 3 repetitions |
| Model | gpt-5.4 |
| Provider | openai |
| Source | subscription |
| Effort | high |
| Route | GEODE AgenticLoop paired 25K and unlimited tool-result-cap arms |
| Harness | eval-sys/mcpmark@cd45b7f, GEODE@02f71fae |
| Diagnostic-only verifier pass rate | 25K 46.7% (7 / 15) · unlimited 66.7% (10 / 15) · Δ +20.0 pp |
| Artifact | geode-eval-artifacts@17133f0c8e893b6d765fcef69712ba0867bd573a/mcpmark/results-paired/mcpmark-gate0b-tool-cap-gpt54-high-20260813t142345z |
- Observed fresh input: 25K 3,782,288 vs unlimited 2,202,725; exact token coverage is 13 / 15 attempts per arm
- All-attempt action wall: 25K 6,076.699s vs unlimited 4,910.217s; runner envelopes 6,119.096s vs 4,951.086s
- All-attempt MCP calls/errors: 25K 443 / 75 vs unlimited 255 / 27
- Repeated read references: 25K 211 vs unlimited 63; tool-result truncations 15 vs 0
- Two score-bearing deadline expirations per arm emitted no native token usage
# Requires the pinned MCPMark checkout and verifier patch, frozen run spec,
# isolated runtime home, and working GPT-5.4 subscription authentication.
.venv/bin/python -m plugins.benchmark_harness.run_mcpmark_pair --profile max-tool-result-tokens --run-spec <frozen-run-spec.json> --mcpmark-root artifacts/eval/harnesses/mcpmark --output-dir <fresh-output-dir> --python .venv/bin/python- The prospective hypothesis was supported, but promotion_authority remains none.
- This five-task, three-repetition subset is diagnostic and is not an MCPMark suite headline.
- Token totals exclude the four timeout arms with empty native usage; wall time and MCP-call totals cover all 15 attempts per arm.
- The prior infrastructure-invalid run contributes no denominator, and scope-incomplete timeout trajectories remain withheld.
filesystem/standard GPT-5.4 GEODE × Codex corrected observation2026-08-13 KSTgpt-5.4subscriptionhigh
| Status | complete |
| Suite/domain | filesystem/standard |
| Model | gpt-5.4 |
| Provider | openai / codex-cli |
| Source | subscription |
| Effort | high |
| Route | GEODE AgenticLoop and isolated Codex CLI paired by task |
| Harness | eval-sys/mcpmark@cd45b7f, GEODE@a8f45f3c9, Codex@dad1db87 |
| Retrospective descriptive verifier outcomes | GEODE 70.0% (21 / 30) · Codex 66.7% (20 / 30) |
| Artifact | geode-eval-artifacts@e5d442f25c9fb4861e28744dbe924a36325c746b/mcpmark/results-paired/mcpmark-filesystem-standard-gpt54-high-geode-codex-k1-boundary-aligned-20260813 |
- Observed delta +1 / 30 (+3.33 pp); prospective hypothesis invalidated
- GEODE 1,646 events / 703 exact tool pairs; Codex 1,735 / 727
- GEODE fresh input 4,204,759 vs Codex 1,444,927; no token-efficiency claim
- Task wall: GEODE 7,842.4s vs Codex 6,970.2s (+12.5%); no causal efficiency claim
- Both releases scope-complete, replay-incomplete, zero orphan tool events
# Illustrative one-task invocation only; exact runner is withheld.
python -m plugins.benchmark_harness.run_mcpmark --mcp filesystem --task-suite standard --tasks <one-frozen-task-id> --models <geode-gpt-5.4|codex-gpt-5.4> --agent <geode|codex> --reasoning-effort high --k 1 --timeout 1200- All 30 pairs share the pinned task tree, fixture reset, model label, effort, and verifier.
- The frozen equal-hard-deadline claim was invalidated: GEODE timed loop.arun, while Codex timed process communication including internal MCP startup.
- Prompt, action budget, retry, compaction, and cache accounting are also not identical; outcomes are retrospective descriptions only.
- The public bundle is not independently executable because the exact runner remains digest-bound but withheld.
- This is one filesystem service slice, not the full 127-task MCPMark Verified leaderboard.
- Raw messages, logs, metadata, and provider diagnostics remain private; only digest-reduced trajectories and validated sidecars are public.
filesystem/easy GPT-5.4 token-efficiency rerun2026-08-12 KSTgpt-5.4subscriptionhigh
| Status | complete |
| Suite/domain | filesystem/easy |
| Model | gpt-5.4 |
| Provider | openai |
| Source | subscription |
| Effort | high |
| Route | GEODE AgenticLoop MCPMark adapter |
| Harness | eval-sys/mcpmark@cd45b7f, GEODE feature@149024e6e |
| Accuracy | 90.0% (9 / 10) |
| Artifact | geode-eval-artifacts@2c2d1f0621f64ff7ceeff8c05d8ebd3449501aaf/trajectories/mcpmark-geode-gpt54-high-token-efficiency-rerun-filesystem-easy-20260812T090254Z-35db8b275a36 |
- Matched input tokens 447,376 → 314,219 (-29.8%)
- Matched output tokens 25,157 → 20,385 (-19.0%)
- Native reasoning tokens 14,174, included within output tokens
- 188 canonical events / 54 exactly paired tool calls and results
- Failure unchanged: file_context/uppercase exact-string mismatch
cd artifacts/eval/harnesses/mcpmark
PYTHONPATH=<geode-feature-tree> .venv/bin/python -m plugins.benchmark_harness.run_mcpmark --mcp filesystem --task-suite easy --models geode-gpt-5.4 --agent geode --reasoning-effort high --k 1 --timeout 1200 --exp-name geode-gpt54-high-token-efficiency-20260812-rerun --output-dir ./results-token-efficiency- The score matched the pre-repair GEODE baseline at 9/10 while input and output tokens fell materially.
- Eight of ten tasks used fewer input tokens; the four tasks with identical round counts fell 12.5%.
- This is one matched diagnostic trial, not MCPMark Verified, a confidence interval, or a subscription billing claim.
- All ten public trajectories are scope-complete and intentionally replay-incomplete; the immutable release was read back from artifact main.
filesystem/easy GPT-5.4 v1.0.12 post-release regression2026-08-03 KSTgpt-5.4subscriptionhigh
| Status | complete |
| Suite/domain | filesystem/easy |
| Model | gpt-5.4 |
| Provider | openai |
| Source | subscription |
| Effort | high |
| Route | GEODE AgenticLoop MCPMark adapter |
| Harness | eval-sys/mcpmark@cd45b7f, GEODE v1.0.12@f99cea63 |
| Accuracy | 90.0% (9 / 10) |
| Artifact | geode-eval-artifacts@04ff1c4a1fee0cd1a3d837ad3a5f5239f1fd9acd/mcpmark/results-geode-agentworld/geode-gpt54-high-v1.0.12-f99cea63-20260803-mcpmark-filesystem-easy |
- Total task execution time 802.182s / average 80.218s
- 53 GEODE turns total / 5.3 average
- 302,984 input / 30,238 output tokens
- 182 canonical events / 56 exactly paired tool calls and results
- Failure: file_context/uppercase left file_01.txt incompletely uppercased
cd artifacts/eval/harnesses/mcpmark
PYTHONPATH=<geode-v1.0.12-release-tree> .venv/bin/python -m plugins.benchmark_harness.run_mcpmark --mcp filesystem --task-suite easy --models geode-gpt-5.4 --agent geode --reasoning-effort high --k 1 --timeout 1200 --exp-name geode-gpt54-high-v1.0.12-f99cea63-20260803-mcpmark-filesystem-easy --output-dir ./results-geode-v1012- The official verifier found all five output files, but file_01.txt was not fully uppercased; the failure is retained without retry.
- No authentication, quota, provider-adapter, MCP transport, or harness exception occurred.
- The v1.0.11 GPT-5.6 10/10 comparison is model-confounded and cannot be attributed to the runtime release alone.
- All ten trajectories are scope-complete and intentionally replay-incomplete; manifest and native receipts are pinned to artifact commit 04ff1c4.
filesystem/easy GPT-5.6 v1.0.11 release regression2026-07-31 KSTgpt-5.6-solsubscriptionhigh
| Status | complete |
| Suite/domain | filesystem/easy |
| Model | gpt-5.6-sol |
| Provider | openai |
| Source | subscription |
| Effort | high |
| Route | GEODE AgenticLoop MCPMark adapter |
| Harness | eval-sys/mcpmark@cd45b7f, GEODE v1.0.11@686ff372 |
| Accuracy | 100.0% (10 / 10) |
| Artifact | geode-eval-artifacts@16a54f08450db771c02e30c73bdc3867f6282f83/mcpmark/results-geode-agentworld/geode-gpt56-sol-high-v1011-686ff372-20260731-mcpmark-filesystem-easy |
- Total task execution time 596.580s / average 59.658s
- 56 GEODE turns total / 5.6 average
- 700,719 input / 12,164 output / 206,848 cache-read tokens
- Recorded estimate $2.937699; not subscription billing
- 226 canonical events / 78 exactly paired tool calls and results / 0 missing required turn IDs
cd artifacts/eval/harnesses/mcpmark
GEODE_HOME=<isolated-v1011-runtime-home> \
PYTHONPATH=<geode-v1.0.11-release-tree> \
.venv/bin/python -m plugins.benchmark_harness.run_mcpmark \
--mcp filesystem \
--task-suite easy \
--models gpt-5.6-sol \
--agent geode \
--reasoning-effort high \
--k 1 \
--timeout 1200 \
--exp-name geode-gpt56-sol-high-v1011-686ff372-20260731-mcpmark-filesystem-easy \
--output-dir ./results-geode-v1011- The earlier file_context/uppercase failure now passes; all ten official filesystem/easy verifiers are green.
- This remains directly comparable only to filesystem/easy, not to the MCPMark Verified standard aggregate.
- The stable geode.trajectory@1 release is scope-complete but intentionally replay-incomplete because dialogue and tool bodies are digested.
- Native receipts and stable trajectories are pinned to geode-eval-artifacts commit 16a54f0.
filesystem/easy GPT-5.6 subscription rerun2026-07-31 KSTgpt-5.6-solsubscriptionhigh
| Status | complete |
| Suite/domain | filesystem/easy |
| Model | gpt-5.6-sol |
| Provider | openai-codex |
| Source | subscription |
| Effort | high |
| Route | GEODE local MCPMark adapter |
| Harness | eval-sys/mcpmark@cd45b7f, GEODE@edb74602b |
| Accuracy | 90.0% (9 / 10) |
| Artifact | geode-eval-artifacts@9c00ecf/mcpmark/results-geode-agentworld/geode-gpt56-sol-high-edb74602b-20260731-mcpmark-filesystem-easy |
- Total task execution time 799.435s / average 79.943s
- 54 GEODE turns total / 5.4 average
- 799,679 input / 10,976 output / 97,792 cache-read tokens
- Recorded estimate $3.887611; not subscription billing
- Failure: file_context/uppercase left file_01.txt incompletely uppercased
cd artifacts/eval/harnesses/mcpmark
GEODE_HOME=<isolated-runtime-home> \
PYTHONPATH=<geode-edb74602b-worktree> \
OPENAI_API_KEY=dummy \
.venv/bin/python -m plugins.benchmark_harness.run_mcpmark \
--mcp filesystem \
--task-suite easy \
--models geode-gpt-5.6-sol \
--agent geode \
--reasoning-effort high \
--k 1 \
--timeout 1200 \
--exp-name geode-gpt56-sol-high-edb74602b-20260731-mcpmark-filesystem-easy \
--output-dir ./results-geode-edb74602b- This is directly comparable to filesystem/easy only, not to the MCPMark Verified standard aggregate.
- The upstream total_tokens and total_reasoning_tokens summary fields were zero despite populated input/output fields; they are not used.
- One response stream disconnected after the first task had already produced its files; that task passed every official integrity check and no 429 occurred.
- Raw receipts and ten normalized tool trajectories are pinned to geode-eval-artifacts commit 9c00ecf.
Verified github standard slice2026-07-04 KSTgpt-5.5subscriptionxhigh
| Status | complete |
| Suite/domain | github/standard |
| Model | gpt-5.5 |
| Provider | openai-codex |
| Source | subscription |
| Effort | xhigh |
| Route | GEODE local MCPMark adapter + GitHub MCP Docker server |
| Harness | eval-sys/mcpmark@cd45b7f, ghcr.io/github/github-mcp-server:v0.15.0 |
| Accuracy | 82.6% (19 / 23) |
| Artifact | artifacts/eval/harnesses/mcpmark/results-geode-agentworld/geode-gpt55-xhigh-20260704-mcpmark-verified-github* |
- Total task execution time 16476.3s
- Average task execution time 716.4s
- Failures: claude-code/label_color_standardization, mcpmark-cicd/deployment_status_workflow, missing-semester/assign_contributor_labels, missing-semester/find_salient_file
- All transient GitHub repositories were deleted by MCPMark cleanup.
cd artifacts/eval/harnesses/mcpmark
# Run each available MCP service through the GEODE adapter.
GEODE_REPO_ROOT=<geode-worktree> \
PYTHONPATH=<geode-worktree>:<geode-site-packages> \
GITHUB_EVAL_ORG=mangowhoiscloud \
GEODE_MCPMARK_GITHUB_REPO_VISIBILITY=public \
.venv/bin/python pipeline.py \
--mcp github \
--task-suite standard \
--models geode-gpt-5.5 \
--agent geode \
--reasoning-effort xhigh \
--k 1 \
--timeout 1500 \
--exp-name geode-gpt55-xhigh-20260704-mcpmark-verified-<service> \
--output-dir ./results-geode-agentworld- The first label_color_standardization record is a fixture setup failure from GitHub state duplication; the retry produced an agent-level verification failure.
- The assign_contributor_labels failure used suffixed transient usernames in labels instead of canonical contributor labels.
- The find_salient_file failure did not create ANSWER.md on the required master branch.
Verified postgres standard slice2026-07-04 KSTgpt-5.5subscriptionxhigh
| Status | complete |
| Suite/domain | postgres/standard |
| Model | gpt-5.5 |
| Provider | openai-codex |
| Source | subscription |
| Effort | xhigh |
| Route | GEODE local MCPMark adapter + postgres-mcp |
| Harness | eval-sys/mcpmark@cd45b7f, postgres-mcp==0.3.0 |
| Accuracy | 95.2% (20 / 21) |
| Artifact | artifacts/eval/harnesses/mcpmark/results-geode-agentworld/geode-gpt55-xhigh-20260704-mcpmark-verified-postgres |
- Total task execution time 8765.7s
- Average task execution time 417.4s
- Failure: employees/employee_performance_analysis
cd artifacts/eval/harnesses/mcpmark
# Run each available MCP service through the GEODE adapter.
GEODE_REPO_ROOT=<geode-worktree> \
PYTHONPATH=<geode-worktree>:<geode-site-packages> \
GITHUB_EVAL_ORG=mangowhoiscloud \
GEODE_MCPMARK_GITHUB_REPO_VISIBILITY=public \
.venv/bin/python pipeline.py \
--mcp postgres \
--task-suite standard \
--models geode-gpt-5.5 \
--agent geode \
--reasoning-effort xhigh \
--k 1 \
--timeout 1500 \
--exp-name geode-gpt55-xhigh-20260704-mcpmark-verified-<service> \
--output-dir ./results-geode-agentworld- The GEODE adapter overrides MCPMark's default postgres server with postgres-mcp==0.3.0 in unrestricted mode.
- A final NoEventLoopError appeared during async cleanup after result writing; it did not affect the recorded verifier result.
Verified filesystem standard slice2026-07-04 KSTgpt-5.5subscriptionxhigh
| Status | complete |
| Suite/domain | filesystem/standard |
| Model | gpt-5.5 |
| Provider | openai-codex |
| Source | subscription |
| Effort | xhigh |
| Route | GEODE local MCPMark adapter |
| Harness | eval-sys/mcpmark@cd45b7f |
| Accuracy | 83.3% (25 / 30) |
| Artifact | artifacts/eval/harnesses/mcpmark/results-geode-agentworld/geode-gpt55-xhigh-20260704-mcpmark-verified-filesystem-* |
- Recorded task execution time 13580.6s over 29 recorded tasks
- Average recorded task execution time 468.3s
- Failures: desktop_template/budget_computation, papers/author_folders, papers/find_math_paper, student_database/english_talent, threestudio/output_analysis
cd artifacts/eval/harnesses/mcpmark
# Run each available MCP service through the GEODE adapter.
GEODE_REPO_ROOT=<geode-worktree> \
PYTHONPATH=<geode-worktree>:<geode-site-packages> \
GITHUB_EVAL_ORG=mangowhoiscloud \
GEODE_MCPMARK_GITHUB_REPO_VISIBILITY=public \
.venv/bin/python pipeline.py \
--mcp filesystem \
--task-suite standard \
--models geode-gpt-5.5 \
--agent geode \
--reasoning-effort xhigh \
--k 1 \
--timeout 1500 \
--exp-name geode-gpt55-xhigh-20260704-mcpmark-verified-<service> \
--output-dir ./results-geode-agentworld- filesystem/standard is a materially harder slice than filesystem/easy.
- papers/author_folders is counted as a failed no-result transport run because both attempts hung before meta output.
- The adapter now aliases file_path to path when the MCP schema expects path, which fixed write_file failures seen in the first filesystem pass.
filesystem/easy category-parallel rerun2026-07-03 05:11 KSTgpt-5.5subscriptionxhigh
| Status | complete |
| Suite/domain | filesystem/easy |
| Model | gpt-5.5 |
| Provider | openai-codex |
| Source | subscription |
| Effort | xhigh |
| Route | GEODE local MCPMark adapter, category-parallel execution |
| Harness | eval-sys/mcpmark@cd45b7f |
| Accuracy | 100.0% (10 / 10) |
| Artifact | artifacts/eval/harnesses/mcpmark/results-geode-live/geode-gpt55-xhigh-20260703-ledger-* |
- Total task execution time 1360.129s
- Average task execution time 136.013s
- 40 GEODE rounds total / 4.0 average
- 429,324 total tokens
- Category rows: file_context 3/3, file_property 2/2, folder_structure 1/1, legal_document 1/1, papers 1/1, student_database 2/2
cd artifacts/eval/harnesses/mcpmark
for category in file_context file_property folder_structure legal_document papers student_database; do
GEODE_REPO_ROOT=<geode-worktree> \
OPENAI_API_KEY=dummy \
FILESYSTEM_TEST_ROOT=./test_environments \
.venv/bin/python pipeline.py \
--mcp filesystem \
--task-suite easy \
--tasks "$category" \
--models geode-gpt-5.5 \
--agent geode \
--reasoning-effort xhigh \
--k 1 \
--timeout 900 \
--exp-name "geode-gpt55-xhigh-20260703-ledger-$category" \
--output-dir ./results-geode-live &
done
wait- This rerun split filesystem/easy by category and executed the six categories in parallel.
- Only filesystem was runnable in the current local environment without additional credentials or Docker services.
- GitHub, Notion, Playwright, and Postgres MCPMark columns remain blocked until their service prerequisites are provisioned.
filesystem/easy full slice2026-07-03 KSTgpt-5.5subscriptionxhigh
| Status | complete |
| Suite/domain | filesystem/easy |
| Model | gpt-5.5 |
| Provider | openai-codex |
| Source | subscription |
| Effort | xhigh |
| Route | GEODE local MCPMark adapter |
| Harness | eval-sys/mcpmark@cd45b7f |
| Accuracy | 100.0% (10 / 10) |
| Artifact | artifacts/eval/harnesses/mcpmark/results-geode-live/geode-gpt55-xhigh-20260703-filesystem-easy/geode-gpt-5-5-xhigh__filesystem-easy/run-1 |
- Total task execution time 1706.044s
- Average task execution time 170.604s
- 40 GEODE rounds total / 4.0 average
- 266,779 total tokens
cd artifacts/eval/harnesses/mcpmark
GEODE_REPO_ROOT=<geode-worktree> \
OPENAI_API_KEY=dummy \
.venv/bin/python pipeline.py \
--mcp filesystem \
--task-suite easy \
--models geode-gpt-5.5 \
--agent geode \
--reasoning-effort xhigh \
--k 1 \
--timeout 900 \
--exp-name geode-gpt55-xhigh-20260703-filesystem-easy \
--output-dir ./results-geode-live- MCPMark filesystem/easy is directly comparable only to the same subset.
- This is not the MCPMark Verified aggregate used by frontier leaderboards.
- OPENAI_API_KEY=dummy satisfied the harness environment check; model calls used the GEODE subscription route.
notion unblock smoke (easy, single task)2026-07-10 KSTgpt-5.5subscriptionxhigh
| Status | complete |
| Suite/domain | notion/easy |
| Model | gpt-5.5 |
| Provider | openai-codex |
| Source | subscription |
| Effort | xhigh |
| Route | GEODE local MCPMark adapter |
| Harness | eval-sys/mcpmark@cd45b7f |
| Accuracy | 1 / 1 |
| Artifact | artifacts/eval/harnesses/mcpmark/results-geode-agentworld/geode-gpt55-xhigh-20260710-notion-smoke-unblock-r2/geode-gpt-5-5-xhigh__notion-easy/run-1 |
- State duplication 58.9s; agent 216.8s over 8 rounds; 62.8k input / 8.0k output tokens.
- The 2026-07-04 stall was an expired browser session: duplication page.goto to app.notion.com timed out at 120s per retry.
- Re-login used a real-Chrome-channel persistent context (Google OAuth rejects automation-flagged browsers); the session cookie lives on .app.notion.com.
set -a; source .mcp_env; set +a
OPENAI_API_KEY=dummy \
.venv/bin/python -m plugins.benchmark_harness.run_mcpmark \
--mcp notion \
--task-suite easy \
--tasks toronto_guide/simple__change_color \
--models geode-gpt-5.5 \
--agent geode \
--reasoning-effort xhigh- Verifier-backed single-task smoke proving the notion service is runnable end to end; not a notion standard score.
- The task embeds a Notion API trap: updating a select option color returns validation_error; the agent passed by redefining options via a database schema update.
notion blocked prerequisite record2026-07-03 KSTgpt-5.5subscriptionxhigh
| Status | blocked |
| Suite/domain | notion/easy |
| Model | gpt-5.5 |
| Provider | openai-codex |
| Source | subscription |
| Effort | xhigh |
| Route | GEODE local MCPMark adapter |
| Harness | eval-sys/mcpmark@cd45b7f |
| Accuracy | blocked |
| Artifact | not created |
- No Notion MCPMark score was produced in this cycle.
- The harness requires source and evaluation Notion workspace credentials.
GEODE_REPO_ROOT=<geode-worktree> \
OPENAI_API_KEY=dummy \
.venv/bin/python pipeline.py \
--mcp notion \
--task-suite easy \
--models geode-gpt-5.5 \
--agent geode \
--reasoning-effort xhigh- Blocked before live execution because the local harness environment has no .mcp_env credentials.
- Record a measured score only after the Notion integration and paired workspaces are provisioned.
playwright blocked prerequisite record2026-07-03 KSTgpt-5.5subscriptionxhigh
| Status | blocked |
| Suite/domain | playwright/easy |
| Model | gpt-5.5 |
| Provider | openai-codex |
| Source | subscription |
| Effort | xhigh |
| Route | GEODE local MCPMark adapter |
| Harness | eval-sys/mcpmark@cd45b7f |
| Accuracy | blocked |
| Artifact | not created |
- No Playwright MCPMark score was produced in this cycle.
- Browser/WebArena service setup was not available in the local benchmark environment.
GEODE_REPO_ROOT=<geode-worktree> \
OPENAI_API_KEY=dummy \
.venv/bin/python pipeline.py \
--mcp playwright \
--task-suite easy \
--models geode-gpt-5.5 \
--agent geode \
--reasoning-effort xhigh- Blocked before live execution because the browser-backed service stack was not running.
- Record a measured score only after the browser environment is provisioned and health checked.
Run 로그
태스크별 meta.json(route, 소요시간, 토큰, verifier 결과)과 messages.json(최종 답변 문자열 또는 빈 목록 placeholder), 생성된 경우 execution.log(순서가 보존된 MCP action/result)는 민감한 로컬 경로를 마스킹한 공개용 copy로 geode-eval-artifacts 레포에 보존됩니다. 이 공개 snapshot에는 전체 model dialogue와 hidden turn이 없으므로 messages.json만으로 대화를 복원할 수 없습니다.
geode-eval-artifacts/mcpmark/results-geode-agentworld: Verified 트랙 run 디렉터리(geode-gpt55-xhigh-20260704-mcpmark-verified-*).geode-eval-artifacts/mcpmark/logs,geode-eval-artifacts/mcpmark/logs-cycle: 파이프라인 stdout 로그(state duplication, verification, cleanup 단계).
run 기록의 artifact 경로는 측정 당시 로컬 harness 경로입니다. 게시된 사본은 위 레포 경로에서 run 이름으로 찾습니다.