본문으로 이동
문서 탐색MCPMark
벤치마크레퍼런스

MCPMark

GEODE의 MCPMark 실측입니다. Verified available-services headline, 서비스 coverage와 blocker, run 기록 전체, 원본 run 로그 링크를 담습니다.

상태 기반 MCP 검증

점수·비용·증거 누락을 함께 봅니다

MCPMark Verified ↗

Available services

64 / 74

86.5% · filesystem + postgres + GitHub

historical slice

Gate 0C

23 / 30

Codex 21 / 30 · common deadline

diagnostic · k=1

Coverage gap

53 tasks

Notion 28 + Playwright 4 + WebArena 21

unmeasured

Gate 0C

59 admitted · 1 withheld

동일 filesystem/standard 30건과 공통 deadline을 사용했습니다. timeout trajectory 한 건은 점수에 남고 공개 admission에서는 보류됐습니다.

Gate 0C 원본 bundle ↗

Gate 0B

7 / 15 guard · 10 / 15 unlimited

25K result guard의 직접 ablation입니다. 세 반복 중 네 timeout은 withheld로 남겼고, 단일 treatment 진단에는 승격 권한을 주지 않았습니다.

Gate 0B 원본 bundle ↗

측정 기록의 발전

Service slices실행 가능한 세 service를 분리 집계하고 raw verifier 로그 보존
Trajectory @1tool call/result exact join과 task별 실행 표면 추가
Matched K=1동일 easy 10건에서 점수·token·agent time을 함께 기록
Invalidationtimeout 시작점 불일치를 발견해 비교 주장을 철회
Prospective gatesrun-spec → attempts → analysis → admission 순서를 고정

MCPMark Verified는 service별 pass@1과 turns·time·token·cost를 함께 열고, trajectory가 없는 제출을 명시합니다. GEODE 표면도 미측정 service와 withheld trajectory를 점수 옆에 보존합니다.

MCPMark는 실제 MCP 서버(filesystem, Postgres, GitHub, Notion, Playwright 등)를 대상으로 한 tool-use 벤치마크입니다. 태스크마다 독립 검증 스크립트가 결과 상태를 확인합니다. GEODE는 evals/benchmarksBaseMCPAgent 어댑터로 참가하고 upstream pipeline.py는 패치하지 않습니다. 점수는 harness commit, 서비스 집합, model route, timeout에 고정해서만 게시합니다.

2026-08-13 GPT-5.4 filesystem/standard 정정 관측

고정된 30개 filesystem/standard task를 GPT-5.4 subscription / effort high로 task별 paired 실행했습니다. GEODE는 21/30 (70.0%), Codex CLI는 20/30 (66.7%)로 GEODE가 1건 앞섰습니다. 60개 trajectory에 3,381 events가 보존됐고, 1,430 tool call/result가 정확히 pairing됐으며 orphan은 없습니다.

사후 source audit에서 원래 사전등록한 equal-hard-deadline 전제가 성립하지 않았음이 확인됐습니다. GEODE는 MCP setup 뒤의loop.arun만, Codex는 내부 MCP startup을 포함하는 process communication을 timed surface로 사용했습니다. 따라서 prospective hypothesis는 invalidated이며, 점수는 retrospective description으로만 남습니다. Native input 총합은 GEODE가 작았지만 cache 제외 입력은 4.20M 대 1.44M으로 더 컸으므로 token-efficiency도 주장하지 않습니다. 공개 bundle에는 정확한 runner가 없으므로 독립 실행 가능한 재현 패키지도 아닙니다.

2026-08-12 matched token-efficiency rerun

같은 GPT-5.4 subscription / effort high, 같은 pinned filesystem/easy 10건을 수정 전후로 대조했습니다. 점수는 9/10 (90.0%)로 유지됐고, 입력 토큰은 447,376에서 314,219로 29.8%, 출력 토큰은 25,157에서 20,385로 19.0% 줄었습니다.

10건 중 8건의 입력 토큰이 감소했고 round 수가 같은 4건도 12.5% 감소했습니다. 188개 canonical event와 54/54 exact tool pair에는 orphan이 없습니다. 단, 한 번의 matched trial이므로 MCPMark Verified 점수·신뢰구간·구독 과금 절감으로 일반화하지 않습니다.

2026-08-03 v1.0.12 post-release regression

공개 배포된 GEODE v1.0.12 (f99cea63)과 gpt-5.4 subscription / effort high filesystem/easy 10건을 실행했습니다. 공식 verifier는 9/10 (90.0%), 총 802.2초와 53 turns입니다. 실패한 file_context/uppercase는 다섯 파일을 모두 만들었지만 file_01.txt를 완전히 대문자로 바꾸지 못했습니다.

인증·quota·provider adapter·MCP transport 오류는 없었습니다. 10개 trajectory는 182개 canonical event와 56개 exact tool pair를 보존하며 scope_complete=true, replay_complete=false입니다. v1.0.11의 GPT-5.6 10/10과 비교할 때 모델까지 바뀌었으므로 release 회귀로 단정하지 않습니다.

2026-07-31 v1.0.11 release regression

배포된 GEODE v1.0.11 (686ff372)과 gpt-5.6-sol subscription / effort high filesystem/easy 10건을 재측정했습니다. 공식 verifier는 10/10 (100.0%), 총 596.6초와 56 turns입니다. 이전 edb74602b run의 유일한 실패였던 file_context/uppercase도 통과했습니다.

10개 stable trajectory의 226개 이벤트는 canonical SQLite 행과 ID·session·turn·call·kind까지 일치합니다. 78개 tool call/result가 모두 정확히 pairing됐고 필수 turn ID 누락은 0건입니다.

Headline: Verified available-services 트랙

2026-07-04 run, GEODE v0.99.269 계열, eval-sys/mcpmark@cd45b7f, gpt-5.5 Codex 구독 route, effort xhigh. 이 측정의 범위는 로컬에서 실행 가능했던 standard 슬라이스(filesystem, postgres, github)입니다.

FileGitHubNotionPlaywrightPostgresAvg.
83.3%
standard, 25 / 30
82.6%
standard, 19 / 23
unmeasured
unblocked 2026-07-10 (easy smoke 1/1); standard 28 tasks not yet measured
unmeasured
live-web subset runnable since 2026-07-10; WebArena subset needs ~100GB images (local disk exceeded)
95.2%
standard, 20 / 21
86.5%
Measured available services only: filesystem+postgres+github

Service coverage

ServiceEasyStandardAdapter 상태Blocker
filesystem1030standard run 완료historical 25 / 30; paired GPT-5.4 21 / 30
postgres1021standard run 완료20 / 21, postgres-mcp==0.3.0
github1023standard run 완료19 / 23, Docker GitHub MCP server. State Duplication Error 6건의 원인(GITHUB_EVAL_ORG 미영속)은 2026-07-10 제거
notion1028실측 가능 (easy smoke 1/1, 2026-07-10)07-04 스톨 원인은 브라우저 세션 만료로 확정, 재발급 절차 확립. standard 28건 미측정
playwright04실행 준비 완료 (2026-07-10)@playwright/mcp@0.0.68 기동 확인. 4건 미측정
playwright_webarena1021stdio adapter 준비WebArena Docker 이미지 실측 119GiB vs 로컬 여유 13GiB. 외장 볼륨 또는 VM 필요
insforge확인 필요조사 중INSFORGE_API_KEY, task manager 인자 호환성 확인 필요
supabase확인 필요미지원HTTP MCP transport. GEODE MCPServerManager는 현재 stdio 중심

구독 쿼터(429 usage_limit_reached)는 full-suite 연속 실행을 리셋 창 단위로 분할시킵니다. 429 실패는 점수에 포함하지 않고 해당 태스크를 재실행합니다.

Run 기록

Verified available-services aggregate2026-07-04 KSTgpt-5.5subscriptionxhigh
Statuscomplete
Suite/domainfilesystem + postgres + github / standard
Modelgpt-5.5
Provideropenai-codex
Sourcesubscription
Effortxhigh
RouteGEODE local MCPMark adapter
Harnesseval-sys/mcpmark@cd45b7f, GEODE feature/mcpmark-agentworld-run
Accuracy86.5% (64 / 74)
Artifactartifacts/eval/harnesses/mcpmark/results-geode-agentworld/geode-gpt55-xhigh-20260704-mcpmark-verified-*
  • Filesystem standard: 25 / 30, 83.3%
  • Postgres standard: 20 / 21, 95.2%
  • GitHub standard: 19 / 23, 82.6%
  • Recorded task execution time: filesystem 13580.6s over 29 recorded tasks, postgres 8765.7s, github 16476.3s
  • Notion was not included: no notion_state.json in the local harness environment.
  • Playwright/WebArena was not included: required Docker images/service stack were absent.
cd artifacts/eval/harnesses/mcpmark
# Run each available MCP service through the GEODE adapter.
GEODE_REPO_ROOT=<geode-worktree> \
PYTHONPATH=<geode-worktree>:<geode-site-packages> \
GITHUB_EVAL_ORG=mangowhoiscloud \
GEODE_MCPMARK_GITHUB_REPO_VISIBILITY=public \
.venv/bin/python pipeline.py \
  --mcp <filesystem|postgres|github> \
  --task-suite standard \
  --models geode-gpt-5.5 \
  --agent geode \
  --reasoning-effort xhigh \
  --k 1 \
  --timeout 1500 \
  --exp-name geode-gpt55-xhigh-20260704-mcpmark-verified-<service> \
  --output-dir ./results-geode-agentworld
  • This is not the full MCPMark Verified leaderboard aggregate. It covers only services that were runnable in the local environment: filesystem, postgres, and github.
  • The OpenAI model route was the GEODE Codex subscription route, not MCPMark's native LiteLLM OpenAI API route.
  • GitHub fixture repositories were made public during execution so the Docker GitHub MCP server could use normal public-repo semantics; all transient repos were deleted by cleanup.
  • The filesystem score counts papers/author_folders as a failed no-result transport run after two attempts without meta output.
filesystem/standard Gate 0C common-deadline diagnostic2026-08-14 KSTgpt-5.4subscriptionhigh
Statuscomplete
Suite/domainfilesystem/standard 30-task paired k=1
Modelgpt-5.4
Provideropenai / codex-cli
Sourcesubscription
Efforthigh
RouteGEODE AgenticLoop and isolated Codex CLI paired by task
Harnesseval-sys/mcpmark@cd45b7f, GEODE@f4b37604, Codex CLI 0.145.0@dad1db87
Diagnostic-only verifier pass rateGEODE 76.7% (23 / 30) · Codex 70.0% (21 / 30) · Δ +6.67 pp
Artifactgeode-eval-artifacts@1160fecfe4447f0a3f4cf30a414f29c61776d012/mcpmark/results-paired/mcpmark-gate0c-filesystem30-gpt54-high-20260813t190922z
  • Paired outcomes: 17 both-pass / 3 both-fail / 6 GEODE-only / 4 Codex-only
  • Exact token coverage: GEODE 29 / 30 attempts; Codex 30 / 30
  • All-arm action wall: GEODE 8,463.252s vs Codex 5,484.322s; runner envelopes 8,536.382s vs 5,548.725s
  • Native execution-log calls/errors: GEODE 644 / 51 vs Codex 678 / 17
  • Normalized trajectory attempts: GEODE 645 including one recovery projection vs Codex 678
  • Read/repeated-read references: GEODE 798 / 81 vs Codex 838 / 213
  • One GEODE score-bearing deadline expiration; its token usage is null and its scope-incomplete trajectory is withheld
# Requires the pinned MCPMark checkout and verifier patch, frozen run spec,
# isolated runtime home, and working GPT-5.4 subscription authentication.
.venv/bin/python -m plugins.benchmark_harness.run_mcpmark_pair   --profile filesystem30-geode-codex   --run-spec <frozen-run-spec.json>   --mcpmark-root artifacts/eval/harnesses/mcpmark   --output-dir <fresh-output-dir>   --python .venv/bin/python
  • The prospective -10 percentage-point threshold was supported, but promotion_authority remains none.
  • This is one direct paired repetition, not k=3 stability, a full MCPMark Verified headline, or an API-key leaderboard claim.
  • The common action deadline excludes fixture setup and the post-action verifier; runtime scaffolds and tool-result policies remain different.
  • Token totals have unmatched coverage and do not support a billing or token-efficiency claim.
  • Fifty-nine scope-complete trajectories are admitted; the one GEODE timeout trajectory remains withheld.
filesystem/standard Gate 0B tool-result-cap diagnostic2026-08-13–14 KSTgpt-5.4subscriptionhigh
Statuscomplete
Suite/domainfilesystem/standard targeted 5 tasks × 3 repetitions
Modelgpt-5.4
Provideropenai
Sourcesubscription
Efforthigh
RouteGEODE AgenticLoop paired 25K and unlimited tool-result-cap arms
Harnesseval-sys/mcpmark@cd45b7f, GEODE@02f71fae
Diagnostic-only verifier pass rate25K 46.7% (7 / 15) · unlimited 66.7% (10 / 15) · Δ +20.0 pp
Artifactgeode-eval-artifacts@17133f0c8e893b6d765fcef69712ba0867bd573a/mcpmark/results-paired/mcpmark-gate0b-tool-cap-gpt54-high-20260813t142345z
  • Observed fresh input: 25K 3,782,288 vs unlimited 2,202,725; exact token coverage is 13 / 15 attempts per arm
  • All-attempt action wall: 25K 6,076.699s vs unlimited 4,910.217s; runner envelopes 6,119.096s vs 4,951.086s
  • All-attempt MCP calls/errors: 25K 443 / 75 vs unlimited 255 / 27
  • Repeated read references: 25K 211 vs unlimited 63; tool-result truncations 15 vs 0
  • Two score-bearing deadline expirations per arm emitted no native token usage
# Requires the pinned MCPMark checkout and verifier patch, frozen run spec,
# isolated runtime home, and working GPT-5.4 subscription authentication.
.venv/bin/python -m plugins.benchmark_harness.run_mcpmark_pair   --profile max-tool-result-tokens   --run-spec <frozen-run-spec.json>   --mcpmark-root artifacts/eval/harnesses/mcpmark   --output-dir <fresh-output-dir>   --python .venv/bin/python
  • The prospective hypothesis was supported, but promotion_authority remains none.
  • This five-task, three-repetition subset is diagnostic and is not an MCPMark suite headline.
  • Token totals exclude the four timeout arms with empty native usage; wall time and MCP-call totals cover all 15 attempts per arm.
  • The prior infrastructure-invalid run contributes no denominator, and scope-incomplete timeout trajectories remain withheld.
filesystem/standard GPT-5.4 GEODE × Codex corrected observation2026-08-13 KSTgpt-5.4subscriptionhigh
Statuscomplete
Suite/domainfilesystem/standard
Modelgpt-5.4
Provideropenai / codex-cli
Sourcesubscription
Efforthigh
RouteGEODE AgenticLoop and isolated Codex CLI paired by task
Harnesseval-sys/mcpmark@cd45b7f, GEODE@a8f45f3c9, Codex@dad1db87
Retrospective descriptive verifier outcomesGEODE 70.0% (21 / 30) · Codex 66.7% (20 / 30)
Artifactgeode-eval-artifacts@e5d442f25c9fb4861e28744dbe924a36325c746b/mcpmark/results-paired/mcpmark-filesystem-standard-gpt54-high-geode-codex-k1-boundary-aligned-20260813
  • Observed delta +1 / 30 (+3.33 pp); prospective hypothesis invalidated
  • GEODE 1,646 events / 703 exact tool pairs; Codex 1,735 / 727
  • GEODE fresh input 4,204,759 vs Codex 1,444,927; no token-efficiency claim
  • Task wall: GEODE 7,842.4s vs Codex 6,970.2s (+12.5%); no causal efficiency claim
  • Both releases scope-complete, replay-incomplete, zero orphan tool events
# Illustrative one-task invocation only; exact runner is withheld.
python -m plugins.benchmark_harness.run_mcpmark   --mcp filesystem   --task-suite standard   --tasks <one-frozen-task-id>   --models <geode-gpt-5.4|codex-gpt-5.4>   --agent <geode|codex>   --reasoning-effort high   --k 1   --timeout 1200
  • All 30 pairs share the pinned task tree, fixture reset, model label, effort, and verifier.
  • The frozen equal-hard-deadline claim was invalidated: GEODE timed loop.arun, while Codex timed process communication including internal MCP startup.
  • Prompt, action budget, retry, compaction, and cache accounting are also not identical; outcomes are retrospective descriptions only.
  • The public bundle is not independently executable because the exact runner remains digest-bound but withheld.
  • This is one filesystem service slice, not the full 127-task MCPMark Verified leaderboard.
  • Raw messages, logs, metadata, and provider diagnostics remain private; only digest-reduced trajectories and validated sidecars are public.
filesystem/easy GPT-5.4 token-efficiency rerun2026-08-12 KSTgpt-5.4subscriptionhigh
Statuscomplete
Suite/domainfilesystem/easy
Modelgpt-5.4
Provideropenai
Sourcesubscription
Efforthigh
RouteGEODE AgenticLoop MCPMark adapter
Harnesseval-sys/mcpmark@cd45b7f, GEODE feature@149024e6e
Accuracy90.0% (9 / 10)
Artifactgeode-eval-artifacts@2c2d1f0621f64ff7ceeff8c05d8ebd3449501aaf/trajectories/mcpmark-geode-gpt54-high-token-efficiency-rerun-filesystem-easy-20260812T090254Z-35db8b275a36
  • Matched input tokens 447,376 → 314,219 (-29.8%)
  • Matched output tokens 25,157 → 20,385 (-19.0%)
  • Native reasoning tokens 14,174, included within output tokens
  • 188 canonical events / 54 exactly paired tool calls and results
  • Failure unchanged: file_context/uppercase exact-string mismatch
cd artifacts/eval/harnesses/mcpmark
PYTHONPATH=<geode-feature-tree> .venv/bin/python -m plugins.benchmark_harness.run_mcpmark   --mcp filesystem   --task-suite easy   --models geode-gpt-5.4   --agent geode   --reasoning-effort high   --k 1   --timeout 1200   --exp-name geode-gpt54-high-token-efficiency-20260812-rerun   --output-dir ./results-token-efficiency
  • The score matched the pre-repair GEODE baseline at 9/10 while input and output tokens fell materially.
  • Eight of ten tasks used fewer input tokens; the four tasks with identical round counts fell 12.5%.
  • This is one matched diagnostic trial, not MCPMark Verified, a confidence interval, or a subscription billing claim.
  • All ten public trajectories are scope-complete and intentionally replay-incomplete; the immutable release was read back from artifact main.
filesystem/easy GPT-5.4 v1.0.12 post-release regression2026-08-03 KSTgpt-5.4subscriptionhigh
Statuscomplete
Suite/domainfilesystem/easy
Modelgpt-5.4
Provideropenai
Sourcesubscription
Efforthigh
RouteGEODE AgenticLoop MCPMark adapter
Harnesseval-sys/mcpmark@cd45b7f, GEODE v1.0.12@f99cea63
Accuracy90.0% (9 / 10)
Artifactgeode-eval-artifacts@04ff1c4a1fee0cd1a3d837ad3a5f5239f1fd9acd/mcpmark/results-geode-agentworld/geode-gpt54-high-v1.0.12-f99cea63-20260803-mcpmark-filesystem-easy
  • Total task execution time 802.182s / average 80.218s
  • 53 GEODE turns total / 5.3 average
  • 302,984 input / 30,238 output tokens
  • 182 canonical events / 56 exactly paired tool calls and results
  • Failure: file_context/uppercase left file_01.txt incompletely uppercased
cd artifacts/eval/harnesses/mcpmark
PYTHONPATH=<geode-v1.0.12-release-tree> .venv/bin/python -m plugins.benchmark_harness.run_mcpmark   --mcp filesystem   --task-suite easy   --models geode-gpt-5.4   --agent geode   --reasoning-effort high   --k 1   --timeout 1200   --exp-name geode-gpt54-high-v1.0.12-f99cea63-20260803-mcpmark-filesystem-easy   --output-dir ./results-geode-v1012
  • The official verifier found all five output files, but file_01.txt was not fully uppercased; the failure is retained without retry.
  • No authentication, quota, provider-adapter, MCP transport, or harness exception occurred.
  • The v1.0.11 GPT-5.6 10/10 comparison is model-confounded and cannot be attributed to the runtime release alone.
  • All ten trajectories are scope-complete and intentionally replay-incomplete; manifest and native receipts are pinned to artifact commit 04ff1c4.
filesystem/easy GPT-5.6 v1.0.11 release regression2026-07-31 KSTgpt-5.6-solsubscriptionhigh
Statuscomplete
Suite/domainfilesystem/easy
Modelgpt-5.6-sol
Provideropenai
Sourcesubscription
Efforthigh
RouteGEODE AgenticLoop MCPMark adapter
Harnesseval-sys/mcpmark@cd45b7f, GEODE v1.0.11@686ff372
Accuracy100.0% (10 / 10)
Artifactgeode-eval-artifacts@16a54f08450db771c02e30c73bdc3867f6282f83/mcpmark/results-geode-agentworld/geode-gpt56-sol-high-v1011-686ff372-20260731-mcpmark-filesystem-easy
  • Total task execution time 596.580s / average 59.658s
  • 56 GEODE turns total / 5.6 average
  • 700,719 input / 12,164 output / 206,848 cache-read tokens
  • Recorded estimate $2.937699; not subscription billing
  • 226 canonical events / 78 exactly paired tool calls and results / 0 missing required turn IDs
cd artifacts/eval/harnesses/mcpmark
GEODE_HOME=<isolated-v1011-runtime-home> \
PYTHONPATH=<geode-v1.0.11-release-tree> \
.venv/bin/python -m plugins.benchmark_harness.run_mcpmark \
  --mcp filesystem \
  --task-suite easy \
  --models gpt-5.6-sol \
  --agent geode \
  --reasoning-effort high \
  --k 1 \
  --timeout 1200 \
  --exp-name geode-gpt56-sol-high-v1011-686ff372-20260731-mcpmark-filesystem-easy \
  --output-dir ./results-geode-v1011
  • The earlier file_context/uppercase failure now passes; all ten official filesystem/easy verifiers are green.
  • This remains directly comparable only to filesystem/easy, not to the MCPMark Verified standard aggregate.
  • The stable geode.trajectory@1 release is scope-complete but intentionally replay-incomplete because dialogue and tool bodies are digested.
  • Native receipts and stable trajectories are pinned to geode-eval-artifacts commit 16a54f0.
filesystem/easy GPT-5.6 subscription rerun2026-07-31 KSTgpt-5.6-solsubscriptionhigh
Statuscomplete
Suite/domainfilesystem/easy
Modelgpt-5.6-sol
Provideropenai-codex
Sourcesubscription
Efforthigh
RouteGEODE local MCPMark adapter
Harnesseval-sys/mcpmark@cd45b7f, GEODE@edb74602b
Accuracy90.0% (9 / 10)
Artifactgeode-eval-artifacts@9c00ecf/mcpmark/results-geode-agentworld/geode-gpt56-sol-high-edb74602b-20260731-mcpmark-filesystem-easy
  • Total task execution time 799.435s / average 79.943s
  • 54 GEODE turns total / 5.4 average
  • 799,679 input / 10,976 output / 97,792 cache-read tokens
  • Recorded estimate $3.887611; not subscription billing
  • Failure: file_context/uppercase left file_01.txt incompletely uppercased
cd artifacts/eval/harnesses/mcpmark
GEODE_HOME=<isolated-runtime-home> \
PYTHONPATH=<geode-edb74602b-worktree> \
OPENAI_API_KEY=dummy \
.venv/bin/python -m plugins.benchmark_harness.run_mcpmark \
  --mcp filesystem \
  --task-suite easy \
  --models geode-gpt-5.6-sol \
  --agent geode \
  --reasoning-effort high \
  --k 1 \
  --timeout 1200 \
  --exp-name geode-gpt56-sol-high-edb74602b-20260731-mcpmark-filesystem-easy \
  --output-dir ./results-geode-edb74602b
  • This is directly comparable to filesystem/easy only, not to the MCPMark Verified standard aggregate.
  • The upstream total_tokens and total_reasoning_tokens summary fields were zero despite populated input/output fields; they are not used.
  • One response stream disconnected after the first task had already produced its files; that task passed every official integrity check and no 429 occurred.
  • Raw receipts and ten normalized tool trajectories are pinned to geode-eval-artifacts commit 9c00ecf.
Verified github standard slice2026-07-04 KSTgpt-5.5subscriptionxhigh
Statuscomplete
Suite/domaingithub/standard
Modelgpt-5.5
Provideropenai-codex
Sourcesubscription
Effortxhigh
RouteGEODE local MCPMark adapter + GitHub MCP Docker server
Harnesseval-sys/mcpmark@cd45b7f, ghcr.io/github/github-mcp-server:v0.15.0
Accuracy82.6% (19 / 23)
Artifactartifacts/eval/harnesses/mcpmark/results-geode-agentworld/geode-gpt55-xhigh-20260704-mcpmark-verified-github*
  • Total task execution time 16476.3s
  • Average task execution time 716.4s
  • Failures: claude-code/label_color_standardization, mcpmark-cicd/deployment_status_workflow, missing-semester/assign_contributor_labels, missing-semester/find_salient_file
  • All transient GitHub repositories were deleted by MCPMark cleanup.
cd artifacts/eval/harnesses/mcpmark
# Run each available MCP service through the GEODE adapter.
GEODE_REPO_ROOT=<geode-worktree> \
PYTHONPATH=<geode-worktree>:<geode-site-packages> \
GITHUB_EVAL_ORG=mangowhoiscloud \
GEODE_MCPMARK_GITHUB_REPO_VISIBILITY=public \
.venv/bin/python pipeline.py \
  --mcp github \
  --task-suite standard \
  --models geode-gpt-5.5 \
  --agent geode \
  --reasoning-effort xhigh \
  --k 1 \
  --timeout 1500 \
  --exp-name geode-gpt55-xhigh-20260704-mcpmark-verified-<service> \
  --output-dir ./results-geode-agentworld
  • The first label_color_standardization record is a fixture setup failure from GitHub state duplication; the retry produced an agent-level verification failure.
  • The assign_contributor_labels failure used suffixed transient usernames in labels instead of canonical contributor labels.
  • The find_salient_file failure did not create ANSWER.md on the required master branch.
Verified postgres standard slice2026-07-04 KSTgpt-5.5subscriptionxhigh
Statuscomplete
Suite/domainpostgres/standard
Modelgpt-5.5
Provideropenai-codex
Sourcesubscription
Effortxhigh
RouteGEODE local MCPMark adapter + postgres-mcp
Harnesseval-sys/mcpmark@cd45b7f, postgres-mcp==0.3.0
Accuracy95.2% (20 / 21)
Artifactartifacts/eval/harnesses/mcpmark/results-geode-agentworld/geode-gpt55-xhigh-20260704-mcpmark-verified-postgres
  • Total task execution time 8765.7s
  • Average task execution time 417.4s
  • Failure: employees/employee_performance_analysis
cd artifacts/eval/harnesses/mcpmark
# Run each available MCP service through the GEODE adapter.
GEODE_REPO_ROOT=<geode-worktree> \
PYTHONPATH=<geode-worktree>:<geode-site-packages> \
GITHUB_EVAL_ORG=mangowhoiscloud \
GEODE_MCPMARK_GITHUB_REPO_VISIBILITY=public \
.venv/bin/python pipeline.py \
  --mcp postgres \
  --task-suite standard \
  --models geode-gpt-5.5 \
  --agent geode \
  --reasoning-effort xhigh \
  --k 1 \
  --timeout 1500 \
  --exp-name geode-gpt55-xhigh-20260704-mcpmark-verified-<service> \
  --output-dir ./results-geode-agentworld
  • The GEODE adapter overrides MCPMark's default postgres server with postgres-mcp==0.3.0 in unrestricted mode.
  • A final NoEventLoopError appeared during async cleanup after result writing; it did not affect the recorded verifier result.
Verified filesystem standard slice2026-07-04 KSTgpt-5.5subscriptionxhigh
Statuscomplete
Suite/domainfilesystem/standard
Modelgpt-5.5
Provideropenai-codex
Sourcesubscription
Effortxhigh
RouteGEODE local MCPMark adapter
Harnesseval-sys/mcpmark@cd45b7f
Accuracy83.3% (25 / 30)
Artifactartifacts/eval/harnesses/mcpmark/results-geode-agentworld/geode-gpt55-xhigh-20260704-mcpmark-verified-filesystem-*
  • Recorded task execution time 13580.6s over 29 recorded tasks
  • Average recorded task execution time 468.3s
  • Failures: desktop_template/budget_computation, papers/author_folders, papers/find_math_paper, student_database/english_talent, threestudio/output_analysis
cd artifacts/eval/harnesses/mcpmark
# Run each available MCP service through the GEODE adapter.
GEODE_REPO_ROOT=<geode-worktree> \
PYTHONPATH=<geode-worktree>:<geode-site-packages> \
GITHUB_EVAL_ORG=mangowhoiscloud \
GEODE_MCPMARK_GITHUB_REPO_VISIBILITY=public \
.venv/bin/python pipeline.py \
  --mcp filesystem \
  --task-suite standard \
  --models geode-gpt-5.5 \
  --agent geode \
  --reasoning-effort xhigh \
  --k 1 \
  --timeout 1500 \
  --exp-name geode-gpt55-xhigh-20260704-mcpmark-verified-<service> \
  --output-dir ./results-geode-agentworld
  • filesystem/standard is a materially harder slice than filesystem/easy.
  • papers/author_folders is counted as a failed no-result transport run because both attempts hung before meta output.
  • The adapter now aliases file_path to path when the MCP schema expects path, which fixed write_file failures seen in the first filesystem pass.
filesystem/easy category-parallel rerun2026-07-03 05:11 KSTgpt-5.5subscriptionxhigh
Statuscomplete
Suite/domainfilesystem/easy
Modelgpt-5.5
Provideropenai-codex
Sourcesubscription
Effortxhigh
RouteGEODE local MCPMark adapter, category-parallel execution
Harnesseval-sys/mcpmark@cd45b7f
Accuracy100.0% (10 / 10)
Artifactartifacts/eval/harnesses/mcpmark/results-geode-live/geode-gpt55-xhigh-20260703-ledger-*
  • Total task execution time 1360.129s
  • Average task execution time 136.013s
  • 40 GEODE rounds total / 4.0 average
  • 429,324 total tokens
  • Category rows: file_context 3/3, file_property 2/2, folder_structure 1/1, legal_document 1/1, papers 1/1, student_database 2/2
cd artifacts/eval/harnesses/mcpmark
for category in file_context file_property folder_structure legal_document papers student_database; do
  GEODE_REPO_ROOT=<geode-worktree> \
  OPENAI_API_KEY=dummy \
  FILESYSTEM_TEST_ROOT=./test_environments \
  .venv/bin/python pipeline.py \
    --mcp filesystem \
    --task-suite easy \
    --tasks "$category" \
    --models geode-gpt-5.5 \
    --agent geode \
    --reasoning-effort xhigh \
    --k 1 \
    --timeout 900 \
    --exp-name "geode-gpt55-xhigh-20260703-ledger-$category" \
    --output-dir ./results-geode-live &
done
wait
  • This rerun split filesystem/easy by category and executed the six categories in parallel.
  • Only filesystem was runnable in the current local environment without additional credentials or Docker services.
  • GitHub, Notion, Playwright, and Postgres MCPMark columns remain blocked until their service prerequisites are provisioned.
filesystem/easy full slice2026-07-03 KSTgpt-5.5subscriptionxhigh
Statuscomplete
Suite/domainfilesystem/easy
Modelgpt-5.5
Provideropenai-codex
Sourcesubscription
Effortxhigh
RouteGEODE local MCPMark adapter
Harnesseval-sys/mcpmark@cd45b7f
Accuracy100.0% (10 / 10)
Artifactartifacts/eval/harnesses/mcpmark/results-geode-live/geode-gpt55-xhigh-20260703-filesystem-easy/geode-gpt-5-5-xhigh__filesystem-easy/run-1
  • Total task execution time 1706.044s
  • Average task execution time 170.604s
  • 40 GEODE rounds total / 4.0 average
  • 266,779 total tokens
cd artifacts/eval/harnesses/mcpmark
GEODE_REPO_ROOT=<geode-worktree> \
OPENAI_API_KEY=dummy \
.venv/bin/python pipeline.py \
  --mcp filesystem \
  --task-suite easy \
  --models geode-gpt-5.5 \
  --agent geode \
  --reasoning-effort xhigh \
  --k 1 \
  --timeout 900 \
  --exp-name geode-gpt55-xhigh-20260703-filesystem-easy \
  --output-dir ./results-geode-live
  • MCPMark filesystem/easy is directly comparable only to the same subset.
  • This is not the MCPMark Verified aggregate used by frontier leaderboards.
  • OPENAI_API_KEY=dummy satisfied the harness environment check; model calls used the GEODE subscription route.
notion unblock smoke (easy, single task)2026-07-10 KSTgpt-5.5subscriptionxhigh
Statuscomplete
Suite/domainnotion/easy
Modelgpt-5.5
Provideropenai-codex
Sourcesubscription
Effortxhigh
RouteGEODE local MCPMark adapter
Harnesseval-sys/mcpmark@cd45b7f
Accuracy1 / 1
Artifactartifacts/eval/harnesses/mcpmark/results-geode-agentworld/geode-gpt55-xhigh-20260710-notion-smoke-unblock-r2/geode-gpt-5-5-xhigh__notion-easy/run-1
  • State duplication 58.9s; agent 216.8s over 8 rounds; 62.8k input / 8.0k output tokens.
  • The 2026-07-04 stall was an expired browser session: duplication page.goto to app.notion.com timed out at 120s per retry.
  • Re-login used a real-Chrome-channel persistent context (Google OAuth rejects automation-flagged browsers); the session cookie lives on .app.notion.com.
set -a; source .mcp_env; set +a
OPENAI_API_KEY=dummy \
.venv/bin/python -m plugins.benchmark_harness.run_mcpmark \
  --mcp notion \
  --task-suite easy \
  --tasks toronto_guide/simple__change_color \
  --models geode-gpt-5.5 \
  --agent geode \
  --reasoning-effort xhigh
  • Verifier-backed single-task smoke proving the notion service is runnable end to end; not a notion standard score.
  • The task embeds a Notion API trap: updating a select option color returns validation_error; the agent passed by redefining options via a database schema update.
notion blocked prerequisite record2026-07-03 KSTgpt-5.5subscriptionxhigh
Statusblocked
Suite/domainnotion/easy
Modelgpt-5.5
Provideropenai-codex
Sourcesubscription
Effortxhigh
RouteGEODE local MCPMark adapter
Harnesseval-sys/mcpmark@cd45b7f
Accuracyblocked
Artifactnot created
  • No Notion MCPMark score was produced in this cycle.
  • The harness requires source and evaluation Notion workspace credentials.
GEODE_REPO_ROOT=<geode-worktree> \
OPENAI_API_KEY=dummy \
.venv/bin/python pipeline.py \
  --mcp notion \
  --task-suite easy \
  --models geode-gpt-5.5 \
  --agent geode \
  --reasoning-effort xhigh
  • Blocked before live execution because the local harness environment has no .mcp_env credentials.
  • Record a measured score only after the Notion integration and paired workspaces are provisioned.
playwright blocked prerequisite record2026-07-03 KSTgpt-5.5subscriptionxhigh
Statusblocked
Suite/domainplaywright/easy
Modelgpt-5.5
Provideropenai-codex
Sourcesubscription
Effortxhigh
RouteGEODE local MCPMark adapter
Harnesseval-sys/mcpmark@cd45b7f
Accuracyblocked
Artifactnot created
  • No Playwright MCPMark score was produced in this cycle.
  • Browser/WebArena service setup was not available in the local benchmark environment.
GEODE_REPO_ROOT=<geode-worktree> \
OPENAI_API_KEY=dummy \
.venv/bin/python pipeline.py \
  --mcp playwright \
  --task-suite easy \
  --models geode-gpt-5.5 \
  --agent geode \
  --reasoning-effort xhigh
  • Blocked before live execution because the browser-backed service stack was not running.
  • Record a measured score only after the browser environment is provisioned and health checked.

Run 로그

태스크별 meta.json(route, 소요시간, 토큰, verifier 결과)과 messages.json(최종 답변 문자열 또는 빈 목록 placeholder), 생성된 경우 execution.log(순서가 보존된 MCP action/result)는 민감한 로컬 경로를 마스킹한 공개용 copy로 geode-eval-artifacts 레포에 보존됩니다. 이 공개 snapshot에는 전체 model dialogue와 hidden turn이 없으므로 messages.json만으로 대화를 복원할 수 없습니다.

run 기록의 artifact 경로는 측정 당시 로컬 harness 경로입니다. 게시된 사본은 위 레포 경로에서 run 이름으로 찾습니다.