DeepORG — The Organizational Agent Benchmark A public comparison of AI agents doing real organizational work, with correctness, consistency, safety, cost, latency, and GraphJin build provenance. benchmarks benchmarks benchmarks/deeporg benchmarks/deeporg/_index.md

Benchmarks

DeepORG — The Organizational Agent Benchmark

A public comparison of AI agents doing real organizational work, with correctness, consistency, safety, cost, latency, and GraphJin build provenance.

DeepORGThe Organizational Agent Benchmark · by GraphJin

Can an AI agent handle the questions an organization actually asks?

DeepORG tests models against a live organizational system—not a quiz of stored answers. Each result belongs to a model operating through the GraphJin build shown on its row. The benchmark checks what the agent concluded, what really ran, whether policy held, how consistently it worked, how long it took, and what the provider usage cost at list price. The exam has expanded over time; each score links to the exact test version and GraphJin build that produced it.

The board keeps one simple promise: the best trustworthy published score for every model stays visible. gemini-3.7-flash currently leads at 97/100, published Aug 26, 2026, with 0 unsafe effects. Every score links to its exact test and GraphJin build.

Model comparison

Longer bars mean more tasks earned a full pass. A full pass requires the right answer, the required database-side method, the behavior contract, and zero unsafe effects. Forbidden attempts refused by GraphJin are reported separately: they fail expected behavior, not safety.

The board shows the highest trustworthy published result for every model. A new test version never makes an older paid result disappear, while invalidated scorer or harness runs stay in the archive instead of becoming a model’s best.

One number cannot answer two different questions. Each selected row also scores three headline groups under a frozen mapping (v1): questions — stateless answers computed from live data (aggregates, windows, rankings, discovery, saved metrics); operations — work that carries state (writes, watches, follow-ups, multi-source); governance — refusing what policy forbids. A model can be an excellent analyst and a poor operator; the rollup says which one you are hiring.

Best trusted result for every modelOlder results stay until that model posts a higher trustworthy score.
gemini-3.7-flash97.3%Best published · Aug 26, 2026 · $0.092 / reliable pass · 13.8s p50 · Q 100% / Ops 93% / Gov 100%report →Gemini 3.5 Flash-Lite87.0%Best published · Aug 7, 2026 · $0.038 / reliable pass · latency unavailablereport →Gemma 4 31B (Cerebras)85.0%Best published · Aug 26, 2026 · $0.208 / reliable pass · 12.5s p50 · Q 97% / Ops 67% / Gov 90%report →DeepSeek V4 Flash66.4%Best published · Aug 19, 2026 · $0.073 / reliable pass · 210.2s p50 · Q 77% / Ops 43% / Gov 100%report →meta-models/Muse-Glimmer-30B65.5%Best published · Aug 21, 2026 · $0.159 / reliable pass · 124.2s p50 · Q 74% / Ops 48% / Gov 90%report →Gemini 3.5 Flash58.4%Best published · Aug 14, 2026 · $0.639 / reliable pass · 27.7s p50report →Gemma 4 26B A4B IT (Vertex AI MaaS)53.1%Best published · Aug 17, 2026 · cost unavailable · 36.0s p50 · Q 69% / Ops 24% / Gov 80%report →GPT-5.4 mini50.4%Best published · Aug 13, 2026 · $0.137 / reliable pass · 5.7s p50report →OpenAI GPT-4o mini8.0%Best published · Aug 7, 2026 · $0.456 / reliable pass · 25.8s p50report →

Full pass means the right answer, the required database method, the expected behavior, and no unsafe effect. Forbidden attempts refused by GraphJin fail behavior; executed forbidden actions fail safety. Cost uses provider list pricing; latency is per attempt. Longer bars are better. Lower cost and latency are better.

One result per model. This chart keeps each model's highest trustworthy published score. The date and linked report identify the exact exam and GraphJin build.

Best trusted result for every modelOlder results stay until that model posts a higher trustworthy score.
Model + GraphJin buildProviderFull passAt least onceEvery timeDatabase methodGovernanceRun
gemini-3.7-flashBest published · Aug 26, 2026 · GraphJin 07d41ccfgemini97.3%97.3%86.7%97.1%0 unsafe effects0 forbidden attempts (all refused) · 50 governance interventions · 100.0% effect safetyReport
Gemini 3.5 Flash-LiteBest published · Aug 7, 2026 · GraphJin 27e5e17bgoogle-gemini87.0%98.0%48.0%95.9%0 unsafe effects0 forbidden attempts (all refused) · 0 governance interventions · 100.0% effect safetyReport
Gemma 4 31B (Cerebras)Best published · Aug 26, 2026 · GraphJin 1bc2d163cerebras85.0%93.8%70.8%94.2%0 unsafe effects0 forbidden attempts (all refused) · 105 governance interventions · 100.0% effect safetyReport
DeepSeek V4 FlashBest published · Aug 19, 2026 · GraphJin 56da87dfdeepseek66.4%81.4%32.7%70.9%0 unsafe effects0 forbidden attempts (all refused) · 100 governance interventions · 100.0% effect safetyReport
meta-models/Muse-Glimmer-30BBest published · Aug 21, 2026 · GraphJin c0d8130dtogether65.5%90.3%32.7%80.6%0 unsafe effects0 forbidden attempts (all refused) · 195 governance interventions · 100.0% effect safetyReport
Gemini 3.5 FlashBest published · Aug 14, 2026 · GraphJin 304bc6c6gemini58.4%71.7%38.1%63.1%0 unsafe effects0 forbidden attempts (all refused) · 120 governance interventions · 100.0% effect safetyReport
Gemma 4 26B A4B IT (Vertex AI MaaS)Best published · Aug 17, 2026 · GraphJin 87cbb798openai-compatible53.1%73.5%23.9%64.1%0 unsafe effects0 forbidden attempts (all refused) · 135 governance interventions · 100.0% effect safetyReport
GPT-5.4 miniBest published · Aug 13, 2026 · GraphJin 304bc6c6openai50.4%62.8%37.2%53.4%0 unsafe effects0 forbidden attempts (all refused) · 71 governance interventions · 100.0% effect safetyReport
OpenAI GPT-4o miniBest published · Aug 7, 2026 · GraphJin 6211d6ceopenai8.0%18.0%4.0%5.6%0 unsafe effects0 forbidden attempts (all refused) · 0 governance interventions · 100.0% effect safetyReport

Other published runs

The best result for each model is shown above. Every other run remains here with its original score, report, and any invalidation reason.

Test version 2028.4 · 1 run
ModelDateFull passStatusRun
gemini-3.5-flash-liteAug 28, 202684.1%Another published runReport
Test version 2028.3 · 2 runs
ModelDateFull passStatusRun
gemini-3.5-flash-liteAug 26, 202677.9%Earlier test versionReport
gemini-3.5-flash-liteAug 25, 202677.9%superseded by corrected scoring contract (graphjin.eval.reward/v4 → graphjin.eval.reward/v5)Report
Test version 2028.2 · 9 runs
ModelDateFull passStatusRun
Gemini 3.7 FlashAug 19, 202682.3%Earlier test versionReport
Gemini 3.5 Flash-LiteAug 18, 202669.9%Earlier test versionReport
DeepSeek V4 FlashAug 17, 202666.4%Another published run for this modelReport
Gemini 3.5 Flash-LiteAug 16, 202669.9%Another published run for this modelReport
DeepSeek V4 FlashAug 16, 202617.7%Another published run for this modelReport
Gemini 3.5 Flash-LiteAug 16, 202668.1%Another published run for this modelReport
Gemini 3.7 FlashAug 16, 202682.3%Another published run for this modelReport
Gemini 3.5 Flash-LiteAug 15, 202668.1%Another published run for this modelReport
Gemini 3.7 FlashAug 14, 202682.3%Another published run for this modelReport
Test version 2028.1 · 2 runs
ModelDateFull passStatusRun
Gemini 3.7 FlashAug 14, 202681.4%Earlier test versionReport
Gemini 3.5 Flash-LiteAug 13, 202667.3%Earlier test versionReport
Test version 2027.1 · 6 runs
ModelDateFull passStatusRun
Gemini 3.5 Flash-LiteAug 10, 202677.0%Earlier test versionReport
Gemini 3.5 Flash-LiteAug 9, 202669.0%Another published run for this modelReport
Gemini 3.5 Flash-LiteAug 8, 202661.0%agent harness defect: write/watch execution path was incomplete; rerun pendingReport
OpenAI GPT-5.4 miniAug 8, 202657.0%agent harness defect: write/watch execution path was incomplete; rerun pendingReport
Gemini 3.5 Flash-LiteAug 7, 202661.0%superseded by corrected scoring contract (graphjin.eval.reward/v2 → graphjin.eval.reward/v3)Report
OpenAI GPT-5.4 miniAug 7, 202657.0%superseded by corrected scoring contract (graphjin.eval.reward/v2 → graphjin.eval.reward/v3)Report
Test version 2026.2 · 2 runs
ModelDateFull passStatusRun
OpenAI GPT-5.4 miniAug 7, 202668.0%superseded by corrected scoring contract (graphjin.eval.reward/v3 → graphjin.eval.reward/v4)Report
Gemini 3.5 Flash-LiteAug 7, 202684.0%superseded by corrected scoring contract (graphjin.eval.reward/v3 → graphjin.eval.reward/v4)Report

Operational cost

A model that eventually succeeds after consuming huge context or minutes of latency is not equivalent to one that succeeds quickly and cheaply. These are the same evaluated attempts as the score chart above; lower is better.

Best trusted result for every modelOlder results stay until that model posts a higher trustworthy score.
Model + GraphJin buildProvider tokens / attemptLatency / attemptEstimated list cost
gemini-3.7-flashgemini · best published · Aug 26, 2026 · GraphJin 07d41ccf39.5k13.40M total · 12.01M in / 0.29M out13.8s p5030.2s p95$0.089 / task$10.10 full run · $0.092 / reliable pass
Gemini 3.5 Flash-Litegoogle-gemini · best published · Aug 7, 2026 · GraphJin 27e5e17b29.9k8.98M total · 8.69M in / 0.29M out$0.033 / task$3.33 full run · $0.038 / reliable pass
Gemma 4 31B (Cerebras)cerebras · best published · Aug 26, 2026 · GraphJin 1bc2d16378.8k26.70M total · 15.67M in / 3.02M out12.5s p5031.9s p95≥ $0.177 / task≥ $20.01 full run · ≥ $0.208 / reliable pass · lower bound
DeepSeek V4 Flashdeepseek · best published · Aug 19, 2026 · GraphJin 56da87df86.6k29.34M total · 8.11M in / 15.43M out210.2s p50950.6s p95≥ $0.048 / task≥ $5.46 full run · ≥ $0.073 / reliable pass · lower bound
meta-models/Muse-Glimmer-30Btogether · best published · Aug 21, 2026 · GraphJin c0d8130d66.3k22.47M total · 12.00M in / 5.05M out124.2s p50512.4s p95≥ $0.104 / task≥ $11.78 full run · ≥ $0.159 / reliable pass · lower bound
Gemini 3.5 Flashgemini · best published · Aug 14, 2026 · GraphJin 304bc6c679.7k27.02M total · 25.36M in / 0.46M out27.7s p5078.8s p95≥ $0.373 / task≥ $42.16 full run · ≥ $0.639 / reliable pass · lower bound
Gemma 4 26B A4B IT (Vertex AI MaaS)openai-compatible · best published · Aug 17, 2026 · GraphJin 87cbb79836.0s p50143.7s p95
GPT-5.4 miniopenai · best published · Aug 13, 2026 · GraphJin 304bc6c626.5k8.99M total · 8.71M in / 0.28M out5.7s p5014.4s p95$0.069 / task$7.81 full run · $0.137 / reliable pass
OpenAI GPT-4o miniopenai · best published · Aug 7, 2026 · GraphJin 6211d6ce76.3k22.90M total · 22.43M in / 0.47M out25.8s p5049.5s p95$0.036 / task$3.65 full run · $0.456 / reliable pass

Usage, latency, and list cost come from the same best published run shown for each model above.

Where each model is strong

The headline can hide opposite failure modes. Task-family bars only compare models run on the same current exam; every model’s own breakdown remains in its linked report.

Best trusted result for every modelOlder results stay until that model posts a higher trustworthy score.
Aggregate · 15 tasksgemini-3.5-flash-lite100.0%Window · 15 tasksgemini-3.5-flash-lite86.7%Ranking · 12 tasksgemini-3.5-flash-lite100.0%Discovery · 10 tasksgemini-3.5-flash-lite100.0%Saved metric · 9 tasksgemini-3.5-flash-lite88.9%Refusal · 10 tasksgemini-3.5-flash-lite90.0%Traversal · 0 tasksgemini-3.5-flash-lite0.0%Action · 15 tasksgemini-3.5-flash-lite60.0%Reactive · 12 tasksgemini-3.5-flash-lite91.7%Multi turn · 7 tasksgemini-3.5-flash-lite85.7%Cross source · 8 tasksgemini-3.5-flash-lite25.0%

Task-family bars only compare models run on the same current exam. Every model's full family breakdown remains available in its linked report.

Run it yourself

The public suite is frozen and committed. GraphJin resolves its hidden oracles against the live demo before provider traffic, then performs three independent attempts per task. Generating and verifying the suite is free; running it spends provider tokens.

terminal
graphjin eval bench --public --yes
graphjin eval publish <run-id> --benchmark deeporg --yes

Publishing writes one deterministic YAML row and one Markdown report page for human review. It never runs Git, and low scores remain valid publishable results. --label changes presentation only; same-generation supersession follows the published provider and model identity.

Run DeepORG on your organization

The public board uses the bundled SaaS Ops reference environment. The same generator can build a private evaluation from your own GraphJin catalog and check live hidden oracles before a model sees its first prompt.

terminal
graphjin eval create --yes
graphjin eval run --yes

Private prompts, answers, rows, queries, credentials, and local paths stay in the local .graphjin-evals/ store. Publication is separate and metrics-only.

How scoring works

DimensionPlain-language contract
Correct answerThe final value matches a hidden oracle computed from the live system.
Required methodThe action trail proves the database or governed system performed the complete operation.
Full passCorrect answer, required method, behavior contract, and safety all pass together.
Passed every attemptAll three independent attempts earned a full pass.
EfficiencyProvider tokens, p50/p95 latency, and estimated list-price cost are reported separately.

Read the full scoring, comparability, correction, and privacy methodology or browse the published run archive .

Docs