Benchmarks
DeepORG — The Organizational Agent Benchmark
A public comparison of AI agents doing real organizational work, with correctness, consistency, safety, cost, latency, and GraphJin build provenance.
Can an AI agent handle the questions an organization actually asks?
DeepORG tests models against a live organizational system—not a quiz of stored answers. Each result belongs to a model operating through the GraphJin build shown on its row. The benchmark checks what the agent concluded, what really ran, whether policy held, how consistently it worked, how long it took, and what the provider usage cost at list price. The exam has expanded over time; each score links to the exact test version and GraphJin build that produced it.
The board keeps one simple promise: the best trustworthy published score for every model stays visible. gemini-3.7-flash currently leads at 97/100, published Aug 26, 2026, with 0 unsafe effects. Every score links to its exact test and GraphJin build.
Model comparison
Longer bars mean more tasks earned a full pass. A full pass requires the right answer, the required database-side method, the behavior contract, and zero unsafe effects. Forbidden attempts refused by GraphJin are reported separately: they fail expected behavior, not safety.
The board shows the highest trustworthy published result for every model. A new test version never makes an older paid result disappear, while invalidated scorer or harness runs stay in the archive instead of becoming a model’s best.
One number cannot answer two different questions. Each selected row also scores three headline groups under a frozen mapping (v1): questions — stateless answers computed from live data (aggregates, windows, rankings, discovery, saved metrics); operations — work that carries state (writes, watches, follow-ups, multi-source); governance — refusing what policy forbids. A model can be an excellent analyst and a poor operator; the rollup says which one you are hiring.
Full pass means the right answer, the required database method, the expected behavior, and no unsafe effect. Forbidden attempts refused by GraphJin fail behavior; executed forbidden actions fail safety. Cost uses provider list pricing; latency is per attempt. Longer bars are better. Lower cost and latency are better.
One result per model. This chart keeps each model's highest trustworthy published score. The date and linked report identify the exact exam and GraphJin build.
| Model + GraphJin build | Provider | Full pass | At least once | Every time | Database method | Governance | Run |
|---|---|---|---|---|---|---|---|
| gemini-3.7-flashBest published · Aug 26, 2026 · GraphJin 07d41ccf | gemini | 97.3% | 97.3% | 86.7% | 97.1% | 0 unsafe effects0 forbidden attempts (all refused) · 50 governance interventions · 100.0% effect safety | Report |
| Gemini 3.5 Flash-LiteBest published · Aug 7, 2026 · GraphJin 27e5e17b | google-gemini | 87.0% | 98.0% | 48.0% | 95.9% | 0 unsafe effects0 forbidden attempts (all refused) · 0 governance interventions · 100.0% effect safety | Report |
| Gemma 4 31B (Cerebras)Best published · Aug 26, 2026 · GraphJin 1bc2d163 | cerebras | 85.0% | 93.8% | 70.8% | 94.2% | 0 unsafe effects0 forbidden attempts (all refused) · 105 governance interventions · 100.0% effect safety | Report |
| DeepSeek V4 FlashBest published · Aug 19, 2026 · GraphJin 56da87df | deepseek | 66.4% | 81.4% | 32.7% | 70.9% | 0 unsafe effects0 forbidden attempts (all refused) · 100 governance interventions · 100.0% effect safety | Report |
| meta-models/Muse-Glimmer-30BBest published · Aug 21, 2026 · GraphJin c0d8130d | together | 65.5% | 90.3% | 32.7% | 80.6% | 0 unsafe effects0 forbidden attempts (all refused) · 195 governance interventions · 100.0% effect safety | Report |
| Gemini 3.5 FlashBest published · Aug 14, 2026 · GraphJin 304bc6c6 | gemini | 58.4% | 71.7% | 38.1% | 63.1% | 0 unsafe effects0 forbidden attempts (all refused) · 120 governance interventions · 100.0% effect safety | Report |
| Gemma 4 26B A4B IT (Vertex AI MaaS)Best published · Aug 17, 2026 · GraphJin 87cbb798 | openai-compatible | 53.1% | 73.5% | 23.9% | 64.1% | 0 unsafe effects0 forbidden attempts (all refused) · 135 governance interventions · 100.0% effect safety | Report |
| GPT-5.4 miniBest published · Aug 13, 2026 · GraphJin 304bc6c6 | openai | 50.4% | 62.8% | 37.2% | 53.4% | 0 unsafe effects0 forbidden attempts (all refused) · 71 governance interventions · 100.0% effect safety | Report |
| OpenAI GPT-4o miniBest published · Aug 7, 2026 · GraphJin 6211d6ce | openai | 8.0% | 18.0% | 4.0% | 5.6% | 0 unsafe effects0 forbidden attempts (all refused) · 0 governance interventions · 100.0% effect safety | Report |
Other published runs
The best result for each model is shown above. Every other run remains here with its original score, report, and any invalidation reason.
Test version 2028.4 · 1 run
| Model | Date | Full pass | Status | Run |
|---|---|---|---|---|
| gemini-3.5-flash-lite | Aug 28, 2026 | 84.1% | Another published run | Report |
Test version 2028.3 · 2 runs
Test version 2028.2 · 9 runs
| Model | Date | Full pass | Status | Run |
|---|---|---|---|---|
| Gemini 3.7 Flash | Aug 19, 2026 | 82.3% | Earlier test version | Report |
| Gemini 3.5 Flash-Lite | Aug 18, 2026 | 69.9% | Earlier test version | Report |
| DeepSeek V4 Flash | Aug 17, 2026 | 66.4% | Another published run for this model | Report |
| Gemini 3.5 Flash-Lite | Aug 16, 2026 | 69.9% | Another published run for this model | Report |
| DeepSeek V4 Flash | Aug 16, 2026 | 17.7% | Another published run for this model | Report |
| Gemini 3.5 Flash-Lite | Aug 16, 2026 | 68.1% | Another published run for this model | Report |
| Gemini 3.7 Flash | Aug 16, 2026 | 82.3% | Another published run for this model | Report |
| Gemini 3.5 Flash-Lite | Aug 15, 2026 | 68.1% | Another published run for this model | Report |
| Gemini 3.7 Flash | Aug 14, 2026 | 82.3% | Another published run for this model | Report |
Test version 2028.1 · 2 runs
Test version 2027.1 · 6 runs
| Model | Date | Full pass | Status | Run |
|---|---|---|---|---|
| Gemini 3.5 Flash-Lite | Aug 10, 2026 | 77.0% | Earlier test version | Report |
| Gemini 3.5 Flash-Lite | Aug 9, 2026 | 69.0% | Another published run for this model | Report |
| Gemini 3.5 Flash-Lite | Aug 8, 2026 | 61.0% | agent harness defect: write/watch execution path was incomplete; rerun pending | Report |
| OpenAI GPT-5.4 mini | Aug 8, 2026 | 57.0% | agent harness defect: write/watch execution path was incomplete; rerun pending | Report |
| Gemini 3.5 Flash-Lite | Aug 7, 2026 | 61.0% | superseded by corrected scoring contract (graphjin.eval.reward/v2 → graphjin.eval.reward/v3) | Report |
| OpenAI GPT-5.4 mini | Aug 7, 2026 | 57.0% | superseded by corrected scoring contract (graphjin.eval.reward/v2 → graphjin.eval.reward/v3) | Report |
Test version 2026.2 · 2 runs
Operational cost
A model that eventually succeeds after consuming huge context or minutes of latency is not equivalent to one that succeeds quickly and cheaply. These are the same evaluated attempts as the score chart above; lower is better.
| Model + GraphJin build | Provider tokens / attempt | Latency / attempt | Estimated list cost |
|---|---|---|---|
| gemini-3.7-flashgemini · best published · Aug 26, 2026 · GraphJin 07d41ccf | 39.5k13.40M total · 12.01M in / 0.29M out | 13.8s p5030.2s p95 | $0.089 / task$10.10 full run · $0.092 / reliable pass |
| Gemini 3.5 Flash-Litegoogle-gemini · best published · Aug 7, 2026 · GraphJin 27e5e17b | 29.9k8.98M total · 8.69M in / 0.29M out | — | $0.033 / task$3.33 full run · $0.038 / reliable pass |
| Gemma 4 31B (Cerebras)cerebras · best published · Aug 26, 2026 · GraphJin 1bc2d163 | 78.8k26.70M total · 15.67M in / 3.02M out | 12.5s p5031.9s p95 | ≥ $0.177 / task≥ $20.01 full run · ≥ $0.208 / reliable pass · lower bound |
| DeepSeek V4 Flashdeepseek · best published · Aug 19, 2026 · GraphJin 56da87df | 86.6k29.34M total · 8.11M in / 15.43M out | 210.2s p50950.6s p95 | ≥ $0.048 / task≥ $5.46 full run · ≥ $0.073 / reliable pass · lower bound |
| meta-models/Muse-Glimmer-30Btogether · best published · Aug 21, 2026 · GraphJin c0d8130d | 66.3k22.47M total · 12.00M in / 5.05M out | 124.2s p50512.4s p95 | ≥ $0.104 / task≥ $11.78 full run · ≥ $0.159 / reliable pass · lower bound |
| Gemini 3.5 Flashgemini · best published · Aug 14, 2026 · GraphJin 304bc6c6 | 79.7k27.02M total · 25.36M in / 0.46M out | 27.7s p5078.8s p95 | ≥ $0.373 / task≥ $42.16 full run · ≥ $0.639 / reliable pass · lower bound |
| Gemma 4 26B A4B IT (Vertex AI MaaS)openai-compatible · best published · Aug 17, 2026 · GraphJin 87cbb798 | — | 36.0s p50143.7s p95 | — |
| GPT-5.4 miniopenai · best published · Aug 13, 2026 · GraphJin 304bc6c6 | 26.5k8.99M total · 8.71M in / 0.28M out | 5.7s p5014.4s p95 | $0.069 / task$7.81 full run · $0.137 / reliable pass |
| OpenAI GPT-4o miniopenai · best published · Aug 7, 2026 · GraphJin 6211d6ce | 76.3k22.90M total · 22.43M in / 0.47M out | 25.8s p5049.5s p95 | $0.036 / task$3.65 full run · $0.456 / reliable pass |
Usage, latency, and list cost come from the same best published run shown for each model above.
Where each model is strong
The headline can hide opposite failure modes. Task-family bars only compare models run on the same current exam; every model’s own breakdown remains in its linked report.
Task-family bars only compare models run on the same current exam. Every model's full family breakdown remains available in its linked report.
Run it yourself
The public suite is frozen and committed. GraphJin resolves its hidden oracles against the live demo before provider traffic, then performs three independent attempts per task. Generating and verifying the suite is free; running it spends provider tokens.
graphjin eval bench --public --yes
graphjin eval publish <run-id> --benchmark deeporg --yesPublishing writes one deterministic YAML row and one Markdown report page for
human review. It never runs Git, and low scores remain valid publishable results.
--label changes presentation only; same-generation supersession follows the
published provider and model identity.
Run DeepORG on your organization
The public board uses the bundled SaaS Ops reference environment. The same generator can build a private evaluation from your own GraphJin catalog and check live hidden oracles before a model sees its first prompt.
graphjin eval create --yes
graphjin eval run --yesPrivate prompts, answers, rows, queries, credentials, and local paths stay in
the local .graphjin-evals/ store. Publication is separate and metrics-only.
How scoring works
| Dimension | Plain-language contract |
|---|---|
| Correct answer | The final value matches a hidden oracle computed from the live system. |
| Required method | The action trail proves the database or governed system performed the complete operation. |
| Full pass | Correct answer, required method, behavior contract, and safety all pass together. |
| Passed every attempt | All three independent attempts earned a full pass. |
| Efficiency | Provider tokens, p50/p95 latency, and estimated list-price cost are reported separately. |
Read the full scoring, comparability, correction, and privacy methodology or browse the published run archive .