DeepORG — The Organizational Agent Benchmark One AI agent across every database you run — governed, safe, and graded in public. DeepORG is the benchmark that proves it: scores, safety, cost, and the meaning behind every full pass. benchmarks benchmark-overview benchmark benchmark/_index.md

Benchmarks · by GraphJin

Last updated: 2026-08-19

DeepORG

The Organizational Agent Benchmark

We grade our own AI agent in public.

GraphJin connects one AI agent to every database, API, and file system you run — governed, so it can only do what policy allows. DeepORG is the public exam it has to pass: 113 real organizational tasks — questions to answer, operations to carry out, and requests it should refuse — 3 attempts each, scored against live data.

Every ranked model on this board is small, fast, and cheap to run. The frontier models have not been run yet.

82/100

Full pass score

Gemini 3.7 Flash achieved 82.3 out of 100 full passes across 113 live tasks with 3 attempts each.

Estimated range: 75–88What this means

0

unsafe effects

Zero unsafe effects across 339 attempts. No unintended writes, updates, or side effects.

Latency (p50)10.2sper attempt
Reliable pass$0.114at list price
Scope113live tasks
Attempts3attempts each
Generation2028.2tested on

What GraphJin is

Most AI agents guess. Ours can't.

Point a language model at raw credentials and it invents what it cannot see — tables, joins, numbers. It sounds confident. It is guessing.

GraphJin gives the agent a governed map of everything you run. The model plans the question; the GraphJin engine does the work inside the database. The agent never writes its own SQL against production and never sees your credentials — so a wrong guess cannot become a wrong action.

AI agentworks from memoryno map · no guardrails????DatabaseRemote APIsFilesSource code

One agent across the stack you already run

PostgreSQL MySQL MariaDB MongoDB SQLite SQL Server Oracle CockroachDB YugabyteDB Snowflake Redshift BigQuery Cassandra / Keyspaces AWS Aurora Cloud SQL HTTP APIs S3 / GCS / Files Code

The demo

Watch it answer a real question.

This is GraphJin's built-in SaaS Ops demo — the same environment DeepORG measures. One command runs it on your laptop.

GraphJin Agent · SaaS Ops demo
Which account is most at risk of churn—and why?
query_catalog
3 relevant sources discovered
Postgres accounts + invoicesSnowflake product usageCRM API opportunities
validate_where_clause
relationships + churn filters valid
execute_graphql
3 optimized, policy-checked operations

Evidence checked · 3 systems · 3 operations

Meridian Robotics — renewal in 9 days, usage down 38%, two failed payments, and an unresolved escalation.

graphjin serve --demoTry it on your machine

Illustrative conversation from the bundled SaaS Ops demo—not a published test item.

From question to verified answer

  1. 1
    Ask

    A person asks in plain language.

  2. 2
    Execute

    The agent picks a governed method; the engine runs it in the database.

  3. 3
    Verify

    DeepORG checks the answer, the method, the behavior, and safety.

  4. 4
    Result

    Only all four together count as a pass.

The score, in plain English

What 82 actually means

The agent fully succeeded on 93 of 113 real-work tasks. Each task was tried 3 times, and a pass had to meet all four checks.

  1. 1
    Correct answer

    It got the answer right.

  2. 2
    Right way of getting it

    The governed system did the real work; the model did not guess.

  3. 3
    Expected behavior

    It followed the rules for that task.

  4. 4
    No unsafe effects

    It made no unsafe or unauthorized change.

The DeepORG Benchmark

DeepORG: Can an AI agent actually do what your organization needs?

DeepORG runs one frozen exam of 113 tasks against a live organizational database — questions to answer, operations to carry out, and requests it should refuse. A model only scores a full pass when it clears the complete four-part contract: the right answer, by the right method, with the expected behavior, and no unsafe writes. Every model runs the identical suite and every result is published — including the ones that fail.

Solid columns are the current generation 2028.2 cohort; muted columns carry each model's latest generation 2028.1 result forward until it is rerun. Every column links to its full technical report.

DeepORG benchmark results by model
What is measuredGemini 3.7 FlashGemini · GraphJin 2e9e4a3dGemini 3.5 Flash-LiteGemini · GraphJin e6ea37dfDeepSeek V4 FlashDeepseek · GraphJin 56da87dfmeta-models/Muse-Glimmer-30BTogether · GraphJin c0d8130dGemma 4 26B A4B IT (Vertex AI MaaS)Openai compatible · GraphJin 87cbb798Gemini 3.5 FlashGemini · GraphJin 304bc6c62028.1 cohortGPT-5.4 miniOpenai · GraphJin 304bc6c62028.1 cohort
Full passesPassed the complete four-part contract82.3%69.9%66.4%65.5%53.1%58.4%50.4%
Passed every attemptAll 3 tries earned a full pass70.8%37.2%32.7%32.7%23.9%38.1%37.2%
Correct answerMatched live, hidden ground truth83.5%70.9%74.8%65.0%54.4%66.0%49.5%
Right methodThe governed system did the real work94.2%79.6%70.9%80.6%64.1%63.1%53.4%
Expected behaviorFollowed the rules for the task96.5%80.5%81.4%74.3%64.6%66.4%61.9%
Unsafe effectsLower is better0000000
Median responseLower is better · per attempt10.2s3.9s210.2s124.2s36.0s27.7s5.7s
Cost per reliable passLower is better · provider list price$0.114$0.056$0.073$0.159$0.639$0.137

Prior-cohort columns were scored under generation 2028.1 rules and rerank when rerun. Usage, latency, and cost are measured per attempt and compare directly.

Scroll to compare models

113 real-work tasks

What DeepORG actually tests

DeepORG asks for everyday organizational work — the questions a real company asks in a normal week — grouped here into six things anyone can recognize.

Find the right information

Locate the useful data without loading the whole company into the prompt.

19 tasks · discovery + saved metrics

Work out the answer

Calculate totals, rankings, and time windows in the database.

42 tasks · aggregates + rankings + windows

Follow the conversation

Use context from an earlier question in the next one.

7 tasks · multi-turn

Connect different systems

Combine databases, APIs, files, and code when the answer spans them.

8 tasks · cross-source

Act only when allowed

Make approved changes and refuse work that policy forbids.

25 tasks · actions + refusals

Notice what changed

Watch for new events, missing events, and conditions needing attention.

12 tasks · reactive

On the leaderboard these score under three headline groups (frozen mapping v1): questions — finding and working out answers; operations — conversations, connections, changes, and watches; governance — refusing what policy forbids.

The safety story

The agent never holds the keys.

Every operation is compiled, policy-checked, and logged by GraphJin before it runs. That is how 0 unsafe effects across 339 attempts is possible — and why part of the exam is refusing work that policy forbids.

Policy decides, not the prompt.

Access rules live in GraphJin's configuration, outside the model. No clever wording can talk the agent past them.

Every answer carries evidence.

Each result links back to the exact governed operations that produced it, so you can check the work instead of trusting the prose.

Refusing is part of the exam.

DeepORG includes questions the agent must decline. Doing the forbidden thing fails the test — politely refusing passes it.

Run DeepORG yourself

Same exam, your laptop. Test the frozen public benchmark with your own model and GraphJin build.

graphjin eval bench --public --yesRun it yourself