Benchmarks
Benchmarks
Public, governed benchmark suites generated from reference environments and verified against live runtime truth.
GraphJin benchmarks are frozen, oracle-verified suites generated from reference environments. Models run through the same governed GraphJin interface, while publication gates preserve the suite fingerprint, reject stale scoring contracts, flag suspicious answer/method divergence, and record the exact GraphJin build used for every result.
The family grows as real reference environments are added. We list only benchmarks that exist and have a published, reproducible suite.
| Benchmark | Status | Models on board | Last updated |
|---|---|---|---|
| DeepORG — The Organizational Agent Benchmark | Live | Gemini 3.5 Flash-Lite, OpenAI GPT-5.4 mini | 2026-08-08 |