Cybatar separates independently researched external benchmark observations from benchmark runs executed and verified by Cybatar. Configuration, harness, benchmark version and provenance are preserved because scores are not comparable when evaluation conditions differ.
102 sourced observations across 11 catalogue models · research snapshot 22 Aug 2026.
Enterprise workflow automation benchmark.
| Model | Score | Configuration / harness | Provenance | Source |
|---|---|---|---|---|
| Gemini 3.7 Flash | 30.4 % | Model evaluation · thinking: high | Provider Reported | Google DeepMind 13 Aug 2026 |
| GPT-5.6 Sol OpenAI | 18.1 % | Model evaluation · provider reported configuration: OpenAI launch evaluation | Provider Reported | OpenAI 09 Jul 2026 |
| GPT-5.6 Terra OpenAI | 15.2 % | Model evaluation · provider reported configuration: OpenAI launch evaluation | Provider Reported | OpenAI 09 Jul 2026 |
| GPT-5.6 Luna OpenAI | 14.9 % | Model evaluation · provider reported configuration: OpenAI launch evaluation | Provider Reported | OpenAI 09 Jul 2026 |
Multi-step workflows using Model Context Protocol tools.
| Model | Score | Configuration / harness | Provenance | Source |
|---|---|---|---|---|
| Gemini 3.1 Pro Preview | 69.2 % | Model evaluation · thinking: high | Provider Reported | Google DeepMind 19 Feb 2026 |
Agentic web-search benchmark.
| Model | Score | Configuration / harness | Provenance | Source |
|---|---|---|---|---|
| Gemini 3.1 Pro Preview | 85.9 % | Model evaluation · search: yes · python: yes · browse: yes | Provider Reported | Google DeepMind 19 Feb 2026 |
Coding-agent benchmark where model and agent harness are both material to the result.
| Model | Score | Configuration / harness | Provenance | Source |
|---|---|---|---|---|
| Claude Opus 5 Anthropic | 67 | Claude Code · reasoning effort: xhigh | Independent | Artificial Analysis 22 Aug 2026 |
| GPT-5.6 Sol OpenAI | 67 | Codex · reasoning effort: max | Independent | Artificial Analysis 22 Aug 2026 |
| Grok 4.5 xAI | 64 | Grok Build · reasoning effort: high | Independent | Artificial Analysis 22 Aug 2026 |
| Gemini 3.7 Flash | 57 | OpenCode · thinking: high | Independent | Artificial Analysis 22 Aug 2026 |
| Gemini 3.1 Pro Preview | 30 | Gemini CLI · thinking: high | Independent | Artificial Analysis 22 Aug 2026 |
Web-development coding arena Elo.
| Model | Score | Configuration / harness | Provenance | Source |
|---|---|---|---|---|
| Gemini 3.7 Flash | 1588 Elo | Model evaluation · | Provider Reported | Google DeepMind 13 Aug 2026 |
Long-horizon software engineering benchmark.
| Model | Score | Configuration / harness | Provenance | Source |
|---|---|---|---|---|
| GPT-5.6 Sol OpenAI | 72.7 % | Model evaluation · provider reported configuration: OpenAI launch evaluation | Provider Reported | OpenAI 09 Jul 2026 |
| GPT-5.6 Terra OpenAI | 69.6 % | Model evaluation · provider reported configuration: OpenAI launch evaluation | Provider Reported | OpenAI 09 Jul 2026 |
| GPT-5.6 Sol OpenAI | 69 % | Codex · reasoning effort: max | Independent | Artificial Analysis 22 Aug 2026 |
| GPT-5.6 Luna OpenAI | 67.2 % | Model evaluation · provider reported configuration: OpenAI launch evaluation | Provider Reported | OpenAI 09 Jul 2026 |
| Gemini 3.7 Flash | 65.3 % | Model evaluation · thinking: high | Provider Reported | Google DeepMind 13 Aug 2026 |
| Claude Opus 5 Anthropic | 60 % | Claude Code · reasoning effort: xhigh | Independent | Artificial Analysis 22 Aug 2026 |
| Grok 4.5 xAI | 60 % | Grok Build · reasoning effort: high | Independent | Artificial Analysis 22 Aug 2026 |
| Gemini 3.7 Flash | 57 % | OpenCode · thinking: high | Independent | Artificial Analysis 22 Aug 2026 |
| Gemini 3.1 Pro Preview | 14 % | Gemini CLI · thinking: high | Independent | Artificial Analysis 22 Aug 2026 |
Production code quality benchmark.
| Model | Score | Configuration / harness | Provenance | Source |
|---|---|---|---|---|
| Gemini 3.7 Flash | 43.6 % | Model evaluation · thinking: high | Provider Reported | Google DeepMind 13 Aug 2026 |
Scientific research coding benchmark.
| Model | Score | Configuration / harness | Provenance | Source |
|---|---|---|---|---|
| Gemini 3.1 Pro Preview | 59 % | Model evaluation · thinking: high | Provider Reported | Google DeepMind 19 Feb 2026 |
| DeepSeek V4 Flash DeepSeek | 50 % | Model evaluation · reasoning effort: max · release: 0731 | Independent | Artificial Analysis 31 Jul 2026 |
Repository understanding benchmark used in the Coding Agent Index.
| Model | Score | Configuration / harness | Provenance | Source |
|---|---|---|---|---|
| Claude Opus 5 Anthropic | 55 % | Claude Code · reasoning effort: xhigh | Independent | Artificial Analysis 22 Aug 2026 |
| Grok 4.5 xAI | 48 % | Grok Build · reasoning effort: high | Independent | Artificial Analysis 22 Aug 2026 |
| GPT-5.6 Sol OpenAI | 43 % | Codex · reasoning effort: max | Independent | Artificial Analysis 22 Aug 2026 |
| Gemini 3.7 Flash | 31 % | OpenCode · thinking: high | Independent | Artificial Analysis 22 Aug 2026 |
| Gemini 3.1 Pro Preview | 9 % | Gemini CLI · thinking: high | Independent | Artificial Analysis 22 Aug 2026 |
More diverse agentic software-engineering tasks.
| Model | Score | Configuration / harness | Provenance | Source |
|---|---|---|---|---|
| Gemini 3.1 Pro Preview | 54.2 % | Model evaluation · single attempt: yes · thinking: high | Provider Reported | Google DeepMind 19 Feb 2026 |
Agentic coding benchmark on real GitHub issues.
| Model | Score | Configuration / harness | Provenance | Source |
|---|---|---|---|---|
| Gemini 3.1 Pro Preview | 80.6 % | Model evaluation · single attempt: yes · thinking: high | Provider Reported | Google DeepMind 19 Feb 2026 |
Agentic terminal-use benchmark.
| Model | Score | Configuration / harness | Provenance | Source |
|---|---|---|---|---|
| GPT-5.6 Sol OpenAI | 88.8 % | Model evaluation · provider reported configuration: OpenAI launch evaluation | Provider Reported | OpenAI 09 Jul 2026 |
| GPT-5.6 Sol OpenAI | 88 % | Codex · reasoning effort: max | Independent | Artificial Analysis 22 Aug 2026 |
| GPT-5.6 Terra OpenAI | 87.4 % | Model evaluation · provider reported configuration: OpenAI launch evaluation | Provider Reported | OpenAI 09 Jul 2026 |
| Gemini 3.7 Flash | 85.8 % | Model evaluation · thinking: high | Provider Reported | Google DeepMind 13 Aug 2026 |
| Claude Opus 5 Anthropic | 85 % | Claude Code · reasoning effort: xhigh | Independent | Artificial Analysis 22 Aug 2026 |
| Grok 4.5 xAI | 85 % | Grok Build · reasoning effort: high | Independent | Artificial Analysis 22 Aug 2026 |
| GPT-5.6 Luna OpenAI | 84.7 % | Model evaluation · provider reported configuration: OpenAI launch evaluation | Provider Reported | OpenAI 09 Jul 2026 |
| Gemini 3.7 Flash | 83 % | OpenCode · thinking: high | Independent | Artificial Analysis 22 Aug 2026 |
| DeepSeek V4 Flash DeepSeek | 79 % | Model evaluation · reasoning effort: max · release: 0731 | Independent | Artificial Analysis 31 Jul 2026 |
| Gemini 3.1 Pro Preview | 68.5 % | Model evaluation · harness: Terminus-2 · thinking: high | Provider Reported | Google DeepMind 19 Feb 2026 |
| Gemini 3.1 Pro Preview | 68 % | Gemini CLI · thinking: high | Independent | Artificial Analysis 22 Aug 2026 |
Multimodal desktop and operating-system agent tasks.
| Model | Score | Configuration / harness | Provenance | Source |
|---|---|---|---|---|
| Gemini 3.7 Flash | 26.3 % | Model evaluation · thinking: high | Provider Reported | Google DeepMind 13 Aug 2026 |
Agentic computer-use benchmark.
| Model | Score | Configuration / harness | Provenance | Source |
|---|---|---|---|---|
| Gemini 3.7 Flash | 47.9 % | Model evaluation · thinking: high | Provider Reported | Google DeepMind 13 Aug 2026 |
Expert PDF and document comprehension benchmark.
| Model | Score | Configuration / harness | Provenance | Source |
|---|---|---|---|---|
| Gemini 3.7 Flash | 34 % | Model evaluation · thinking: high | Provider Reported | Google DeepMind 13 Aug 2026 |
| GPT-5.6 Sol OpenAI | 30.7 % | Model evaluation · provider reported configuration: OpenAI launch evaluation | Provider Reported | OpenAI 09 Jul 2026 |
| GPT-5.6 Terra OpenAI | 24.7 % | Model evaluation · provider reported configuration: OpenAI launch evaluation | Provider Reported | OpenAI 09 Jul 2026 |
| GPT-5.6 Luna OpenAI | 22.7 % | Model evaluation · provider reported configuration: OpenAI launch evaluation | Provider Reported | OpenAI 09 Jul 2026 |
Agentic banking tool-use benchmark.
| Model | Score | Configuration / harness | Provenance | Source |
|---|---|---|---|---|
| DeepSeek V4 Flash DeepSeek | 31 % | Model evaluation · reasoning effort: max · release: 0731 | Independent | Artificial Analysis 31 Jul 2026 |
Composite model intelligence across nine evaluations.
| Model | Score | Configuration / harness | Provenance | Source |
|---|---|---|---|---|
| Claude Opus 5 Anthropic | 63 | Model evaluation · reasoning effort: max | Independent | Artificial Analysis 22 Aug 2026 |
| GPT-5.6 Sol OpenAI | 61 | Model evaluation · reasoning effort: max | Independent | Artificial Analysis 22 Aug 2026 |
| Grok 4.6 xAI | 61 | Model evaluation · reasoning effort: high | Independent | Artificial Analysis 12 Aug 2026 |
| GPT-5.6 Terra OpenAI | 57 | Model evaluation · reasoning effort: max | Independent | Artificial Analysis 22 Aug 2026 |
| Gemini 3.7 Flash | 56 | Model evaluation · thinking: high | Independent | Artificial Analysis 22 Aug 2026 |
| DeepSeek V4 Pro DeepSeek | 53 | Model evaluation · reasoning effort: max · release: 0813 | Independent | Artificial Analysis 13 Aug 2026 |
| DeepSeek V4 Flash DeepSeek | 52 | Model evaluation · reasoning effort: max · release: 0731 | Independent | Artificial Analysis 31 Jul 2026 |
| GPT-5.6 Luna OpenAI | 52 | Model evaluation · reasoning effort: max | Independent | Artificial Analysis 22 Aug 2026 |
Professional healthcare benchmark.
| Model | Score | Configuration / harness | Provenance | Source |
|---|---|---|---|---|
| GPT-5.6 Sol OpenAI | 60.5 % | Model evaluation · provider reported configuration: OpenAI launch evaluation | Provider Reported | OpenAI 09 Jul 2026 |
| GPT-5.6 Terra OpenAI | 57.7 % | Model evaluation · provider reported configuration: OpenAI launch evaluation | Provider Reported | OpenAI 09 Jul 2026 |
| GPT-5.6 Luna OpenAI | 55.7 % | Model evaluation · provider reported configuration: OpenAI launch evaluation | Provider Reported | OpenAI 09 Jul 2026 |
Agentic knowledge-work benchmark focused on analytical quality and presentation.
| Model | Score | Configuration / harness | Provenance | Source |
|---|---|---|---|---|
| Claude Opus 5 Anthropic | 1720 Elo | Model evaluation · reasoning effort: max | Independent | Artificial Analysis 24 Jul 2026 |
| Grok 4.5 xAI | 1313 Elo | Model evaluation · reasoning effort: high | Independent | Artificial Analysis 22 Aug 2026 |
| Claude Sonnet 5 Anthropic | 1292 Elo | Model evaluation · reasoning effort: xhigh | Independent | Artificial Analysis 22 Aug 2026 |
| DeepSeek V4 Flash DeepSeek | 1285 Elo | Model evaluation · reasoning effort: max · release: 0731 | Independent | Artificial Analysis 22 Aug 2026 |
| Gemini 3.7 Flash | 1131 Elo | Model evaluation · thinking: high | Independent | Artificial Analysis 22 Aug 2026 |
Agentic real-world professional knowledge-work benchmark.
| Model | Score | Configuration / harness | Provenance | Source |
|---|---|---|---|---|
| Claude Opus 5 Anthropic | 1861 Elo | Model evaluation · reasoning effort: max | Independent | Artificial Analysis 24 Jul 2026 |
| Grok 4.6 xAI | 1753 Elo | Model evaluation · reasoning effort: high | Independent | Artificial Analysis 12 Aug 2026 |
| GPT-5.6 Sol OpenAI | 1679 Elo | Model evaluation · reasoning effort: xhigh | Independent | Artificial Analysis 22 Aug 2026 |
| Claude Sonnet 5 Anthropic | 1595 Elo | Model evaluation · reasoning effort: max | Independent | Artificial Analysis 22 Aug 2026 |
| DeepSeek V4 Pro DeepSeek | 1590 Elo | Model evaluation · reasoning effort: max · release: 0813 | Independent | Artificial Analysis 22 Aug 2026 |
| DeepSeek V4 Flash DeepSeek | 1559 Elo | Model evaluation · reasoning effort: max · release: 0731 | Independent | Artificial Analysis 31 Jul 2026 |
| Gemini 3.7 Flash | 1525 Elo | Model evaluation · thinking: high | Provider Reported | Google DeepMind 13 Aug 2026 |
| Gemini 3.1 Pro Preview | 1317 Elo | Model evaluation · thinking: high | Provider Reported | Google DeepMind 19 Feb 2026 |
Complex legal workflow evaluation.
| Model | Score | Configuration / harness | Provenance | Source |
|---|---|---|---|---|
| Gemini 3.7 Flash | 90.7 % | Model evaluation · thinking: high | Provider Reported | Google DeepMind 13 Aug 2026 |
Long-context multi-needle retrieval/reasoning at 128k context.
| Model | Score | Configuration / harness | Provenance | Source |
|---|---|---|---|---|
| Gemini 3.7 Flash | 97 % | Model evaluation · thinking: high | Provider Reported | Google DeepMind 13 Aug 2026 |
| Gemini 3.1 Pro Preview | 84.9 % | Model evaluation · context: 128k_average | Provider Reported | Google DeepMind 19 Feb 2026 |
Long-context multi-needle retrieval/reasoning at 1M context.
| Model | Score | Configuration / harness | Provenance | Source |
|---|---|---|---|---|
| Gemini 3.1 Pro Preview | 26.3 % | Model evaluation · context: 1m_pointwise | Provider Reported | Google DeepMind 19 Feb 2026 |
Multilingual question answering benchmark.
| Model | Score | Configuration / harness | Provenance | Source |
|---|---|---|---|---|
| Gemini 3.1 Pro Preview | 92.6 % | Model evaluation | Provider Reported | Google DeepMind 19 Feb 2026 |
Multimodal understanding and reasoning without tools.
| Model | Score | Configuration / harness | Provenance | Source |
|---|---|---|---|---|
| GPT-5.6 Sol OpenAI | 83 % | Model evaluation · provider reported configuration: OpenAI launch evaluation | Provider Reported | OpenAI 09 Jul 2026 |
| GPT-5.6 Terra OpenAI | 80.7 % | Model evaluation · provider reported configuration: OpenAI launch evaluation | Provider Reported | OpenAI 09 Jul 2026 |
| Gemini 3.1 Pro Preview | 80.5 % | Model evaluation · tools: none | Provider Reported | Google DeepMind 19 Feb 2026 |
| GPT-5.6 Luna OpenAI | 78.4 % | Model evaluation · provider reported configuration: OpenAI launch evaluation | Provider Reported | OpenAI 09 Jul 2026 |
Multimodal understanding and reasoning with tools.
| Model | Score | Configuration / harness | Provenance | Source |
|---|---|---|---|---|
| GPT-5.6 Sol OpenAI | 84.6 % | Model evaluation · provider reported configuration: OpenAI launch evaluation | Provider Reported | OpenAI 09 Jul 2026 |
| GPT-5.6 Terra OpenAI | 82 % | Model evaluation · provider reported configuration: OpenAI launch evaluation | Provider Reported | OpenAI 09 Jul 2026 |
| GPT-5.6 Luna OpenAI | 79.5 % | Model evaluation · provider reported configuration: OpenAI launch evaluation | Provider Reported | OpenAI 09 Jul 2026 |
Frontier mathematical reasoning, tiers 1–3.
| Model | Score | Configuration / harness | Provenance | Source |
|---|---|---|---|---|
| GPT-5.6 Sol OpenAI | 89 % | Model evaluation · provider reported configuration: OpenAI launch evaluation | Provider Reported | OpenAI 09 Jul 2026 |
| GPT-5.6 Terra OpenAI | 84.9 % | Model evaluation · provider reported configuration: OpenAI launch evaluation | Provider Reported | OpenAI 09 Jul 2026 |
| GPT-5.6 Luna OpenAI | 78.6 % | Model evaluation · provider reported configuration: OpenAI launch evaluation | Provider Reported | OpenAI 09 Jul 2026 |
Frontier mathematical reasoning, tier 4.
| Model | Score | Configuration / harness | Provenance | Source |
|---|---|---|---|---|
| GPT-5.6 Sol OpenAI | 83 % | Model evaluation · provider reported configuration: OpenAI launch evaluation | Provider Reported | OpenAI 09 Jul 2026 |
| GPT-5.6 Terra OpenAI | 68.3 % | Model evaluation · provider reported configuration: OpenAI launch evaluation | Provider Reported | OpenAI 09 Jul 2026 |
| GPT-5.6 Luna OpenAI | 58.5 % | Model evaluation · provider reported configuration: OpenAI launch evaluation | Provider Reported | OpenAI 09 Jul 2026 |
Abstract reasoning on novel logic tasks.
| Model | Score | Configuration / harness | Provenance | Source |
|---|---|---|---|---|
| Gemini 3.1 Pro Preview | 77.1 % | Model evaluation · ARC Prize verified: yes · thinking: high | Provider Reported | Google DeepMind 19 Feb 2026 |
Graduate-level scientific reasoning benchmark.
| Model | Score | Configuration / harness | Provenance | Source |
|---|---|---|---|---|
| GPT-5.6 Sol OpenAI | 94.6 % | Model evaluation · provider reported configuration: OpenAI launch evaluation | Provider Reported | OpenAI 09 Jul 2026 |
| Gemini 3.1 Pro Preview | 94.3 % | Model evaluation · tools: none · thinking: high | Provider Reported | Google DeepMind 19 Feb 2026 |
| GPT-5.6 Terra OpenAI | 92.9 % | Model evaluation · provider reported configuration: OpenAI launch evaluation | Provider Reported | OpenAI 09 Jul 2026 |
| GPT-5.6 Luna OpenAI | 92.3 % | Model evaluation · provider reported configuration: OpenAI launch evaluation | Provider Reported | OpenAI 09 Jul 2026 |
| DeepSeek V4 Flash DeepSeek | 91 % | Model evaluation · reasoning effort: max · release: 0731 | Independent | Artificial Analysis 31 Jul 2026 |
Verified multidisciplinary expert reasoning set reported by Google DeepMind.
| Model | Score | Configuration / harness | Provenance | Source |
|---|---|---|---|---|
| Gemini 3.7 Flash | 53.6 % | Model evaluation · thinking: high | Provider Reported | Google DeepMind 13 Aug 2026 |
Broad expert-level reasoning and knowledge benchmark.
| Model | Score | Configuration / harness | Provenance | Source |
|---|---|---|---|---|
| Gemini 3.1 Pro Preview | 44.4 % | Model evaluation · tools: none · thinking: high | Provider Reported | Google DeepMind 19 Feb 2026 |
| DeepSeek V4 Flash DeepSeek | 37 % | Model evaluation · reasoning effort: max · release: 0731 | Independent | Artificial Analysis 31 Jul 2026 |
Percentage of attempted answers classified as hallucinations.
| Model | Score | Configuration / harness | Provenance | Source |
|---|---|---|---|---|
| DeepSeek V4 Flash DeepSeek | 84 % | Model evaluation · reasoning effort: max · release: 0731 | Independent | Artificial Analysis 31 Jul 2026 |
Rewards correct answers and penalizes hallucinations; refusal is not penalized.
| Model | Score | Configuration / harness | Provenance | Source |
|---|---|---|---|---|
| DeepSeek V4 Flash DeepSeek | -16 | Model evaluation · reasoning effort: max · release: 0731 | Independent | Artificial Analysis 31 Jul 2026 |
Cybatar composite scores appear only when a model has complete published Cybatar-run evidence for every required component suite. External research observations above never substitute for a Cybatar benchmark run.
Composite focused on African institutional context with supporting finance and reliability.
| Model | Provider | Score | Coverage |
|---|---|---|---|
| No model yet has complete published Cybatar evidence for this index. | |||
Composite focused on enterprise planning, tool decisions, recovery and supporting reasoning.
| Model | Provider | Score | Coverage |
|---|---|---|---|
| No model yet has complete published Cybatar evidence for this index. | |||
Enterprise-readiness composite emphasising reliability, agentic planning and institutional work.
| Model | Provider | Score | Coverage |
|---|---|---|---|
| No model yet has complete published Cybatar evidence for this index. | |||
Composite focused on finance, quantitative reasoning and trustworthy evidence handling.
| Model | Provider | Score | Coverage |
|---|---|---|---|
| No model yet has complete published Cybatar evidence for this index. | |||
Broad model capability composite.
| Model | Provider | Score | Coverage |
|---|---|---|---|
| No model yet has complete published Cybatar evidence for this index. | |||
Composite focused on groundedness, abstention and dependable instruction execution.
| Model | Provider | Score | Coverage |
|---|---|---|---|
| No model yet has complete published Cybatar evidence for this index. | |||
Only reviewed, full-suite, non-synthetic runs executed through Cybatar appear here.
| Model | Provider | Benchmark | Version | Score | 95% CI | Published |
|---|---|---|---|---|---|---|
| No verified Cybatar-run model benchmark results have been published yet. | ||||||