Cybatar Model Intelligence

AI model benchmark evidence.

Cybatar separates independently researched external benchmark observations from benchmark runs executed and verified by Cybatar. Configuration, harness, benchmark version and provenance are preserved because scores are not comparable when evaluation conditions differ.

External Benchmark Research Baseline

Current published model evidence

102 sourced observations across 11 catalogue models · research snapshot 22 Aug 2026.

How to read this: these figures were reported by the named independent evaluator or model provider; they were not executed by Cybatar and do not feed Cybatar proprietary indexes. Always compare benchmark version, model configuration, tools and agent harness before drawing conclusions.
AutomationBench 2026

Enterprise workflow automation benchmark.

Higher is better
ModelScoreConfiguration / harnessProvenanceSource
Gemini 3.7 Flash
Google
30.4 %Model evaluation · thinking: highProvider ReportedGoogle DeepMind
13 Aug 2026
GPT-5.6 Sol
OpenAI
18.1 %Model evaluation · provider reported configuration: OpenAI launch evaluationProvider ReportedOpenAI
09 Jul 2026
GPT-5.6 Terra
OpenAI
15.2 %Model evaluation · provider reported configuration: OpenAI launch evaluationProvider ReportedOpenAI
09 Jul 2026
GPT-5.6 Luna
OpenAI
14.9 %Model evaluation · provider reported configuration: OpenAI launch evaluationProvider ReportedOpenAI
09 Jul 2026
MCP Atlas 2026

Multi-step workflows using Model Context Protocol tools.

Higher is better
ModelScoreConfiguration / harnessProvenanceSource
Gemini 3.1 Pro Preview
Google
69.2 %Model evaluation · thinking: highProvider ReportedGoogle DeepMind
19 Feb 2026
BrowseComp 2026

Agentic web-search benchmark.

Higher is better
ModelScoreConfiguration / harnessProvenanceSource
Gemini 3.1 Pro Preview
Google
85.9 %Model evaluation · search: yes · python: yes · browse: yesProvider ReportedGoogle DeepMind
19 Feb 2026
Artificial Analysis Coding Agent Index v1.3

Coding-agent benchmark where model and agent harness are both material to the result.

Higher is better
ModelScoreConfiguration / harnessProvenanceSource
Claude Opus 5
Anthropic
67 Claude Code · reasoning effort: xhighIndependentArtificial Analysis
22 Aug 2026
GPT-5.6 Sol
OpenAI
67 Codex · reasoning effort: maxIndependentArtificial Analysis
22 Aug 2026
Grok 4.5
xAI
64 Grok Build · reasoning effort: highIndependentArtificial Analysis
22 Aug 2026
Gemini 3.7 Flash
Google
57 OpenCode · thinking: highIndependentArtificial Analysis
22 Aug 2026
Gemini 3.1 Pro Preview
Google
30 Gemini CLI · thinking: highIndependentArtificial Analysis
22 Aug 2026
Code Arena 2026-08

Web-development coding arena Elo.

Higher is better
ModelScoreConfiguration / harnessProvenanceSource
Gemini 3.7 Flash
Google
1588 EloModel evaluation · Provider ReportedGoogle DeepMind
13 Aug 2026
DeepSWE 1.1

Long-horizon software engineering benchmark.

Higher is better
ModelScoreConfiguration / harnessProvenanceSource
GPT-5.6 Sol
OpenAI
72.7 %Model evaluation · provider reported configuration: OpenAI launch evaluationProvider ReportedOpenAI
09 Jul 2026
GPT-5.6 Terra
OpenAI
69.6 %Model evaluation · provider reported configuration: OpenAI launch evaluationProvider ReportedOpenAI
09 Jul 2026
GPT-5.6 Sol
OpenAI
69 %Codex · reasoning effort: maxIndependentArtificial Analysis
22 Aug 2026
GPT-5.6 Luna
OpenAI
67.2 %Model evaluation · provider reported configuration: OpenAI launch evaluationProvider ReportedOpenAI
09 Jul 2026
Gemini 3.7 Flash
Google
65.3 %Model evaluation · thinking: highProvider ReportedGoogle DeepMind
13 Aug 2026
Claude Opus 5
Anthropic
60 %Claude Code · reasoning effort: xhighIndependentArtificial Analysis
22 Aug 2026
Grok 4.5
xAI
60 %Grok Build · reasoning effort: highIndependentArtificial Analysis
22 Aug 2026
Gemini 3.7 Flash
Google
57 %OpenCode · thinking: highIndependentArtificial Analysis
22 Aug 2026
Gemini 3.1 Pro Preview
Google
14 %Gemini CLI · thinking: highIndependentArtificial Analysis
22 Aug 2026
FrontierCode 1.1 Main 1.1

Production code quality benchmark.

Higher is better
ModelScoreConfiguration / harnessProvenanceSource
Gemini 3.7 Flash
Google
43.6 %Model evaluation · thinking: highProvider ReportedGoogle DeepMind
13 Aug 2026
SciCode 2026

Scientific research coding benchmark.

Higher is better
ModelScoreConfiguration / harnessProvenanceSource
Gemini 3.1 Pro Preview
Google
59 %Model evaluation · thinking: highProvider ReportedGoogle DeepMind
19 Feb 2026
DeepSeek V4 Flash
DeepSeek
50 %Model evaluation · reasoning effort: max · release: 0731IndependentArtificial Analysis
31 Jul 2026
SWE-Atlas-QnA 2026

Repository understanding benchmark used in the Coding Agent Index.

Higher is better
ModelScoreConfiguration / harnessProvenanceSource
Claude Opus 5
Anthropic
55 %Claude Code · reasoning effort: xhighIndependentArtificial Analysis
22 Aug 2026
Grok 4.5
xAI
48 %Grok Build · reasoning effort: highIndependentArtificial Analysis
22 Aug 2026
GPT-5.6 Sol
OpenAI
43 %Codex · reasoning effort: maxIndependentArtificial Analysis
22 Aug 2026
Gemini 3.7 Flash
Google
31 %OpenCode · thinking: highIndependentArtificial Analysis
22 Aug 2026
Gemini 3.1 Pro Preview
Google
9 %Gemini CLI · thinking: highIndependentArtificial Analysis
22 Aug 2026
SWE-Bench Pro 2026

More diverse agentic software-engineering tasks.

Higher is better
ModelScoreConfiguration / harnessProvenanceSource
Gemini 3.1 Pro Preview
Google
54.2 %Model evaluation · single attempt: yes · thinking: highProvider ReportedGoogle DeepMind
19 Feb 2026
SWE-Bench Verified 2026

Agentic coding benchmark on real GitHub issues.

Higher is better
ModelScoreConfiguration / harnessProvenanceSource
Gemini 3.1 Pro Preview
Google
80.6 %Model evaluation · single attempt: yes · thinking: highProvider ReportedGoogle DeepMind
19 Feb 2026
Terminal-Bench 2.1

Agentic terminal-use benchmark.

Higher is better
ModelScoreConfiguration / harnessProvenanceSource
GPT-5.6 Sol
OpenAI
88.8 %Model evaluation · provider reported configuration: OpenAI launch evaluationProvider ReportedOpenAI
09 Jul 2026
GPT-5.6 Sol
OpenAI
88 %Codex · reasoning effort: maxIndependentArtificial Analysis
22 Aug 2026
GPT-5.6 Terra
OpenAI
87.4 %Model evaluation · provider reported configuration: OpenAI launch evaluationProvider ReportedOpenAI
09 Jul 2026
Gemini 3.7 Flash
Google
85.8 %Model evaluation · thinking: highProvider ReportedGoogle DeepMind
13 Aug 2026
Claude Opus 5
Anthropic
85 %Claude Code · reasoning effort: xhighIndependentArtificial Analysis
22 Aug 2026
Grok 4.5
xAI
85 %Grok Build · reasoning effort: highIndependentArtificial Analysis
22 Aug 2026
GPT-5.6 Luna
OpenAI
84.7 %Model evaluation · provider reported configuration: OpenAI launch evaluationProvider ReportedOpenAI
09 Jul 2026
Gemini 3.7 Flash
Google
83 %OpenCode · thinking: highIndependentArtificial Analysis
22 Aug 2026
DeepSeek V4 Flash
DeepSeek
79 %Model evaluation · reasoning effort: max · release: 0731IndependentArtificial Analysis
31 Jul 2026
Gemini 3.1 Pro Preview
Google
68.5 %Model evaluation · harness: Terminus-2 · thinking: highProvider ReportedGoogle DeepMind
19 Feb 2026
Gemini 3.1 Pro Preview
Google
68 %Gemini CLI · thinking: highIndependentArtificial Analysis
22 Aug 2026
Agent's Last Exam 2026

Multimodal desktop and operating-system agent tasks.

Higher is better
ModelScoreConfiguration / harnessProvenanceSource
Gemini 3.7 Flash
Google
26.3 %Model evaluation · thinking: highProvider ReportedGoogle DeepMind
13 Aug 2026
OSWorld 2.0

Agentic computer-use benchmark.

Higher is better
ModelScoreConfiguration / harnessProvenanceSource
Gemini 3.7 Flash
Google
47.9 %Model evaluation · thinking: highProvider ReportedGoogle DeepMind
13 Aug 2026
GDP.pdf 2026

Expert PDF and document comprehension benchmark.

Higher is better
ModelScoreConfiguration / harnessProvenanceSource
Gemini 3.7 Flash
Google
34 %Model evaluation · thinking: highProvider ReportedGoogle DeepMind
13 Aug 2026
GPT-5.6 Sol
OpenAI
30.7 %Model evaluation · provider reported configuration: OpenAI launch evaluationProvider ReportedOpenAI
09 Jul 2026
GPT-5.6 Terra
OpenAI
24.7 %Model evaluation · provider reported configuration: OpenAI launch evaluationProvider ReportedOpenAI
09 Jul 2026
GPT-5.6 Luna
OpenAI
22.7 %Model evaluation · provider reported configuration: OpenAI launch evaluationProvider ReportedOpenAI
09 Jul 2026
Tau3-Banking 1.0.1

Agentic banking tool-use benchmark.

Higher is better
ModelScoreConfiguration / harnessProvenanceSource
DeepSeek V4 Flash
DeepSeek
31 %Model evaluation · reasoning effort: max · release: 0731IndependentArtificial Analysis
31 Jul 2026
Artificial Analysis Intelligence Index v4.1.1

Composite model intelligence across nine evaluations.

Higher is better
ModelScoreConfiguration / harnessProvenanceSource
Claude Opus 5
Anthropic
63 Model evaluation · reasoning effort: maxIndependentArtificial Analysis
22 Aug 2026
GPT-5.6 Sol
OpenAI
61 Model evaluation · reasoning effort: maxIndependentArtificial Analysis
22 Aug 2026
Grok 4.6
xAI
61 Model evaluation · reasoning effort: highIndependentArtificial Analysis
12 Aug 2026
GPT-5.6 Terra
OpenAI
57 Model evaluation · reasoning effort: maxIndependentArtificial Analysis
22 Aug 2026
Gemini 3.7 Flash
Google
56 Model evaluation · thinking: highIndependentArtificial Analysis
22 Aug 2026
DeepSeek V4 Pro
DeepSeek
53 Model evaluation · reasoning effort: max · release: 0813IndependentArtificial Analysis
13 Aug 2026
DeepSeek V4 Flash
DeepSeek
52 Model evaluation · reasoning effort: max · release: 0731IndependentArtificial Analysis
31 Jul 2026
GPT-5.6 Luna
OpenAI
52 Model evaluation · reasoning effort: maxIndependentArtificial Analysis
22 Aug 2026
HealthBench Professional 2026

Professional healthcare benchmark.

Higher is better
ModelScoreConfiguration / harnessProvenanceSource
GPT-5.6 Sol
OpenAI
60.5 %Model evaluation · provider reported configuration: OpenAI launch evaluationProvider ReportedOpenAI
09 Jul 2026
GPT-5.6 Terra
OpenAI
57.7 %Model evaluation · provider reported configuration: OpenAI launch evaluationProvider ReportedOpenAI
09 Jul 2026
GPT-5.6 Luna
OpenAI
55.7 %Model evaluation · provider reported configuration: OpenAI launch evaluationProvider ReportedOpenAI
09 Jul 2026
AA-Briefcase 2026-08

Agentic knowledge-work benchmark focused on analytical quality and presentation.

Higher is better
ModelScoreConfiguration / harnessProvenanceSource
Claude Opus 5
Anthropic
1720 EloModel evaluation · reasoning effort: maxIndependentArtificial Analysis
24 Jul 2026
Grok 4.5
xAI
1313 EloModel evaluation · reasoning effort: highIndependentArtificial Analysis
22 Aug 2026
Claude Sonnet 5
Anthropic
1292 EloModel evaluation · reasoning effort: xhighIndependentArtificial Analysis
22 Aug 2026
DeepSeek V4 Flash
DeepSeek
1285 EloModel evaluation · reasoning effort: max · release: 0731IndependentArtificial Analysis
22 Aug 2026
Gemini 3.7 Flash
Google
1131 EloModel evaluation · thinking: highIndependentArtificial Analysis
22 Aug 2026
GDPval-AA v2

Agentic real-world professional knowledge-work benchmark.

Higher is better
ModelScoreConfiguration / harnessProvenanceSource
Claude Opus 5
Anthropic
1861 EloModel evaluation · reasoning effort: maxIndependentArtificial Analysis
24 Jul 2026
Grok 4.6
xAI
1753 EloModel evaluation · reasoning effort: highIndependentArtificial Analysis
12 Aug 2026
GPT-5.6 Sol
OpenAI
1679 EloModel evaluation · reasoning effort: xhighIndependentArtificial Analysis
22 Aug 2026
Claude Sonnet 5
Anthropic
1595 EloModel evaluation · reasoning effort: maxIndependentArtificial Analysis
22 Aug 2026
DeepSeek V4 Pro
DeepSeek
1590 EloModel evaluation · reasoning effort: max · release: 0813IndependentArtificial Analysis
22 Aug 2026
DeepSeek V4 Flash
DeepSeek
1559 EloModel evaluation · reasoning effort: max · release: 0731IndependentArtificial Analysis
31 Jul 2026
Gemini 3.7 Flash
Google
1525 EloModel evaluation · thinking: highProvider ReportedGoogle DeepMind
13 Aug 2026
Gemini 3.1 Pro Preview
Google
1317 EloModel evaluation · thinking: highProvider ReportedGoogle DeepMind
19 Feb 2026
Harvey LAB-AA 2026

Complex legal workflow evaluation.

Higher is better
ModelScoreConfiguration / harnessProvenanceSource
Gemini 3.7 Flash
Google
90.7 %Model evaluation · thinking: highProvider ReportedGoogle DeepMind
13 Aug 2026
MRCR v2 — 128k v2

Long-context multi-needle retrieval/reasoning at 128k context.

Higher is better
ModelScoreConfiguration / harnessProvenanceSource
Gemini 3.7 Flash
Google
97 %Model evaluation · thinking: highProvider ReportedGoogle DeepMind
13 Aug 2026
Gemini 3.1 Pro Preview
Google
84.9 %Model evaluation · context: 128k_averageProvider ReportedGoogle DeepMind
19 Feb 2026
MRCR v2 — 1M v2

Long-context multi-needle retrieval/reasoning at 1M context.

Higher is better
ModelScoreConfiguration / harnessProvenanceSource
Gemini 3.1 Pro Preview
Google
26.3 %Model evaluation · context: 1m_pointwiseProvider ReportedGoogle DeepMind
19 Feb 2026
MMMLU 2026

Multilingual question answering benchmark.

Higher is better
ModelScoreConfiguration / harnessProvenanceSource
Gemini 3.1 Pro Preview
Google
92.6 %Model evaluationProvider ReportedGoogle DeepMind
19 Feb 2026
MMMU-Pro (no tools) 2026

Multimodal understanding and reasoning without tools.

Higher is better
ModelScoreConfiguration / harnessProvenanceSource
GPT-5.6 Sol
OpenAI
83 %Model evaluation · provider reported configuration: OpenAI launch evaluationProvider ReportedOpenAI
09 Jul 2026
GPT-5.6 Terra
OpenAI
80.7 %Model evaluation · provider reported configuration: OpenAI launch evaluationProvider ReportedOpenAI
09 Jul 2026
Gemini 3.1 Pro Preview
Google
80.5 %Model evaluation · tools: noneProvider ReportedGoogle DeepMind
19 Feb 2026
GPT-5.6 Luna
OpenAI
78.4 %Model evaluation · provider reported configuration: OpenAI launch evaluationProvider ReportedOpenAI
09 Jul 2026
MMMU-Pro (with tools) 2026

Multimodal understanding and reasoning with tools.

Higher is better
ModelScoreConfiguration / harnessProvenanceSource
GPT-5.6 Sol
OpenAI
84.6 %Model evaluation · provider reported configuration: OpenAI launch evaluationProvider ReportedOpenAI
09 Jul 2026
GPT-5.6 Terra
OpenAI
82 %Model evaluation · provider reported configuration: OpenAI launch evaluationProvider ReportedOpenAI
09 Jul 2026
GPT-5.6 Luna
OpenAI
79.5 %Model evaluation · provider reported configuration: OpenAI launch evaluationProvider ReportedOpenAI
09 Jul 2026
FrontierMath Tier 1-3 v2

Frontier mathematical reasoning, tiers 1–3.

Higher is better
ModelScoreConfiguration / harnessProvenanceSource
GPT-5.6 Sol
OpenAI
89 %Model evaluation · provider reported configuration: OpenAI launch evaluationProvider ReportedOpenAI
09 Jul 2026
GPT-5.6 Terra
OpenAI
84.9 %Model evaluation · provider reported configuration: OpenAI launch evaluationProvider ReportedOpenAI
09 Jul 2026
GPT-5.6 Luna
OpenAI
78.6 %Model evaluation · provider reported configuration: OpenAI launch evaluationProvider ReportedOpenAI
09 Jul 2026
FrontierMath Tier 4 v2

Frontier mathematical reasoning, tier 4.

Higher is better
ModelScoreConfiguration / harnessProvenanceSource
GPT-5.6 Sol
OpenAI
83 %Model evaluation · provider reported configuration: OpenAI launch evaluationProvider ReportedOpenAI
09 Jul 2026
GPT-5.6 Terra
OpenAI
68.3 %Model evaluation · provider reported configuration: OpenAI launch evaluationProvider ReportedOpenAI
09 Jul 2026
GPT-5.6 Luna
OpenAI
58.5 %Model evaluation · provider reported configuration: OpenAI launch evaluationProvider ReportedOpenAI
09 Jul 2026
ARC-AGI-2 verified-2026

Abstract reasoning on novel logic tasks.

Higher is better
ModelScoreConfiguration / harnessProvenanceSource
Gemini 3.1 Pro Preview
Google
77.1 %Model evaluation · ARC Prize verified: yes · thinking: highProvider ReportedGoogle DeepMind
19 Feb 2026
GPQA Diamond 2026

Graduate-level scientific reasoning benchmark.

Higher is better
ModelScoreConfiguration / harnessProvenanceSource
GPT-5.6 Sol
OpenAI
94.6 %Model evaluation · provider reported configuration: OpenAI launch evaluationProvider ReportedOpenAI
09 Jul 2026
Gemini 3.1 Pro Preview
Google
94.3 %Model evaluation · tools: none · thinking: highProvider ReportedGoogle DeepMind
19 Feb 2026
GPT-5.6 Terra
OpenAI
92.9 %Model evaluation · provider reported configuration: OpenAI launch evaluationProvider ReportedOpenAI
09 Jul 2026
GPT-5.6 Luna
OpenAI
92.3 %Model evaluation · provider reported configuration: OpenAI launch evaluationProvider ReportedOpenAI
09 Jul 2026
DeepSeek V4 Flash
DeepSeek
91 %Model evaluation · reasoning effort: max · release: 0731IndependentArtificial Analysis
31 Jul 2026
HLE-Verified 2026-08

Verified multidisciplinary expert reasoning set reported by Google DeepMind.

Higher is better
ModelScoreConfiguration / harnessProvenanceSource
Gemini 3.7 Flash
Google
53.6 %Model evaluation · thinking: highProvider ReportedGoogle DeepMind
13 Aug 2026
Humanity's Last Exam 2026

Broad expert-level reasoning and knowledge benchmark.

Higher is better
ModelScoreConfiguration / harnessProvenanceSource
Gemini 3.1 Pro Preview
Google
44.4 %Model evaluation · tools: none · thinking: highProvider ReportedGoogle DeepMind
19 Feb 2026
DeepSeek V4 Flash
DeepSeek
37 %Model evaluation · reasoning effort: max · release: 0731IndependentArtificial Analysis
31 Jul 2026
AA-Omniscience Hallucination Rate 2026-08

Percentage of attempted answers classified as hallucinations.

Lower is better
ModelScoreConfiguration / harnessProvenanceSource
DeepSeek V4 Flash
DeepSeek
84 %Model evaluation · reasoning effort: max · release: 0731IndependentArtificial Analysis
31 Jul 2026
AA-Omniscience Index 2026-08

Rewards correct answers and penalizes hallucinations; refusal is not penalized.

Higher is better
ModelScoreConfiguration / harnessProvenanceSource
DeepSeek V4 Flash
DeepSeek
-16 Model evaluation · reasoning effort: max · release: 0731IndependentArtificial Analysis
31 Jul 2026

Cybatar composite indexes

Cybatar composite scores appear only when a model has complete published Cybatar-run evidence for every required component suite. External research observations above never substitute for a Cybatar benchmark run.

Cybatar African Context Intelligence Index

Composite focused on African institutional context with supporting finance and reliability.

ModelProviderScoreCoverage
No model yet has complete published Cybatar evidence for this index.
Cybatar Agentic Readiness Index

Composite focused on enterprise planning, tool decisions, recovery and supporting reasoning.

ModelProviderScoreCoverage
No model yet has complete published Cybatar evidence for this index.
Cybatar Enterprise Model Index

Enterprise-readiness composite emphasising reliability, agentic planning and institutional work.

ModelProviderScoreCoverage
No model yet has complete published Cybatar evidence for this index.
Cybatar Financial Intelligence Index

Composite focused on finance, quantitative reasoning and trustworthy evidence handling.

ModelProviderScoreCoverage
No model yet has complete published Cybatar evidence for this index.
Cybatar Model Intelligence Index

Broad model capability composite.

ModelProviderScoreCoverage
No model yet has complete published Cybatar evidence for this index.
Cybatar Trust & Reliability Index

Composite focused on groundedness, abstention and dependable instruction execution.

ModelProviderScoreCoverage
No model yet has complete published Cybatar evidence for this index.

Verified Cybatar benchmark runs

Only reviewed, full-suite, non-synthetic runs executed through Cybatar appear here.

ModelProviderBenchmarkVersionScore95% CIPublished
No verified Cybatar-run model benchmark results have been published yet.