Cost per successful task
Runtime cost is incomplete unless failures, retries, human review and exception handling are included. The useful unit is work that reaches an acceptable outcome.
Cybatar connects model pricing, workload cost, enterprise outcomes and value-engineering methods so institutions can compare the full cost of machine work with the value it actually creates.
AI economics is the discipline of connecting the full cost of AI activity—including model usage, infrastructure, human review, failure, governance and implementation—to measurable outcomes, validated value and investment decisions.
A system-level method for comparing the full cost of machine work with measurable institutional value.
Runtime cost is incomplete unless failures, retries, human review and exception handling are included. The useful unit is work that reaches an acceptable outcome.
Where comparable capability evidence exists, CEES relates capability to a standardised workload basket and normalises the strongest observed capability-per-dollar result to 100. It is a research comparator—not a forecast of every production workload.
Methodology status: active · effective 23 Aug 2026
Use a standardised reference basket to compare current model economics without pretending the basket predicts your own production workload.
Knowledge Assistant. A bounded enterprise knowledge interaction used for price comparison across models. Basket: 8,000 input + 2,000 output tokens; 25% cached-input assumption.
| Provider | Model | Cost / task | Cost / 1,000 tasks | Capability evidence | CEES |
|---|---|---|---|---|---|
| DeepSeek | DeepSeek V4 Flash | USD 0.0014 | USD 1.41 |
52.0 AA Intelligence Index |
100.0 / 100 |
| OpenAI | GPT-5.6 Luna | USD 0.0036 | USD 3.64 |
52.0 AA Intelligence Index |
38.6 / 100 |
| DeepSeek | DeepSeek V4 Pro | USD 0.0044 | USD 4.36 |
53.0 AA Intelligence Index |
32.9 / 100 |
| Gemini 3.7 Flash | USD 0.0122 | USD 12.15 |
56.0 AA Intelligence Index |
12.5 / 100 | |
| OpenAI | GPT-5.6 Terra | USD 0.0364 | USD 36.40 |
57.0 AA Intelligence Index |
4.2 / 100 |
| xAI | Grok 4.6 | USD 0.0500 | USD 50.00 |
61.0 AA Intelligence Index |
3.3 / 100 |
| OpenAI | GPT-5.6 Sol | USD 0.0720 | USD 72.00 |
61.0 AA Intelligence Index |
2.3 / 100 |
| Anthropic | Claude Opus 5 | USD 0.0900 | USD 90.00 |
63.0 AA Intelligence Index |
1.9 / 100 |
| Anthropic | Claude Sonnet 5 | USD 0.0360 | USD 36.00 | No common capability observation | — |
| xAI | Grok 4.5 | USD 0.0492 | USD 49.20 | No common capability observation | — |
| Cohere | Command A | USD 0.0400 | USD 40.00 | No common capability observation | — |
| Gemini 3.6 Flash | USD 0.0122 | USD 12.15 | No common capability observation | — | |
| Gemini 3.5 Flash | USD 0.0273 | USD 27.30 | No common capability observation | — | |
| Gemini 3.5 Flash-Lite | USD 0.0069 | USD 6.86 | No common capability observation | — | |
| Mistral AI | Mistral Large 3 | USD 0.0061 | USD 6.10 | No common capability observation | — |
| Mistral AI | Mistral Medium 3.5 | USD 0.0243 | USD 24.30 | No common capability observation | — |
| Mistral AI | Mistral Small 4 | USD 0.0021 | USD 2.13 | No common capability observation | — |
Published observations remain connected to their source, date and confidence so they can inform decisions without being mistaken for universal truths.
88% of surveyed organisations reported using AI in at least one business function in 2025.
Stanford HAI
Open source79% of surveyed organisations reported regular generative AI use in at least one function in 2025.
Stanford HAI
Open source62% of respondents in McKinsey’s Enterprise AI FinOps survey had moved beyond experimentation into active AI deployment.
McKinsey & Company · 20 Jul 2026
Open sourceMcKinsey cites research indicating that context-heavy agentic tasks can consume roughly 1,000× more tokens than simpler code-reasoning or chat tasks. Exact ratios vary by workload.
McKinsey & Company · 08 Jul 2026
Open sourceMcKinsey cites production coding-workflow research indicating about 60% of agentic task cost can be tied to checking, repairing and re-verifying outputs. Exact shares vary by workload.
McKinsey & Company · 08 Jul 2026
Open source93% of qualified respondents in McKinsey’s May 2026 Enterprise AI FinOps survey reported exceeding their AI budgets.
McKinsey & Company · 20 Jul 2026
Open sourceMcKinsey reports AI spend increasing nearly fourfold as organisations move from isolated use cases to enterprise-wide adoption.
McKinsey & Company · 20 Jul 2026
Open sourceMcKinsey reports that 20–30% of AI spend is often unaccounted for because investments are fragmented across vendors, tools and commercial models.
McKinsey & Company · 20 Jul 2026
Open sourceMcKinsey estimates only about 20–25% of companies have mature AI FinOps practices.
McKinsey & Company · 20 Jul 2026
Open sourceMcKinsey reports organisations with high forecasting maturity saving 10% more on AI spend than peers on average.
McKinsey & Company · 20 Jul 2026
Open sourceAbout one-third of surveyed organisations had achieved 20–30% savings through active AI-spend optimisation actions.
McKinsey & Company · 20 Jul 2026
Open sourceThe 2026 AI Index summarizes studies reporting roughly 14–15% gains in customer support, 26% in software development and 50% in marketing output, while noting smaller gains in deeper-reasoning work.
Stanford HAI
Open sourceMcKinsey cites evidence that token usage can vary by up to 30× when an agent executes the same task, reinforcing the need to model cost as a distribution rather than a fixed unit price.
McKinsey & Company · 20 Jul 2026
Open sourcePwC’s 2026 Global AI Jobs Barometer reports a 62% average wage premium for workers with AI skills compared with comparable roles without those skills.
PwC
Open sourcedocumentation time saved per patient
documentation time saved per patient: 4 minutes
planned users
planned users: 505,000 people
trial users
trial users: 30,000 people
administrative time saved per user daily
administrative time saved per user daily: 43 minutes
sepsis identification improvement
sepsis identification improvement: 80 percent
average patient stay reduction
average patient stay reduction: 1 day
false positive reduction
false positive reduction: 95 percent
security operations productivity gain
security operations productivity gain: 46.7 percent
travel mailbox email reduction
travel mailbox email reduction: 90 percent
complex query time before
complex query time before: 5 days
complex query time after
complex query time after: 30 seconds
annual salary cost avoidance
annual salary cost avoidance: 120,000 CAD
manual biomarker validation effort expected to be saved
manual biomarker validation effort expected to be saved: 5 years
first-call resolution
first-call resolution: 91 percent
cost per chat reduction
cost per chat reduction: 77 percent
complex interactions handled concurrently
complex interactions handled concurrently: 3 x
self-service engagement growth
self-service engagement growth: 49 percent
queries answered weekly
queries answered weekly: 10,000 queries
Model task volume, human baseline cost, runtime consumption, review, failure economics, implementation investment, revenue uplift and risk avoidance before the institution scales AI spend.