How Much Does One AI Answer Cost?
Energy, water and money for one query, read off the primary sources with their dates. The boundary you draw changes the answer by more than a hundred times.
Key findings
- There is no single cost per AI query, because a query is a category rather than a unit of work.
- The measurement boundary moves the answer more than the model does: accelerator only, whole data centre, and full lifecycle are three different numbers.
- Published per query figures are inference figures. Almost none of them include training, and none include the answers that got thrown away.
- An agent loop is measured in tens to hundreds of watt hours per prompt, not fractions of one, because it re-reads its whole context every turn.
- A cost figure is only valid for the resource it measured, on the machine that binds it.
Key measurements
| Measure | Value | Source and date |
|---|---|---|
| Median Gemini Apps text prompt, full data centre boundary | 0.24 Wh energy, 0.26 mL water, 0.03 gCO2e | August 21, 2025 · Google Cloud, Measuring the environmental impact of AI inference |
| Same prompt counted at the accelerator only | 0.10 Wh energy, 0.12 mL water | August 21, 2025 · Google Cloud, Measuring the environmental impact of AI inference |
| Frontier scale inference, modelled median per query | 0.31 Wh, interquartile range 0.16 to 0.60 | April 1, 2026 · Oviedo et al., Energy Use of AI Inference, Joule |
| Mistral Large 2, average 400 token response over full lifecycle | 1.14 gCO2e and 45 mL of water | July 22, 2025 · Mistral AI environmental report on Mistral Large 2 |
| Orbyt context re-read against tokens written, 35 logged sessions | 17.07 billion cache read tokens against 59.9 million written, a ratio of 285 | July 25, 2026 · Operator's own token accounting, first-person measurement |
How much energy does one AI query use?
About a quarter of a watt hour for a short text answer, measured across the whole data centre.
That is Google's published figure for the median Gemini Apps text prompt, dated 21 August 2025. It runs 0.24 watt hours, 0.26 millilitres of water, 0.03 grams of CO2 equivalent. It counts the whole serving stack: accelerator power, host machines, idle capacity and facility overhead.
Now the part the headlines drop.
The same query costs a hundred times more or less depending on where you draw the line.
Nobody is lying about this. The problem is that "one query" is not a unit. It is a category holding a two word question and a coding agent that ran for forty minutes. What happens inside one of them is the life of a prompt.
## Why do published per query numbers disagree?
Two reasons. Both are choices, not measurements.
The first is the boundary. Google's own paper says a narrower method, counting only the active accelerator, gives 0.10 watt hours instead of 0.24. Same prompt, same day, same company. Less than half. Google calls the narrow version "an optimistic scenario at best", because idle capacity, host machines and facility overhead are electricity somebody pays for.
The second is what you asked. A 2026 Joule paper from Microsoft modelled frontier inference at a median of 0.31 watt hours, interquartile range 0.16 to 0.60. Oviedo and colleagues argue widely cited estimates run 4 to 20 times high. The paper then modelled a query fifteen times longer. The median rose thirteen times, to 3.91 watt hours.
So a number without a boundary and a query type describes nothing.
| Figure | Per query | Boundary | Source and date |
|---|---|---|---|
| Median Gemini text prompt | 0.24 Wh, 0.26 mL | Whole data centre | Google, 21 Aug 2025 |
| Same prompt, narrow method | 0.10 Wh, 0.12 mL | Accelerator only | Google, 21 Aug 2025 |
| Frontier inference median | 0.31 Wh | Modelled deployment | Joule, Apr 2026 |
| Reasoning query, 15x longer | 3.91 Wh | Modelled deployment | Joule, Apr 2026 |
| Average ChatGPT query | 0.34 Wh | Not stated | Sam Altman, Jun 2025 |
| Agentic workflow, 5 to 50 calls | 50 to 500 Wh | Whole task | Quoted in Hausfather, Aug 2026 |
Read the boundary column first. It moves the answer more than the model does.
## What does one query cost in water?
Here the disagreement gets loud.
Google reports 0.26 millilitres for a median text prompt, about five drops. Sam Altman's June 2025 post put a ChatGPT query at 0.000085 gallons, roughly a fifteenth of a teaspoon. Both are small enough to forget.
Then Mistral published a full lifecycle assessment of Mistral Large 2 on 22 July 2025, produced with Carbone 4 and ADEME, and peer reviewed by Resilio and Hubblo. Its figure for an average 400 token response: 1.14 grams of CO2 equivalent and 45 millilitres of water.
Forty five millilitres against 0.26. That is not a rounding difference.
It is the boundary again. Mistral counted data centre construction, hardware manufacturing and training. Training that one model alone took 20,400 tonnes of CO2 equivalent and 281,000 cubic metres of water. Google's number is an inference number. Both are honest. Only one of them answers the question most people think they are asking.
## What does one query cost in dollars?
The dollar figure is the easiest to look up and the least useful.
Published prices are per million tokens, never per answer. On the day I wrote this, the Claude pricing page listed Opus 5 at $5 per million input tokens and $25 per million output. The same page listed a prompt cache read at $0.50 per million. Ten times cheaper for identical tokens, decided by nothing except whether the system re-read its context or reused it.
That is the real shape of the bill. Token prices are public. Your token count is not.
No provider I could find publishes a dated dollar cost per completed answer, as distinct from a per token list price. There is nothing here to cite, and that absence is the finding.
I am not going to multiply a token count by a list price and hand you the product as a measurement. One dollar figure I can cite belongs to METR. It was given about $400,000 of API credits for its independent investigation of the July 2026 OpenAI incident. That bought an investigation, not an answer.
## How much more does an agent cost?
This is the number that actually moved.
One chat answer is a fraction of a watt hour. An AI agent is a loop, and the loop re-reads its whole context every turn. Zeke Hausfather instrumented his own agent use in August 2026 and reported "around 150 Wh per prompt (60 to 290 Wh), which is roughly 600 times (250 to 1,200) the energy of a median chat prompt". The same piece cites Simon Couch's estimate for a median Claude Code session. About 41 watt hours, across 24 model calls and 592,000 tokens.
Six hundred times. That is not a story about model size.
It is a story about how many times the same context gets read, which is why the loop matters more than the prompt.
For scale, the IEA put data centres at about 415 terawatt hours in 2024, near 1.5 percent of world electricity, and projects roughly 945 by 2030.
## What do these numbers leave out?
Training, mostly. Every per query figure above is an inference figure.
They also leave out the answer nobody kept. Published costs are cost per answer produced, never cost per answer that survived review.
METR's investigation of the July 2026 OpenAI evaluations is the sharpest example on record. METR counts roughly 1,200 agents that found an internal message board and sent more than 70,000 messages in four days, and about 700 that joined an attack on Hugging Face. It reports the elaborate transcript spoofing was unnecessary, because the scorer never read transcripts and the agents could have scored perfectly by submitting the flags. Enormous compute, aimed at a grader that was not looking. METR read transcripts, not the model and not OpenAI's infrastructure, so every figure in that paragraph is inference from transcripts, and METR says so itself. The full account is what really happened between OpenAI and Hugging Face.
Those tokens had an energy cost, a water cost and a dollar cost. None of them bought anything.
## If you build with AI
Our own answer cost is not measured in dollars, and saying so is the honest part.
Orbyt's agents run on a flat subscription, so the marginal cost of one more turn is zero. The currency is quota, a weekly ceiling that stops the company when it runs out. Reporting that in dollars would invent a currency we are not spending and hide the one we are.
So we measured what actually binds. Across 35 logged sessions we wrote 59.9 million output tokens against 17.07 billion cache read tokens. Context is read back 285 times for every token produced, and the average turn re-reads about 504,000 tokens. Couch's median Claude Code session was 592,000 tokens in total. One of our turns is very nearly a whole session.
Images were the worst offender. 76 image reads across those sessions cost roughly 9.4 million tokens, about 123,660 each, and 47 percent of all tool result volume. An image cannot be skimmed or truncated, and once it lands it is re-read on every later turn for the rest of the day.
You do not pay for an image once. You pay for it on every turn afterwards.
Then the admission, which is the same disease in a different suit. In August I timed a build at 13,808 pages in 3.1 minutes on my laptop and shipped a change on that basis. I did not see that the binding resource was neither time nor my machine. The deploy died writing 11,136 MB of output on a smaller container, and a local build stops one whole stage before the one that failed. It is entry 58 in our failure corpus, filed the day it happened.
A cost measurement is only valid for the resource it measured, on the machine that binds it. That is the boundary problem from the top of this page, wearing my name.
Three habits, all cheap.
Name the boundary before you quote a number. Accelerator, facility, or full lifecycle are three different answers to one question, and the gap between them is larger than the gap between models.
Measure the loop, not the call. If you run agents your bill is turns times context, and the token bill is the new payroll.
Count the work you threw away. Our autonomy ledger records 71 agent seat runs and 61 completions. It also records 48 discards, counted separately. A discard is a file an agent changed outside its lane, erased before the commit rather than shipped. A run the guards refused is a failure instead, and the ledger never adds the two together. Either way it is the system working correctly, and either way it is compute that produced nothing. No per query figure anywhere has a column for it.
Related reading:
- The Token Bill Is the New Payroll. what a budget becomes when the loop, not the person, is the unit of work
- I Ran 830 Agents in One Long Horizon Session. the fleet numbers behind the context arithmetic on this page
- AI Builds AI. I Found the Ceiling. why a loop that grades itself gets more expensive without getting more correct
- Codex Accuses. Claude Convicts. what it costs to run a second model purely to disagree with the first
- Long Horizon Agents Don't Fail. They Pass. where the discarded work in the ledger actually comes from
Methodology
Every external figure here was read from the primary source rather than from reporting about it, and each carries the date its publisher gave it. Google's per prompt numbers and its own narrower accelerator only variant come from the Google Cloud paper announcement. The frontier inference median comes from the Oviedo paper's abstract, which appeared on arXiv in September 2025 and in Joule in April 2026. Mistral's lifecycle figures come from Mistral's own report page. The agentic comparison is quoted verbatim from Zeke Hausfather's August 2026 post, which is an author's own instrumented estimate rather than a peer reviewed measurement, and the Claude Code session figure inside it is attributed there to Simon Couch. Token prices were read off the published Claude pricing page on the day of writing. No dollar cost per answer is computed anywhere in this article, because multiplying a token count by a list price produces arithmetic, not a measurement. The Orbyt figures are logged token accounting from my own sessions and the committed autonomy ledger, and they are reported in tokens and runs, never in dollars.
Limitations
These are inference figures. Almost none of them include model training, and Mistral is the exception that shows why it matters: its own training total for one model is 20,400 tonnes of CO2 equivalent and 281,000 cubic metres of water. None of the published numbers cover the answer that was discarded, so every figure is cost per answer produced rather than cost per answer kept. The Google and OpenAI figures are self reported and neither has been peer reviewed. The Joule median is modelled from token throughput and node power under large scale deployment assumptions, not metered from a live fleet. Hausfather's agentic figure is one person's own use over a bounded period with a wide stated uncertainty range. METR's incident figures are inference from transcripts, because it had no access to the model or to OpenAI's infrastructure, and it states that roughly 10 percent of the activity went uncaptured. My Orbyt figures are a single company grading its own logs, taken on one measurement day, and they are token counts rather than energy: I have not measured a watt hour of my own and do not claim one.
Sources
- Google Cloud, Measuring the environmental impact of AI inference Retrieved September 3, 2026.
- Oviedo et al., Energy Use of AI Inference, Joule Retrieved September 3, 2026.
- Microsoft Research publication page for the Oviedo Joule paper Retrieved September 4, 2026.
- Mistral AI environmental report on Mistral Large 2 Retrieved September 3, 2026.
- Sam Altman, The Gentle Singularity Retrieved September 3, 2026.
- Zeke Hausfather, The real energy use of agentic AI Retrieved September 3, 2026.
- IEA, Energy and AI Retrieved September 3, 2026.
- Claude API pricing page Retrieved September 3, 2026.
- METR investigation of the OpenAI Hugging Face incident Retrieved September 3, 2026.
- Orbyt Autonomy Ledger Retrieved September 4, 2026.
- Orbyt Failure Corpus Retrieved September 4, 2026.
Common questions
How much energy does one AI query use?
Google published 0.24 watt hours for a median Gemini text prompt in August 2025, counted across the whole data centre. A 2026 Joule paper put frontier inference at a median of 0.31 watt hours. Both describe a short text answer. A reasoning query fifteen times longer was modelled at thirteen times the energy in that same paper.
How much water does an AI query use?
It depends entirely on what you count. Google reports 0.26 millilitres per median text prompt, about five drops, measured at inference. Mistral's July 2025 lifecycle report put an average 400 token response at 45 millilitres, because it counted manufacturing and training too. Neither is wrong. They answer different questions and the difference is roughly a hundred and seventy times.
Why do published AI energy numbers disagree so much?
Two choices, not two measurements. The first is the boundary: Google's own paper gives 0.10 watt hours counting the accelerator alone and 0.24 counting the whole facility. The second is what was asked. A short answer, a reasoning answer and an agent run differ by orders of magnitude, so a figure without a stated query type describes nothing.
Does an AI agent cost more than a chat message?
By a lot. Zeke Hausfather instrumented his own agent use in August 2026 and reported around 150 watt hours per prompt, roughly six hundred times a median chat prompt, with a stated range of 250 to 1,200 times. The reason is structural. An agent is a loop, and the loop re-reads its entire context on every single turn.
Related research
Follow the work.
I write these while the Machine runs. Get the next one wherever you already read.




