AI Infrastructure & MLOps · entry 06/09 · 3 min read
What it costs the grid
Electricity, water and carbon are real constraints on AI — and most of the numbers in circulation are extrapolations from thin disclosure, which is worth knowing before you quote one.
Training is the headline; inference is the bill
Public discussion fixates on one enormous training run, because it is a discrete, dramatic, quotable event. For models people actually use at scale it is the smaller half: the only two first-party lifecycle accountings published put training at roughly a third to two-fifths of total machine-learning energy. When a given model passes its training energy depends entirely on request volume — weeks at consumer-assistant scale, never for a model serving one company — so there is no universal crossover to quote. The lever is how many tokens you serve and on what, the same lever that decides your invoice.
Why the per-prompt numbers span a factor of a hundred
Published figures run from a tenth of a watt-hour for a small model at low precision to tens of watt-hours for a long reasoning trace on a frontier one. Mostly that is not disagreement about one quantity; it is real variation between events that are not alike and should not share a unit. Disclosure has improved from nothing to something — one provider published a per-prompt methodology in 2025, itemising data-center overhead and idle capacity, landing near a quarter of a watt-hour for a median text prompt. What is still withheld decides the answer: model size, batch size, and what the "median" prompt was.
Firmer ground than that
Data centers as a whole take roughly one and a half percent of world electricity, AI perhaps a sixth of that — a fraction of one percent today, heading toward just under three percent for data centers overall by 2030. The growth rate is the striking figure, not the level. (Even the historical number is contested; there is no common definition of what counts as a data center.) The binding constraint is local rather than global: interconnection queues measured in years, transformer lead times measured in multiples of one, and the politics of siting. Water is more complicated than the headline suggests — evaporative cooling consumes a great deal on site, while closed-loop and air-cooled designs consume little on site and buy that with more electricity, which evaporates water at the power plant instead. Indirect water, through generation, is usually the larger share.
The two levers that dominate
Where the electricity comes from, and when you use it. Grid carbon intensity spans well over an order of magnitude between regions and swings by a large factor across the day, so identical computation can differ enormously in emissions with no change to the model. Efficiency improves quickly too — better hardware, quantization, smaller models doing more — and also makes the thing cheaper to do more of. So far the total has risen while the per-token cost fell.
Failure mode
Comparing one prompt to a household activity. Seconds of television, drops of water, a lightbulb for a minute: memorable, and usually rigged by boundary — the AI side counts data-center energy only, the television side counts the whole appliance. It also collapses a marginal, regional, time-varying quantity into one number that cannot be right for everyone. If the footprint bears on your decision, ask the three things that determine it: which model, on which hardware, on whose grid. Treat a figure missing all three as a rhetorical device, not a measurement.