A point of view for regulated water and wastewater utilities.
In January 2026, New York’s State Comptroller released an audit of how the state tracks its lead service lines. Auditors pulled a sample of records and checked them against what had actually happened in the field. More than a quarter were wrong — 105 of the 371 lines reviewed. And they were wrong in one direction: pipes that crews had already dug up and replaced were still on the books as “lead” or “unknown.”
The work had been done. The record never caught up.
That is one finding in one state audit, but it is the whole problem in miniature. A crew closed out a replacement. An inventory reported a status. No system’s job was to make the first update the second — so the official answer drifted from the truth and stayed there until an auditor went looking. This is the record that now has to satisfy a federal rule, a primacy agency, and eventually a resident who wants to know what their water runs through.
Widen the lens and it is not an outlier. Nearly three thousand New York water systems owed an initial service line inventory by October 2024; a third missed the deadline, and many that filed on time leaned on “unknown” — not because the material is unknowable, but because the basis for knowing it was never written down anywhere a system could reach. The utility often has the answer. It cannot always show its work.
This is the situation into which the water sector is now buying AI. And it is why most of what is being sold today will not survive contact with a regulated utility.
The double squeeze
Two curves are crossing in North American water, and their intersection defines the next decade.
The first is the regulatory ratchet. The Lead and Copper Rule Improvements take effect in November 2027, and with them a decade of service line replacement built on inventories that must stay accurate, current, and provable. Revised Consumer Confidence Report requirements land in 2027. PFAS is the instructive case: the treatment standards themselves are being fought over — the compliance deadline pushed to 2031, some of the 2024 limits proposed for rescission — and yet the monitoring and reporting obligations underneath them have survived every rollback, with initial monitoring still due in 2027. The paperwork outlasts the politics. Layer on consent decrees — the EPA has entered scores of them with municipal systems since the 1990s, carrying tens of billions of dollars in mandated work and years of progress reporting — and the risk assessments and response plans that AWIA requires utilities to recertify on a rolling cycle.
Look closely at each of these and you find the same thing underneath: a documentation obligation. Prove the inventory. Prove the sampling. Prove the maintenance happened, on this asset, at this time, for this reason.
And regulators have started asking about AI itself. In March 2026, the Arizona Corporation Commission opened the first formal state inquiry into utilities’ use of artificial intelligence — how it is paid for, how it is secured, and how its decisions can be explained. Other commissions will follow. They always do.
The second curve is the workforce cliff. Up to a third of the water sector’s workforce is expected to retire within the decade, and most utilities report continuing losses of operators, engineers, and managers year over year. What walks out the door is not headcount. Headcount can be replaced. What walks out is the why — the valve that needs a quarter turn past closed, the lift station that trips in wet weather but only from the north cell, the reason a sampling site was chosen in 2009. The knowledge that answers an auditor’s question years from now is, right now, mostly in the heads of people who will not be there to answer it.
That is the squeeze: obligations that demand more memory, more reconstruction, more proof — landing on organizations that hold less of all three every year.
There is a dollar version of this, and it is worth saying plainly, because the rest of this argument lives on the axis of proof rather than price. Every audit answered by archaeology is billable hours — staff and consultants rebuilding a basis that should have been a query. Every truck rolled twice because the work order and the GIS feature disagreed about which asset is which is margin lost to fragmentation. Every retirement that takes undocumented judgment with it gets re-derived later at the cost of a failure, a callback, or a consultant. Defensibility can read like insurance against an audit half a decade away. It is also the same foundation that stops paying the fragmentation tax this year. The squeeze is what fuses the two: obligations rising while the people who held the answers leave is precisely the condition under which “nice to be able to prove it” becomes “cannot operate without proving it.”
Why the current wave of AI won’t survive contact
Walk any water conference floor and the offer is capabilities: leak detection, predictive maintenance, a copilot for your hydraulic model, a chatbot trained on your O&M manuals. Some of these are good tools. None of them addresses the squeeze, and three patterns explain why.
The point solution adds an eleventh silo to ten. Each pilot arrives with its own data ingestion, its own asset naming, its own store of conclusions. The leak model doesn’t know what the work order system knows. The chatbot doesn’t know what the leak model found. A utility that buys five AI tools ends up with five more places where the truth is partially held and nowhere it is fully held.
The chatbot over documents answers confidently and unaccountably. It retrieves passages that resemble the question and composes a fluent response. It cannot tell you whether the SOP it quoted is the current revision. It cannot connect the manual to the asset, the asset to the work order, the work order to the crew. Ask it why, and it will produce an explanation — not evidence. Those are different things, and a regulator knows the difference.
The pilot stalls where the systems meet. Ask utilities why AI initiatives die between pilot and production and the answer is rarely the model. SCADA, GIS, the CMMS, LIMS, and billing each run stably on their own and disagree with each other about everything: asset identities, naming, timestamps. The AI needs one continuous picture. The utility has five partial ones. This is not a data-quality nuisance to be cleaned up after the pilot proves value. It is the reason the pilot cannot prove value.
Beneath all three patterns is one distinction that separates water from nearly every industry buying AI right now. In most sectors, a wrong AI answer is a bad user experience. In a regulated utility, an AI-assisted answer becomes part of the operating record — and the eventual audience for that record is a primacy agency, a commission, a court. The question that matters is not whether the tool helped today. It is whether the answer can be defended in five years, by people who weren’t in the room, after the person who knew why has retired.
“Useful” is the wrong bar. Every vendor clears it in a demo. The bar is defensible.
Two clarifications, because this argument is easy to misread. This is not a case against AI capability in water — leak detection that works is worth having, and a good predictive model earns its keep. It is a case about what has to be true underneath the capability for it to hold up in a regulated environment. And it is not a pitch for compliance software. Compliance is not the product; it is what you get for free when doing the work and recording the evidence are the same act.
What we believe
1. The binding constraint is context, not models. A utility’s operating knowledge is split across five systems of record and some number of retiring heads, and no model — however capable — can reason with context it cannot reach. We have made the general version of this argument elsewhere; in water, the sector’s own pilot graveyard makes it empirically. Better models read the wrong context more eloquently. The work is connecting the context — and connecting it does not mean another migration. The instinct this sector has learned to fear, rip out the systems and hope the replacement holds, is the wrong instinct here, and so is its modern cousin, the multi-year data-lake program. The systems of record stay where they are and stay authoritative. What is missing is the layer of meaning above them: which asset in the maintenance system is which feature in GIS is which tag in the historian, which procedure governs it, and why the crew does it the way they do. Facts can stay where they live. Meaning needs a home.
2. Defensibility is a design requirement, not a feature. Every answer an AI produces inside a utility should carry its evidence: which record, from which system, as of when, touched by whom. Not as an export generated for the audit, but as the way the answer exists in the first place. And there is an inversion to demand here: the record that satisfies the auditor should be produced by doing the work — who did what, on which asset, under which procedure, with what result — not reconstructed from fragments years later. When execution itself generates the evidence, compliance stops being a season and becomes a by-product. An answer that cannot survive a records request in 2031 should not be informing a decision in 2026.
3. Institutional knowledge is an asset class — manage it like one. Utilities already run the best asset management discipline in the public sector: inventory, condition assessment, criticality, renewal before failure. The knowledge in a senior operator’s head deserves exactly that discipline, and almost nowhere receives it. It is uninventoried, its condition is unassessed, and its failure mode — retirement — is scheduled and ignored. Capture cannot be an exit-interview binder. It has to happen at the moment of work: the closeout note, the inspection, the judgment call and the reason for it, recorded where the next person — human or machine — will find it attached to the asset it concerns.
4. Autonomy is earned, not configured. The right posture for AI in a utility is proposal, not action. An agent notices stock running short against the maintenance schedule and drafts the requisition. It assembles the monitoring report from the operational record and queues it for the certifier’s signature. It proposes next quarter’s preventive maintenance plan with the reasoning attached. In every case, an accountable operator approves, and every step is logged. Autonomy can expand as reliability is demonstrated — but it expands inside the audit trail, never ahead of it. And some boundaries do not belong on the ladder at all. There are things an AI system in a water utility should be structurally incapable of doing — touching process control is the first and clearest. Not configured not to. Not instructed not to. Incapable. A utility should be able to show its commission not just what its AI did, but who let it.
Five questions to ask before you buy
Any AI investment a utility considers — including anything we would ever put in front of one — should clear these:
- Where did this answer come from? Ask to see the chain: answer to record, record to system, system to timestamp. If the vendor shows you an explanation instead of evidence, you have your answer.
- Can this be reconstructed in five years? After the product has shipped three versions and your staff has turned over — is the basis for today’s answer still reachable?
- Does it capture our knowledge or just consume our documents? What does it learn from a closed work order? From a correction an operator makes? If nothing, it will know as little in year three as on day one.
- Does it work across our systems or beside them? If it cannot hold one picture across the CMMS, GIS, SCADA, and the lab, it is another partial view — you have enough of those.
- Who approved what it did? Ask to see the log you would hand your commission — and ask what the system is incapable of doing, not merely configured not to.
A vendor that stumbles on the first question has answered the other four.
As for where to begin: not with the flashiest prediction. Start where work is recurring, deadline-bearing, and cross-system — preventive maintenance, stock replenishment, the reporting calendar — because that is where fragmentation costs the most and where the evidence obligations already exist. A foundation that can carry that work can carry the rest; one that cannot, cannot.
The stake
By the end of the decade, North American utilities will have sorted into two postures.
Some will have bought capabilities. They will own more tools, more silos, and a growing explainability debt — coming due precisely as lead service line reporting matures, monitoring and reporting obligations compound, and commissions start asking how the algorithms work.
Some will have built the foundation: their systems connected into one governed picture, their retiring operators’ knowledge captured where work happens, every AI answer born with its evidence attached. For them the curve bends the other way. Every work order enriches the record. Every retirement costs a little less. Every audit becomes a query instead of an archaeology dig. And for operators running more than one site, the compounding is steeper still: a lesson confirmed at one plant — a substitution that works, a failure mode caught early — becomes fleet knowledge instead of dying in a site silo.
In New York, the work was done and the record never caught up. The next decade of AI in water is about making those the same act — so that doing the work is what proves it.
Leave a Reply