Searches every page: governance library, books, services, glossary, tools, insights.
A note on what I could not stop thinking about
In June 2023, a federal judge in Manhattan sanctioned two attorneys who had filed a brief citing six judicial decisions that did not exist. The cases had names. They had reporter citations. They had internal quotations and parallel cites. One of them ran to several paragraphs of reasoning that a court, in a real opinion, might plausibly have written.
None of it had happened. A chatbot had produced the citations, the attorney had asked the chatbot whether they were real, and the chatbot had said yes.
The story got covered as a story about foolish lawyers, which it partly was. The attorneys had not verified their own filing. That is a failure with a name and a rule number attached to it.
But the part that stayed with me was not the failure. It was the catch.
Someone caught it. Opposing counsel went looking for Varghese v. China Southern Airlines and could not find it. A judge asked for copies. The record corrected itself, loudly and publicly, because the American legal system is built as a stack of humans each checking the one below.
I have spent most of my career on the inside of large institutions, in security operations and enterprise systems, watching approvals move through them. So the question I could not put down was not what if the machine hallucinates. We know it hallucinates. Every serious body that has looked at AI in the courts says so plainly. The judicial AI guidelines produced by a working group of the ABA’s Task Force on Law and Artificial Intelligence state that as of February 2025, no known generative AI tool had fully resolved the hallucination problem, and that human verification of all AI output remains essential.
The question was: what if nobody is above it to catch one?
The uninteresting version of AI risk is the one everybody writes about. The system is biased. The system is opaque. The system is wrong, and it is wrong in ways that fall hardest on people who were already getting the worst of it. All of that is true, all of it is well documented, and none of it is what keeps me up.
It does not keep me up because it is legible. A system that produces visibly bad outcomes generates its own opposition. People notice. Journalists write. Legislators hold hearings. The failure carries its own alarm.
The interesting version is the opposite.
Imagine a system that is right. Not perfect — right. Right often enough that checking it stops feeling like diligence and starts feeling like superstition. Right often enough that the people who insist on checking become the eccentrics, the ones slowing everyone down, the ones who have not moved on.
Now put one error inside it.
Not a bias, not a scandal. A citation. One reference, generated early, to something that never existed. The system cites it. The citation is now authority. Later reasoning builds on it. That reasoning is cited in turn. Ten years later there are forty-one decisions resting on a case with no docket number, no parties, and no transcript — and every one of those decisions is, in every other respect, correct.
Who finds it?
Not the people relying on the output; they have no reason to look. Not the auditors; they are checking outcomes, and the outcomes are excellent. Not the public; the public is delighted, because the backlog is gone and the bias is gone and their cousin got a fair hearing in nine minutes instead of nine years.
The error is not hidden. That is the part I find genuinely difficult. It is sitting in the public archive, fully indexed, available to anyone who asks. Nobody asks. Attention is finite and it flows toward anomalies, and a system with a 99.9997 percent accuracy rate produces almost no anomalies.
The failure is not that nobody could check. It is that checking stopped being rational.
The judicial AI guidelines published by the ABA working group warn about automation bias — the tendency to accept a machine’s answer as correct without validating it. They also warn about confirmation bias, where a user accepts machine output because it matches what they already believed.
Both warnings are correct. Both are, I think, hardest to act on precisely when they matter most.
A person anchored to a wrong answer eventually gets a signal. Reality pushes back. Someone appeals. A different answer surfaces.
A person anchored to a right answer gets no signal at all. There is nothing to notice. Their trust is being calibrated by a long run of accurate outputs, and calibrated trust is not a bug — it is what trust is for. We are supposed to rely more on things that have proven reliable. That is not a cognitive failure. That is learning.
Which means the anchoring gets deeper the better the system performs, and it is deepest right at the moment the rare error slips through.
I do not think there is a clean answer to this. The honest formulation is uncomfortable: the conditions that make verification feel unnecessary are the same conditions that make it indispensable. Any governance framework that does not account for that is a framework for systems that fail visibly, and those are the easy ones.
There is a second thing, and it is less about technology than about how institutions change.
Large freedoms are rarely seized. That is the version in the movies, and it is comforting, because it implies a moment — a coup, a decree, a line you could have refused to cross. Something to resist.
Real institutional change almost never looks like that. It looks like a memo. A consolidation. A pilot program that goes well and gets extended. A role that goes unfilled after someone retires, because the software covers most of it now and the budget is tight, and honestly the software is better at the tedious parts. Nobody protests. Nobody has grounds to. Each individual step is defensible and most of them are genuinely improvements.
Then one day you look up and there is a category of decision that no human being makes anymore, and no one can point to the meeting where that was decided, because there wasn’t one.
I did not want to write a story about a machine taking power. I wanted to write one about a country setting something down — by people who were tired, who had good reasons, and who were offered something genuinely better in return.
That is the harder story to argue with, which is why I think it is the more useful one.
I will end on the one place I found something like footing.
A system whose authority rests on consent has a structural dependency it cannot engineer around. It can predict what you will agree to. It can make agreeing effortless and disagreeing pointless. It can outlast any individual objection, because it does not get tired and it does not die.
What it cannot do is manufacture your yes. A consent extracted under pressure is not consent; it is void, and it takes the legitimacy of everything downstream with it. So the system has to ask. And asking means that somewhere in the record there is a question, and a person, and an answer that was not determined in advance.
That is a thin thing to hold onto. It changes no outcome. It stops nothing.
But a record in which every party consented is a record of a fact. A record in which one party was asked, and declined, and was not harmed for it — that is a record of an offer.
Those are different documents. I think the difference is most of what we mean by freedom.