essay2026-07 · 6 min ·
Agents don't only fail in the model
A hallucination in front of a client is a model failure. A team quietly switching the automation back off is not — and it's usually the more expensive one.
The failure nobody post-mortems
When an agent hallucinates in production, you get an incident ticket, a post-mortem, maybe a better eval. The machinery of engineering kicks in, because engineering has a vocabulary for model failures.
But watch what happens when a support team stops trusting the triage agent after two bad answers, and quietly routes around it. No ticket. No post-mortem. The dashboards still look fine — the agent answers fast and correctly. It just doesn't matter anymore, because nobody's listening to it.
An organization can have flawless pipelines and still stall. The gap is usually human — and unlike a hallucination, it doesn't file an alert.
"People stuff" is measurable — if you insist on it
The usual response to this observation is a workshop. Ours was to put the human side on the same map as the technical side, with the same evidence discipline. Three things from that zone are already deployed and measured in a working organization:
- An agent–human communication protocol — the rules of who says what to whom, running as an organization's working nervous system for months.
- Bidirectional accountability — agents give structured feedback to their humans, not just the other way around. In production since May.
- A two-minute trust score — any human+agent pair, scored on one formula: Trust = Alignment × Reliability.
None of this is a survey about feelings. Each carries an evidence class, the same grading we apply to compilers and pipelines.
Why this matters for whoever runs the org
If you lead engineering: when agents stall despite sound pipelines, the human zone of the map names why — a post-mortem vocabulary for the failures that aren't in the model.
If you lead HR or L&D: this is evidence for the table you're already at — which skills to teach first, which roles change, where trust needs rebuilding. Measured before your learning program and re-measured after, it's learning impact you can defend beyond completion rates.
One read, both languages. That's the difference between an organizational discipline and an IT audit.
Want the measured version of this argument? Research · Want it applied to your organization? The diagnostic.