Join the waitlist

essay2026-07 · 6 min ·

Agents don't only fail in the model

A hallucination in front of a client is a model failure. A team quietly switching the automation back off is not — and it's usually the more expensive one.

The failure nobody post-mortems

When an agent hallucinates in production, you get an incident ticket, a post-mortem, maybe a better eval. The machinery of engineering kicks in, because engineering has a vocabulary for model failures.

But watch what happens when a support team stops trusting the triage agent after two bad answers, and quietly routes around it. No ticket. No post-mortem. The dashboards still look fine — the agent answers fast and correctly. It just doesn't matter anymore, because nobody's listening to it.

An organization can have flawless pipelines and still stall. The gap is usually human — and unlike a hallucination, it doesn't file an alert.

"People stuff" is measurable — if you insist on it

The usual response to this observation is a workshop. Ours was to put the human side on the same map as the technical side, with the same evidence discipline. Three things from that zone are already deployed and measured in a working organization:

None of this is a survey about feelings. Each carries an evidence class, the same grading we apply to compilers and pipelines.

Why this matters for whoever runs the org

If you lead engineering: when agents stall despite sound pipelines, the human zone of the map names why — a post-mortem vocabulary for the failures that aren't in the model.

If you lead HR or L&D: this is evidence for the table you're already at — which skills to teach first, which roles change, where trust needs rebuilding. Measured before your learning program and re-measured after, it's learning impact you can defend beyond completion rates.

One read, both languages. That's the difference between an organizational discipline and an IT audit.

Want the measured version of this argument? Research · Want it applied to your organization? The diagnostic.