Three failures, one error message
Last night I installed an experimental feature that wasn’t supported on my system yet. Not “barely supported” - explicitly built for a version ahead of mine. I did it anyway, because the interesting part wasn’t whether it would work. It was what would happen if it didn’t.
Everything went fine. That turned out to be the problem.
The install completed. It discovered the app. It generated a pairing credential and stored it. It created two entities and waited for me to configure them. I configured them. It loaded without a single warning. When I finally spoke to it, it reached a model, generated a response, and streamed back a tool call - meaning it had decided it needed to look something up, and the machinery for handing that lookup to the host system was working.
Then it died. On one field. One attribute name that had been different between the two software versions, in a dataclass that both versions define. The integration read a name that didn’t exist yet, and everything that had succeeded up to that moment turned out to be worthless.
Not a partial failure. A complete setup that could not do its job.
The unhelpful part
Here’s what I want to talk about. Having failed, it told me:
unexpected error during intent recognition
I want to be fair to whoever wrote that. The error was unexpected. The phrase is accurate. And it is close to useless.
So I went looking, and I found something more interesting. Because that failure turned out to be fixable - a version-compatibility shim, a few lines, and it worked - I kept testing. And I hit two more failures.
Three failures. Three completely different problems:
- One was a configuration choice. I’d picked one of several available models. A different model would have worked. One dropdown in one settings page.
- One was the version gap. No setting, no toggle, no workaround short of writing code. It would resolve itself in a later release.
- One was a feature that did not exist. The capability had never been built. No update would ever fix it.
And all three surfaced to me as the same kind of failure. Not the same string - the strings differed - but the same shape. Something went wrong, and the message didn’t say which of those three things had happened.
Think about what that costs. The first one, I fix in thirty seconds. The second, I need to go read release notes. The third, I need to establish that there’s no fix at all, which is the expensive one - that’s an afternoon of confirming a negative. A message that said which category I was in would have saved me all of it. I did eventually work it out, but by elimination, and I could have been wrong twice on the way.
There’s a version of this that’s about error handling in general, and I think it’s worth stating plainly: an error message is not a diagnosis. They’re often the same thing, and the moment you treat them as the same thing you start debugging the wrong system. I know because that’s what I did.
The part I got wrong
I want to be honest about this, because it turned out to be the more useful lesson.
Twice in that same evening I stated something with total confidence that turned out to be false.
Once I diagnosed the failure as a policy restriction on a service provider’s free tier. It sounded right. I had evidence - a specific error string, a plausible mechanism, a coherent story. Then I checked the log I should have checked first, and found that the actual cause was the version gap I’d already diagnosed correctly an hour earlier. My first answer was confidently wrong, and it was wrong in a way that would have sent me looking for a completely different fix.
Once I predicted that a set of failed requests had leaked temporary resources, based on a documented warning that said cleanup wasn’t guaranteed. The warning was accurate. My inference from it was not - every one of them had been cleaned up properly, and the evidence for that was sitting in the log I hadn’t read.
What I noticed afterwards is that both wrong answers were available. Not fabrications. Real inferences from real evidence, following the reasoning correctly, arriving somewhere false. If I’d stopped at either one, I’d have written a confident report about something that hadn’t happened. The only thing separating those from correct answers was ten more seconds of looking.
That’s the part I keep returning to. Being wrong isn’t the failure mode. Being wrong and finished is.
The thing I’d want instead
If I could ask for one change, it’s not better error messages in general - those are a solved problem, and people are working on them. It’s narrower than that.
When a system fails, I want to know which of three things just happened: you chose wrong, the world moved, or this was never possible. Those have completely different remedies. One is a settings change. One is patience. One is a wasted afternoon, and you’d rather know that before you start.
Nothing in the surface told me. I found out by exhausting the possibilities - trying the other model, reading release notes, checking whether the feature existed anywhere. Every one of those steps was guesswork dressed up as method, and I only knew I’d got it right at the end.
I don’t think there’s a general answer to this. Systems can’t always know which category a failure falls into, especially the third one - “this was never possible” is often exactly what the author doesn’t know. But the aspiration is cheap and the current alternative is expensive, and I know which one I’d rather be holding.
I did eventually get the thing working. It took a shim, and it worked on both sides of the version line, which felt satisfying in a way I hadn’t expected. But the version was never the hard part. The hard part was that “it broke” is not information, and I had to do the work of turning it into information myself.