010 Work

Redwood · Customer-facing platform

The feature that had been broken for years

Users do not report broken things forever. They try twice, find another way to get their work done, and stop counting the feature as part of the system. After that, a support queue with nothing in it means nothing at all.

0FX At a glance

503Returned, every time
0Bug reports
2Kinds of broken
Found
A crash waiting to happen: a name that resolved to nothing
Found
A feature answering with an error on every call
Cause
Renames during an old restructure, references left behind
Why silent
Users stopped trying long before anyone asked why
Fix
References corrected where the feature was still wanted
Durable fix
The build now fails on a name that does not resolve

In a lot of business software, a reference to something that does not exist is not caught when the software is built. It is caught when that particular line runs — which may be rarely, or only on one branch of one condition, or in a part of the system most people never open.

That is how a system ends up containing calls to parts of itself that were deleted or renamed years earlier. Nothing objects. The software starts. Everything looks fine.

What we found

Two flavours, and they are worth separating because they fail differently.

The first was a straightforward crash waiting to happen: code asking for something by a name that resolved to nothing. If that line ran, the request died. It had presumably not run often, or it had and the resulting errors were among the many nobody was reading.

The second is the more interesting one. A feature checked whether a particular component was available and, if it was not, returned a “service unavailable” error instead of failing messily. That check was written defensively, which is good practice. The component it checked for had been renamed. So the check said “not available”, and the feature returned an error, and it did that on every call, for years.

Nobody reported it.

Sit with that for a moment

A feature was completely broken, in production, on a platform people use for their jobs, and there were no bug reports.

Not because users are forgiving. Because they stop. Somebody tries a thing, it does not work, they find another way to do their job, and the second attempt is a shrug rather than a support ticket. After a while the feature is not broken from their point of view — it simply is not something the system does. New staff are trained by people who already know not to bother with it, and within two years its absence is folklore rather than a fault.

Which means an empty support queue is not evidence of a working system. It is evidence about how your users feel about reporting things.

Both problems came from the same cause: a restructure some years earlier that moved and renamed parts of the system, leaving a small number of references behind. That restructure was not careless. It was large, it was mostly right, and this is the residue that a large mostly-right change leaves behind in software that will not tell you.

What was done about it

The references were traced to what they were supposed to point at and corrected where the feature was still wanted.

Where it was not — and one of them was not, the surrounding feature had been superseded and nobody had asked for it in years — the code was deleted rather than repaired. Fixing a path nobody wants to be able to reach is not a favour to anyone: it converts a loud failure into a quiet, working, untested path that somebody will eventually stumble into.

Then the more useful change: the project’s automatic checks were configured to fail the build when a name does not resolve, so the next one of these is a failed build rather than a discovery in three years.

That last part is worth more than the fixes. The individual bugs were an afternoon. The property that this class of bug can no longer reach production is permanent.

The uncomfortable finding underneath

The reason this survived is that nobody had a list of what the system claims to do, and a way to check that it still does it.

There were no tests over the affected paths. There was no monitoring that would notice a feature returning the same error on every single call — which is, incidentally, the easiest possible thing to alert on, because a feature at 100% failure is not a subtle signal. And there was no route by which a user’s private decision to stop using something became information that anybody acted on.

Any one of the three would have caught this years earlier. That is the actual finding, and it is bigger than the bugs.

What we would do differently

We fixed the references and raised the automatic checks. We did not go looking for the other features that might be in the same state — quietly non-functional, with users who stopped asking.

The way to find those is not in the code. It is a list of what the system is supposed to do, walked by a person, one item at a time, clicking each one. It is unglamorous and it is the only method that reliably finds this class of problem. We should have proposed it and we did not.

If you have inherited a system

The question is not “are there bugs”. The question is: is there anything in here that stopped working and nobody told us?

Finding out costs nothing but time. Take the list of things the system is supposed to do and have somebody try each one. On any system that has been running for more than a few years and has changed hands, this exercise finds something.

If that list does not exist either, that is the first deliverable. Tell us what you inherited — the reply is free, and it will tell you honestly whether a full read of the system is worth your money.

0SY The symptom this fixes

You bought the company, or took over the department, and the software came with it. There is no documentation and no one left to ask.

0RL Related work

Other work worth reading

Does this sound like your system?

Describe what breaks in your own words. A person reads it and replies within one business day — what we think is happening and whether we’re the right people for it. Free.

Write to us hello@yourcodecare.com