Ask a business why it has not updated its software in eight months and you will usually be told something about risk. Push a little and the real answer turns up, and it is almost never about risk in the abstract. It is that updating is unpleasant in a specific, concrete way, and everybody has quietly agreed to do it as seldom as possible.
Here the unpleasantness was precise: releasing a change signed every user out.
Why that happened
When you sign in to a system, it has to remember that it is you on the next page you open. It keeps a small note to that effect. Where it keeps that note is a detail nobody thinks about until it matters.
This software kept the note on the machine it was running on. That works perfectly, right up to the moment you replace the machine — which is exactly what installing an update does. New machine, no notes, everybody is a stranger again.
In practice that meant: a release at 3pm threw twenty-odd people out of the system, each of them losing whatever half-finished form was open at the time. So releases moved to the evening, which meant somebody had to be there in the evening. Evening releases are tiring, so they became fortnightly, then monthly, then whenever there was something urgent enough to justify the evening.
The part that costs money
That schedule has a cost, and it is not the evenings.
Changes accumulate. A release that carries one small fix is easy to reason about and easy to undo. A release that has been waiting three months carries forty changes from four different pieces of work, and when something goes wrong afterwards nobody knows which of the forty did it. So the response to a bad release is a long, expensive investigation rather than a two-minute reversal.
And because that is true, the fear becomes rational. Everyone knows that releasing is dangerous, so they release less, which makes each release more dangerous. The organization ends up describing itself as careful when what it actually is, is stuck.
Meanwhile the things that should be routine — a security update, a small fix somebody asked for in March — sit in a queue behind the fear.
What changed
Three pieces, and none of them touched the features.
The note that says who is signed in moved out of the software’s own machine and into a small shared service that all copies of the software can read. Now the software can be replaced underneath a signed-in user without them noticing, because the thing that remembers them is not the thing being replaced.
Two copies of the software run at all times. A release does not stop and restart the system; it starts new copies alongside the old ones, waits until the new ones are answering properly, moves traffic across, and only then retires the old ones. If the new copies never answer properly, traffic never moves, and the old ones keep serving. That is the difference between a bad release and an outage.
And the release watches itself. If the new version fails its checks, the system puts the previous version back on its own, without waiting for somebody to notice, log in and decide. Underneath both, the database can be restored to any point in the previous 35 days, which is the backstop for the class of problem that no deployment mechanism can catch — a change that works perfectly and does the wrong thing to your data.
What it is like now
Releases happen in the middle of the working day. Nobody is signed out. Nobody is told in advance, because there is nothing to tell them.
The change nobody expects is what that does to everything else. When shipping a small fix is a ten-minute non-event, small fixes get shipped. The queue of “we should probably do that at some point” stops being a queue and turns back into ordinary work, and the software starts moving again after years of standing still.
What we would do differently
There is no separate rehearsal environment here, and that was a deliberate decision rather than an omission: a permanent copy of production costs real money every month and, in our experience, drifts far enough from the real thing that it stops answering the question you built it to answer. What stands in its place is a full local copy any developer can run, the automatic rollback, and the 35-day recovery window.
That trade has a sharp edge and it should be said out loud: if the local copy stops being easy to run, nothing else is checking releases before they ship. Keeping it working is not housekeeping, it is the safety net. Somebody has to own that, and it should be written down as a job rather than assumed.
If this sounds like your system
Ask the person who releases your software two questions.
What happens to somebody who is signed in and typing when we release? If the answer is “they get logged out”, you have this page.
If a release goes wrong at 4pm on a Friday, what puts it back, and how long does that take? If the answer involves a person, a decision and a login, you are one bad afternoon away from finding out how long it really takes.
What this kind of work costs is published here, with everything else.