006 Work

Aster · Internal business system

Updates that no longer throw everyone out

The reason this company updated its software rarely was not caution. It was that every update signed every member of staff out, in the middle of whatever they were doing. So updates waited for the evening, and then for a quiet week, and then for a quarter.

0FX At a glance

0People signed out by a release
2Copies serving at all times
1Automatic undo, no human
Found
Sign-in state stored inside the software's own machine
Consequence
Releases evenings only; then rarely; then in batches
Fix
Sign-in state moved outside, shared by every copy
Fix
Two copies serving; a release replaces them one at a time
Safety
A bad release rolls itself back without anyone deciding
Also
Point-in-time database recovery for 35 days

Ask a business why it has not updated its software in eight months and you will usually be told something about risk. Push a little and the real answer turns up, and it is almost never about risk in the abstract. It is that updating is unpleasant in a specific, concrete way, and everybody has quietly agreed to do it as seldom as possible.

Here the unpleasantness was precise: releasing a change signed every user out.

Why that happened

When you sign in to a system, it has to remember that it is you on the next page you open. It keeps a small note to that effect. Where it keeps that note is a detail nobody thinks about until it matters.

This software kept the note on the machine it was running on. That works perfectly, right up to the moment you replace the machine — which is exactly what installing an update does. New machine, no notes, everybody is a stranger again.

In practice that meant: a release at 3pm threw twenty-odd people out of the system, each of them losing whatever half-finished form was open at the time. So releases moved to the evening, which meant somebody had to be there in the evening. Evening releases are tiring, so they became fortnightly, then monthly, then whenever there was something urgent enough to justify the evening.

The part that costs money

That schedule has a cost, and it is not the evenings.

Changes accumulate. A release that carries one small fix is easy to reason about and easy to undo. A release that has been waiting three months carries forty changes from four different pieces of work, and when something goes wrong afterwards nobody knows which of the forty did it. So the response to a bad release is a long, expensive investigation rather than a two-minute reversal.

And because that is true, the fear becomes rational. Everyone knows that releasing is dangerous, so they release less, which makes each release more dangerous. The organization ends up describing itself as careful when what it actually is, is stuck.

Meanwhile the things that should be routine — a security update, a small fix somebody asked for in March — sit in a queue behind the fear.

What changed

Three pieces, and none of them touched the features.

The note that says who is signed in moved out of the software’s own machine and into a small shared service that all copies of the software can read. Now the software can be replaced underneath a signed-in user without them noticing, because the thing that remembers them is not the thing being replaced.

Two copies of the software run at all times. A release does not stop and restart the system; it starts new copies alongside the old ones, waits until the new ones are answering properly, moves traffic across, and only then retires the old ones. If the new copies never answer properly, traffic never moves, and the old ones keep serving. That is the difference between a bad release and an outage.

And the release watches itself. If the new version fails its checks, the system puts the previous version back on its own, without waiting for somebody to notice, log in and decide. Underneath both, the database can be restored to any point in the previous 35 days, which is the backstop for the class of problem that no deployment mechanism can catch — a change that works perfectly and does the wrong thing to your data.

What it is like now

Releases happen in the middle of the working day. Nobody is signed out. Nobody is told in advance, because there is nothing to tell them.

The change nobody expects is what that does to everything else. When shipping a small fix is a ten-minute non-event, small fixes get shipped. The queue of “we should probably do that at some point” stops being a queue and turns back into ordinary work, and the software starts moving again after years of standing still.

What we would do differently

There is no separate rehearsal environment here, and that was a deliberate decision rather than an omission: a permanent copy of production costs real money every month and, in our experience, drifts far enough from the real thing that it stops answering the question you built it to answer. What stands in its place is a full local copy any developer can run, the automatic rollback, and the 35-day recovery window.

That trade has a sharp edge and it should be said out loud: if the local copy stops being easy to run, nothing else is checking releases before they ship. Keeping it working is not housekeeping, it is the safety net. Somebody has to own that, and it should be written down as a job rather than assumed.

If this sounds like your system

Ask the person who releases your software two questions.

What happens to somebody who is signed in and typing when we release? If the answer is “they get logged out”, you have this page.

If a release goes wrong at 4pm on a Friday, what puts it back, and how long does that take? If the answer involves a person, a decision and a login, you are one bad afternoon away from finding out how long it really takes.

What this kind of work costs is published here, with everything else.

0SY The symptom this fixes

They are not lazy and they are not bad at the job. They are one person holding up something that needs two kinds of work, and only one of them is visible to you.

0RL Related work

Other work worth reading

Does this sound like your system?

Describe what breaks in your own words. A person reads it and replies within one business day — what we think is happening and whether we’re the right people for it. Free.

Write to us hello@yourcodecare.com