001 Work

Aster · Internal business system

The server that could not be updated

One machine ran everything the staff used. Its operating system had stopped receiving security fixes three years earlier, and no amount of maintenance could change that. Here is what that costs a business, and how you get off it without rewriting the software.

0FX At a glance

3 yrsSince the last security fix
45,878Failed login attempts
0Lines rewritten
Found
Operating system past end of life; no fixes exist
Found
Never restarted in almost three years
Found
Main disk 96% full, weeks from stopping
Approach
Move the software as it is; change it later
Result
Hosting that receives updates, and a way to apply them
Old machine
Verified empty, then powered off and deleted

Software runs on a machine. That machine runs an operating system, and whoever makes that operating system supports it for a set number of years. During those years, when somebody discovers a way to break in, a fix is published and you install it. After the support date passes, people keep finding ways in and keep publishing them, and no fix is ever written for your version again.

This machine had passed that date three years before we saw it.

What that actually means

Nothing visible. The business worked. Work was booked, records were kept, staff logged in every morning, and nothing on their screens suggested a problem. That is the difficulty with this situation: it produces no symptoms until it produces a catastrophe.

Underneath, three separate things were true.

The operating system had received no security fix in three years, and hundreds of the components installed on it were behind the versions they should have been on. There is no path to fix that. You cannot update an operating system that is no longer being updated. The only real move is onto one that is.

The machine had been running continuously for almost three years without a restart. That gets presented as a point of pride. It means the opposite: no update that requires a restart had ever been applied.

And the front door was open. The administrative login for the machine itself was reachable from anywhere on the internet, protected by a password and nothing else. The log recorded 45,878 failed attempts. None of those were a person sitting at a keyboard; they are automated programs that sweep the internet all day looking for exactly this. Almost all of them move on. It takes one that does not.

The disk was also nearly full

A separate problem, and the more urgent one, because it had a date attached.

The main disk was 96% full. Most of that was a single file: an error log 48 gigabytes deep, written over and over by a feature that had been quietly failing for years. Nobody was reading the log, so nobody knew the feature was broken, and nobody knew the log was eating the disk.

When a disk fills completely the software stops. Not slowly, and not with a warning — it simply cannot write anything more, and everything that depends on writing fails at once. On the day we looked, that was weeks away.

We archived what was worth keeping, cleared what was not, and recovered 54 gigabytes. Then we added the thing that should have been there from the beginning: a rule that stops any log from growing without limit again.

Moving without rewriting

The instinct when you find important software on a dying machine is to rebuild the software. That is the most expensive option available, it takes a year or more, and it does not address the thing that is actually dangerous. The dangerous thing is the machine.

So the software moved as it was.

We packaged the application exactly as it ran — same code, same behaviour, same screens — and put it on hosting that is maintained by people whose whole business is maintaining it. Nothing about how the system works changed. Staff saw the same screens on the Monday as they had on the Friday, which is the point.

That is not the end state anyone would design from scratch. The software itself is still old, and modernizing it is separate work with its own cost and its own schedule. What the move does is untangle two problems that had been stuck together. The machine cannot be secured is now solved. The software is old is now a decision you take on a calendar, in your own time, instead of under pressure.

Making sure the old machine was really empty

The last step is the one that usually gets skipped, and it is what turns a migration into a finished migration.

Before the old machine was switched off, we went through it looking for anything that existed only there: files nobody had mentioned, a database running quietly in the background, scheduled jobs that nothing else was doing. It turned out the machine had also been serving a second, older website that everybody had forgotten was running on it.

Only when every item was accounted for was the machine powered off, and then deleted. There is now no forgotten box with company data on it sitting in a rack somewhere. That is its own quiet risk, and it survives most migrations.

What we would do differently

We cleaned the disk before we fully understood what was filling it. The 48-gigabyte log turned out to be a dead feature failing on repeat, and knowing that first would have made the cleanup a five-minute decision rather than a careful afternoon. Read the thing that is growing before you decide what to do with it.

We would also have asked earlier what else lived on that machine. Finding a second live website three weeks in was fine. Finding it the day before the move would not have been.

If this sounds like your system

Ask whoever maintains your software one question: is the operating system on our server still receiving security updates, and when does that stop?

Three answers are possible. A date in the future is the one you want. A date in the past is this page. “I don’t know” is the most common, and finding out takes about ten minutes.

Your hosting company can answer it in a support ticket, and the same rough version information is visible to any stranger who looks. It costs nothing to find out which of the three you are.

0SY The symptom this fixes

The fear is rational. In a system with no backups you trust and no way to test, every change really is a gamble. That is fixable, and it is the cheapest work there is.

0RL Related work

Other work worth reading

Does this sound like your system?

Describe what breaks in your own words. A person reads it and replies within one business day — what we think is happening and whether we’re the right people for it. Free.

Write to us hello@yourcodecare.com