002 Work

Aster · Internal business system

Six hundred gigabytes that existed in one place

Every file the business had ever received from a customer sat on a single machine, with no copy anywhere else. Nobody had decided that. It had simply never been anyone's job, which is how it usually happens.

0FX At a glance

640 GBWith no second copy
3Kinds of file at risk
36Restore points today
Found
Uploads, images and legal documents on one machine only
Found
An old copy routine failing into a destination long deleted
First step
Copy it off — which was also the first real backup
Method
Resumable copy, then a second pass for anything that changed
Today
Daily backup, restore points kept 98 days
Proof
A restore was actually performed, not assumed

Ask a business where its files are and you usually get an answer about the software. The files are in the system. The system is on the server. Everyone nods, and nobody has asked the next question, which is: and where else?

Here the answer was nowhere.

What was at stake

Three sets of files, about 640 gigabytes in total, all on the same machine.

Customer uploads — everything people had sent in over the years. A large library of processed images generated from those uploads, which the software recreates on demand but only if the originals still exist. And a third set that matters differently: documents, agreements and payment records. That last one is not just operationally important. It is the paperwork you would need if somebody disputed something two years from now.

There had been an attempt at protecting this. A routine ran on a schedule, copying images to storage somewhere else. When we traced it, the destination it was copying into had been deleted at some point in the past. The routine still ran. It still reported that it had run. It had been copying into nothing for a long time.

That is not unusual. A backup that nobody has ever restored from is a belief, not a backup, and the belief survives right up until the day it is tested for real.

What “gone” would have meant

If that machine had failed — a disk, a fire, a mistake at the hosting company — the position on the following morning would have been: the software can be reinstalled, the database can be recovered from its own separate backup, and every file any customer ever sent is gone permanently.

Not degraded. Not partially recoverable. Gone.

For a business whose work is producing and holding those files, that is not a technical incident. It is the end of the ability to do the work, plus a legal problem for every agreement in the third folder.

Copying it off, carefully

The first task was therefore not the migration. It was making a second copy exist at all, as early as possible, because until that was true everything else was being done on top of a single point of failure.

Two things make this less simple than it sounds.

The data is large enough that a copy takes the better part of a day, and the business does not stop while it runs. So the copy had to be the kind that can be interrupted and resumed, and that can be run a second time to pick up only what changed since the first pass. That is exactly what happened: one long copy early, then a short second pass at the end to catch everything added in between.

The other issue is that files and records refer to each other. The database does not store where a file is; it stores its name, and the software works out the location by a fixed convention. Break that convention during a move — put things in a slightly different arrangement — and every file becomes unreachable while still existing. So the arrangement was preserved exactly, and the software did not need to be told anything had changed. It could not tell.

What is true now

The files live on storage that is separate from any single machine, with a daily backup taken automatically and restore points kept for 98 days. There are 36 of them at the time of writing.

More usefully: a restore has actually been carried out. Not a check that the backup job reported success — an actual recovery of actual files, to prove that the path from “we have backups” to “the files are back” exists and that somebody has walked it. That is the difference between a backup and a recovery plan, and it is why the first is worth so much less than people think.

The storage also moved older files into a cheaper tier automatically, which took the monthly bill for it down by roughly half without anyone deciding which files matter.

What we would do differently

We should have asked to see a restore on day one, before anything else. The failing copy routine had been reporting success for years and nobody had reason to doubt it; ten minutes of “show me a file coming back” would have surfaced the real position immediately rather than in week two.

We also kept the full set rather than working out which files the software still refers to. It is very likely that some of that 640 gigabytes is no longer referenced by anything. Working out which is a slow, careful job with a real chance of deleting something that mattered, and storage is cheap enough that we did not do it. That remains true and it remains a small ongoing cost.

If this sounds like your system

Two questions, in this order.

Where is the second copy of our files, and when did somebody last bring one back from it? Not “do we have backups”. Anyone can say yes to that. The useful answer names a date and a person.

If the main machine were gone tomorrow morning, what would still exist? Most people are surprised by their own answer.

If neither has an answer today, write and tell us what you are running. The reply is free and it will tell you honestly whether this is worth spending money on.

0SY The symptom this fixes

You bought the company, or took over the department, and the software came with it. There is no documentation and no one left to ask.

0RL Related work

Other work worth reading

Does this sound like your system?

Describe what breaks in your own words. A person reads it and replies within one business day — what we think is happening and whether we’re the right people for it. Free.

Write to us hello@yourcodecare.com