Application logs are the record of what your software did. When something breaks, the log is how anybody finds out why. You want plenty of it, and you want to keep it long enough to be useful.
That sits in direct tension with the other thing that is true: the data flowing through the software belongs to your customers, and a log is a copy of that data, sitting somewhere with weaker access control than the database it came from.
What we found
The software logged errors with context, which is the right instinct. Context means: when something failed, record enough about the situation to understand it later.
The context being recorded was the request. All of it. A request that fails while creating a record carries whatever the form carried — a person’s name, a date of birth, an identifier, and everything else somebody typed into the screen.
There was a second, quieter version of the same problem. Somebody debugging a stubborn issue had added direct writes to a file — not through the logging system, just an append to a path on disk. Those writes had shipped and stayed. They were not covered by log rotation, not covered by retention, and not visible to anybody looking at the logging settings, because as far as the settings were concerned they did not exist.
Neither of these is negligence. Both are what happens when the decision “what is safe to write down” is made hundreds of times, by different people, at the moment they are trying to fix something else.
Why “be careful what you log” does not work
Because it is advice, and advice fails silently.
The rule has to be remembered by every developer, in every error path, forever, including the person debugging at the end of a long day with a customer waiting. It gets followed almost always. Almost always is the problem: you cannot tell which of a million log lines is the one where it was not, and a single line with a person’s details in it is the same finding as a thousand.
The property you want is different in kind. You want it to be impossible to write a customer’s details into the log by accident, not merely discouraged.
What was built
One filter, sitting where every error already passes.
The software already routed failures through a single point — the thing that turns a failure into a response. That is the choke point. The filter was attached there, so context passes through it before it reaches the log, without any of the hundreds of places that log errors knowing anything about it.
It works by field name and by shape. Known-sensitive fields are replaced with a marker rather than deleted, so the log still shows that a field was present and what it was called, which is most of the diagnostic value. Anything that looks like a key or a token is removed by pattern regardless of what it is called, because the field a secret arrives in is not always predictable.
Then the direct-to-disk writes were removed and replaced with ordinary logging calls, so everything the software records now goes through the same path — and therefore through the same filter. That change is small and boring and it is the one that closes the hole permanently, because it removes the second road.
Sixty-four tests cover the filter. They matter more than the count suggests: a filter that silently stops working is worse than none, because the log looks clean and is not. The tests check the removal on each known field, check that unknown fields carrying key-shaped values are caught, and check that the surrounding diagnostic context survives — a filter that strips everything is also a failure, just a less dangerous one.
What it changed downstream
Two things became decidable that had not been before.
Retention could be set deliberately. You cannot answer “how long should we keep logs” while the honest description of the log is “we are not sure what is in it”. With the contents known, the retention period becomes an ordinary operational decision instead of an open question.
And the logs became usable by more people. A log that might contain personal data has to be treated as though it does, which means the people who most need it to debug a production problem are the people least likely to be allowed near it. That is a tax on every incident, paid in the minutes when minutes are expensive.
What we would do differently
We purged the historic logs rather than cleaning them. That was the pragmatic call — the volume was large, the format was inconsistent across years, and writing a reliable retrospective cleaner would have cost more than the logs were worth. It is still a loss: some of that history was genuinely useful for understanding old behaviour, and it is gone.
We also did not add anything that would catch a new direct-to-disk write appearing in future. A check in the build that fails on writes outside the logging system would take an afternoon and would make this permanent rather than current. It is not there.
If you run software that handles personal data
Two questions, in this order.
What ends up in our logs when a request fails? If the answer is “whatever was in the request”, you have this. It is extremely common and it is not a sign of a bad team.
Who can read those logs, and how long are they kept? People are usually surprised by both answers. Logs tend to be readable by more people than the database, and kept longer than anyone chose, because nobody ever decided — a default did.
If neither question has an answer today, our audit establishes both at a fixed price. Or write and tell us what you are running and we will say whether it is worth looking at.