Automatic updates between companies are how modern business software works. Your supplier ships an order, and their system sends a short message to yours saying so. Your system updates the order and moves on. Nobody is involved, which is the point.
Those messages travel over the open internet. Anyone who learns the address can send one.
That is why they are signed. The two sides share a secret; the sender stamps each message with it, the receiver recomputes the stamp and compares. A message with the wrong stamp is thrown away. It is the same idea as a signature on a letter, and it is the only thing standing between “our supplier told us” and “somebody told us”.
What we found
This platform accepted messages from six outside systems.
The code that checked them was a list — one entry per sender, each saying how to verify that sender’s messages. At the end of the list was a fallback for anything not named. The fallback answered “valid”.
Two senders were not on the list. Their messages were accepted on arrival, unverified, because the fallback said yes.
This is not a strange kind of mistake. Somebody writes the check for the first partner, adds a permissive fallback so nothing breaks while the others are wired up, and moves on. The fallback is meant to be temporary. Nothing in the code says so, no test fails while it stays, and the platform works perfectly either way. It survives every code review after that, because reviewers read the change in front of them and the fallback is not in the change.
What it was worth to somebody
Not what people assume, and the difference matters.
Nobody was going to extract records this way. These addresses do not return data — they accept it. What a stranger could do was write.
An unsigned status update lets somebody mark orders shipped that were never shipped. An unsigned fulfilment message lets them attach the wrong tracking details to real orders. On a platform where those events feed what the customer sees and what the business reports on, the damage is quiet, and it is expensive to unpick — by the time anybody notices, the false events are mixed in with thousands of true ones and nothing distinguishes them.
There was no evidence any of this had happened. There was also no way to prove it had not, which is a finding in its own right, and we said so rather than rounding it to “no evidence of compromise”.
How it was fixed without an outage
The dangerous way to fix this is to flip the fallback to “reject” and release. One line, obviously correct, and it takes the platform down.
It takes it down because you do not yet know that every real sender is signing correctly. Partner documentation goes stale. One might sign a slightly different version of the message. Another might have been sending unsigned messages for a year precisely because that fallback allowed it. Flip the switch and their messages start bouncing, silently, and orders stop moving.
So it went in three stages.
Check, but do not enforce. Verification was written for both missing senders and switched on in a mode that computed the answer, recorded it, and accepted the message either way. That produced the fact nobody had: how many real messages would have been rejected.
Read what was recorded. Every genuine sender was signing correctly, and the count of would-be rejections was zero across the observation window. That is the evidence you need, and there is no other way to get it.
Then refuse by default. The fallback was inverted. An unnamed sender is now rejected, so the next partner integration cannot be quietly trusted the way these two were — it fails loudly on the first message and somebody wires it up properly.
What the tests are for
Twenty-eight tests went in with the change, and they are worth describing, because “we added tests” means nothing on its own.
They cover each sender separately, because each one signs differently. They cover a message altered after signing — the signature must fail. And they cover a replayed message: a genuine, correctly signed message captured and sent again an hour later, which must be rejected on its timestamp, because a valid signature on a stale message is still an attack.
That last one is the test people skip and the one that matters most. A signature proves who sent a message. It does not prove when, and it does not prove the message was sent once.
What we would do differently
We fixed the fallback where it stood. The better shape is one that cannot have a fallback at all — a register where adding a partner means registering how to verify them, and an unregistered sender is not a decision the code makes but a state that cannot exist. That is a larger change than this piece of work was scoped for, and we did not make it. The current shape depends on a default that a future developer could flip back in a hurry.
The other thing: we could not tell anyone whether anything had been forged before the fix, because the request logs were kept for less time than the gap had existed. If you keep thirty days of logs, understand what that buys you. It answers “is this happening now”. It never answers “did this ever happen”.
If this sounds like your system
The question to ask whoever maintains your software is short: for each system that sends us data automatically, how do we prove the message came from them?
There are three honest answers. “It is signed and we check the signature” is the one you want. “It comes from an address we allow” is weaker but real. “We don’t check” is common, and it is worth knowing rather than assuming.
If you would rather establish this yourself first, ask for the list of systems that send you data automatically and, next to each, how the message is proved to have come from them. The list is short and the answers are one line each.