Most modern websites are built on tools that do a great deal for you automatically. You describe the shape of the thing you want to collect — a name, a date, a phone number — and the tool builds the storage, the form handling and the address the form posts to.
It also, by default, builds the other half: a way to read what has been collected. That is a sensible default when you are building the private admin screen for your own staff. It is not a sensible default when the thing sitting in front of it is a public form.
Nobody decides this. It arrives switched on.
What was reachable
The form collected applications. Each record held a name, a date of birth, a photograph, and the phone number and email address of the person submitting.
The same address that accepted a submission would answer a request to list the submissions. There were 3,788 of them. A person who understood how these tools are built could have retrieved the lot, in order, in a couple of minutes, without a password, without an account, and without anything in the logs that looked unusual — because from the software’s point of view, nothing unusual had happened. It had been asked a question it was configured to answer.
This is one of the small number of findings where the harm does not need explaining to anybody. Names, dates of birth and contact numbers, held by an organization that collected them for one purpose, available in bulk to anyone. That is the kind of thing that ends up in a regulator’s inbox, and the organization’s honest defence — we never turned that on — is true and is not a defence.
How it was closed
There were two ways to fix it.
The quick one is to go into the tool’s settings and untick the permission that allows the public to read. That takes a minute, and it works, and we did not do it.
We did not do it because a permission that can be unticked can be re-ticked. Not maliciously — by somebody a year from now who is debugging a different problem, sees a permission that looks like it belongs to a working feature, switches it on to test something, and gets distracted. There would be no error, no alert, and nothing to notice. The setting would simply be back on.
So the reading behaviour was removed from the software itself. The address still accepts a submission, because that is the form’s job. There is now no code anywhere that answers a request to list or read them back, and a settings screen cannot conjure code into existence. The failure mode is now a plain “not found” rather than a decision somebody could reverse.
The old, now-meaningless permission was left in place on purpose, with a comment next to the code explaining why. Someone will find that permission one day and think it is a loose end. It is a second lock: even if they tick it, there is nothing behind it.
The other door in the same wall
While we were there, one related default got the same treatment.
The site also accepted file uploads from the public — necessary, because the form takes a photograph. As shipped, it accepted any file type, at up to fifty megabytes. That combination is how a public form becomes free file hosting for strangers, and occasionally how it becomes something worse.
Uploads are now restricted to actual images, and the size limit was brought down to something a photograph fits inside comfortably and a stolen film does not.
What we would do differently
We found this by reading the code and then checking it against what the live site actually answered. That is the right method, and it took time.
The faster method would have been to write down, at the start, every address the public can reach and what each one does when you ask it the six most obvious questions. That list is a day of work, it is useful forever, and it is the thing that turns “we think it is fine” into “here is what is exposed, on one page”. We produced it eventually. It should have been the first deliverable.
We also want to be exact about one thing, because it is the part people get wrong when they read a page like this and go looking at their own systems. Not everything a public form touches should be locked down. Elsewhere on the same site there is a public directory that the business exists to publish, and an earlier reading of this finding suggested closing that too. It would have broken the live site and removed the reason the site exists. Knowing the difference between this data was never meant to be readable and this data is the product is the actual skill, and getting it wrong in the cautious direction is still getting it wrong.
If this sounds like your system
Ask whoever maintains your website: for every form on our site, what happens if somebody asks the form’s own address to list what it has collected?
The answer should be “nothing, that does not exist”. If the answer is “the public role doesn’t have permission”, ask the follow-up: what stops that permission changing?
If you would rather establish this yourself first, ask whoever built the site to show you what each form’s own address returns when it is asked to list what it has collected. It is a five-minute demonstration, and watching it is more convincing than being told.