There is a category of problem that does not fit the word “security”, because nothing is being broken into. Everything the visitor asks for is something they are allowed to have. The problem is the rate, and the volume, and what somebody does with it afterwards.
This business publishes a directory. Publishing it is the point — it is what the site is for, it is why people visit, and locking it behind a login would be the same as switching the business off. So the question was never “how do we stop people reading this”. It was “how do we stop one program taking all of it, every week, while leaving the several thousand people who read it normally completely alone”.
How easy it was
Pages of a directory are usually served in batches: twenty at a time, with a next button. That limit was in place and worked.
What had no limit was a second, quieter path. The same request could be asked to include related information alongside each entry, and there was no ceiling on how much of that it would return. Asked the right way, one request came back with 1,167 distinct records in it, in a single response of about half a megabyte.
At that rate the entire directory — a little over eighteen thousand entries — comes out in roughly sixteen requests. Sixteen. That is not a scraping campaign, it is a coffee break.
And it was being taken. Not theoretically: the traffic logs showed a recurring pattern, from changing addresses, returning on a schedule, doing precisely this.
Why it is worth money to somebody
Every business with a public directory eventually meets this, and the reason is always one of three.
Someone is building a competing directory and would rather have yours than compile one. Someone is harvesting contact details to sell. Or someone is training something on it. In all three cases the value to them is the completeness — a hundred entries are worthless, all of them are a product — and the cost lands on you as bandwidth, as slower pages for real visitors, and as the odd position of having your own catalogue republished somewhere you did not choose.
There is also the plain operational cost. Those oversized requests were among the heaviest the system served. Removing them made the site faster for everybody, which was not the goal but was a real result.
What was done
Three things, in an order that mattered.
First, a guard was put in front of every public request that bounds what one request may return: which related information can be included, how many entries, how deep. The important design choice inside it is that it trims rather than refuses. A request asking for something unreasonable gets a reasonable answer, not an error. That distinction is what makes it safe to deploy on a live site — a guard that rejects unexpected requests is one unforeseen parameter away from breaking a page for real visitors on a Saturday.
Second, the edge — the layer in front of the site that sees traffic before the site does — was given rules for the pattern: rate limits and known bad behaviours. These were run in counting mode for three days first, recording what they would have blocked without blocking anything. Then the matches were read, one by one. They were attacks, without exception, so the rules were switched to blocking.
Third, and this is the part that makes it defensible: before the guard went live, every kind of request the real site makes — sixteen distinct shapes, the ones behind actual pages people load — was captured. After the guard went live, all sixteen were replayed. Same responses, same counts, same totals, same records, and no errors. The bulk path had gone from 1,167 records per request to 24. Taking the whole directory went from sixteen requests to something closer to eight hundred, which is not impossible but is now visible, slow and blockable.
What we would do differently
We spent time early trying to work out who was doing it. It is an interesting question and it does not change the answer: the fix is the same whoever it is, and identity work on rotating addresses is a hobby, not a deliverable.
We also left one edge rule deliberately in counting mode rather than blocking, because it was triggering on genuine uploads from real users. That is written down with its reason, and the same risk is handled properly elsewhere. It is the kind of exception that becomes a mystery in two years if nobody records why it exists, so it is recorded.
If this sounds like your system
Ask one question about any list your site publishes: what is the largest number of entries a single request can return, and who checked?
If the answer is “twenty, that is the page size”, ask whether there is another way to ask — a search, an export, a related-items parameter. There usually is, and it usually has no limit on it, because limits get written for the path somebody was thinking about.
If you would rather have someone establish that for you, write and tell us what you publish. The reply is free and it will say honestly whether this applies to you.