A simple "running queries over the whole dataset can cause significant costs due to the size of the dataset" should be enough. And I think that's a valid and fair point.
The whole part of accusing Google should just be ignored.
https://github.com/HTTPArchive/httparchive.org/blob/main/doc...
I don't think that's a proper warning on costs.
I don't know. Google could trivially solve this problem by imposing an opt-out warning on potentially expensive queries.
"It looks like your query might cost $14k. Are you sure?"
But money.
AWS charges $27/hour for a server with 3TB of memory. Enough to run the queries in memory.
Again, you could fit the whole dataset in memory in an EC2 instance and do your thing.
A regex query on response_bodies would churn through 2.5TB of data every time it's run.