- Collection: Oxylabs, $130M from Warburg at $3.6B (~10x ARR) - Index: Keenable, $26M seed, Accel - Knowledge: Firecrawl, $75M Series B, 1.5M users
...money followed the data.
4 karma · joined December 20, 2025
- Collection: Oxylabs, $130M from Warburg at $3.6B (~10x ARR) - Index: Keenable, $26M seed, Accel - Knowledge: Firecrawl, $75M Series B, 1.5M users
...money followed the data.
;]
You're right that scraping has a bad reputation (still, although it's one of the top topics on google words), and some of it is well-deserved.
The moral framing is fair in the training-crawler context, but the article's point is about collateral damage to legitimate use cases. Price comparison, research, public data pipelines... these aren't the bad actors, they just look like them.
That's the gap worth closing in my opinion.
Curious to know have you had success with that approach at scale, or more for one-off access agreements?
The problem: tarpits don't check intent. They detect automated request patterns. If your price tracker follows links systematically, skips JS execution, or hits pages at regular intervals — it looks identical to GPTBot. The trap fires anyway.
The collateral damage is real. One Rutgers/Wharton study found sites with aggressive crawler blocking saw a 23% drop in total traffic, including human visitors.
The escalation ladder is now at step 4: 1. robots.txt (gentleman's agreement) 2. User-agent filtering 3. Behavioural detection 4. Active tarpits — waste your compute, poison your data
If you're running any data pipeline at scale, you need to validate responses now. Tarpits serve plausible-looking Markov garbage. If you're not checking, it's already in your database.
Full writeup: https://foura.ai/blog/web-scraping-tarpits-collateral-damage