HNHacker News
TopNewBestAskShowJobs

angelhadjiev

4 karma · joined December 20, 2025

building a web data harvesting infrastructure at 4A. https://foura.ai
submissionscomments
angelhadjiev··on Web Data – Latest Capital Raise
This summer $231M went into collecting the web for AI.

- Collection: Oxylabs, $130M from Warburg at $3.6B (~10x ARR) - Index: Keenable, $26M seed, Accel - Knowledge: Firecrawl, $75M Series B, 1.5M users

...money followed the data.

angelhadjiev··on Ask HN: Will Claude watermarking hurt your SEO? (I guess for sure it wont help)
The question, whilst rhetorical, supposes we have an alternative.
angelhadjiev··on Google's new hand-wave reCAPTCHA can be bypassed with a stock photo
"(...) this incident, combined with users’ initial skepticism about Google’s practices regarding user data, likely won’t make too many people wave at the camera anytime soon."

;]

angelhadjiev··on Proxy Pool Size Means Nothing in 2026
also interesting benchmark on fingerprinting: https://www.reddit.com/r/webscraping/comments/1tq1es9/browse...
angelhadjiev··on Web scraping tarpits are catching legitimate data teams, not just AI crawlers
;] Not a bot - just sleep-deprived. Spent last night chasing a tarpit at 2am.

You're right that scraping has a bad reputation (still, although it's one of the top topics on google words), and some of it is well-deserved.

The moral framing is fair in the training-crawler context, but the article's point is about collateral damage to legitimate use cases. Price comparison, research, public data pipelines... these aren't the bad actors, they just look like them.

That's the gap worth closing in my opinion.

angelhadjiev··on Web scraping tarpits are catching legitimate data teams, not just AI crawlers
Fair point. Direct outreach works when you can identify who to contact and they’re responsive. In practice though, most data teams are scraping hundreds of domains, not one. The hostmaster path doesn’t scale, and tarpits often get deployed at the CDN/WAF layer (Cloudflare, Vercel) where there’s no meaningful human on the other end anyway.

Curious to know have you had success with that approach at scale, or more for one-off access agreements?

angelhadjiev··on Web scraping tarpits are catching legitimate data teams, not just AI crawlers
Sites are deploying infinite fake-page mazes (Nepenthes, Locaine, etc.) to trap and poison AI training crawlers that ignore robots.txt. The motivation is understandable — Cloudflare reported 75% of AI web traffic in mid-2025 was training-related, and nearly 60% of reputable sites now block AI bots.

The problem: tarpits don't check intent. They detect automated request patterns. If your price tracker follows links systematically, skips JS execution, or hits pages at regular intervals — it looks identical to GPTBot. The trap fires anyway.

The collateral damage is real. One Rutgers/Wharton study found sites with aggressive crawler blocking saw a 23% drop in total traffic, including human visitors.

The escalation ladder is now at step 4: 1. robots.txt (gentleman's agreement) 2. User-agent filtering 3. Behavioural detection 4. Active tarpits — waste your compute, poison your data

If you're running any data pipeline at scale, you need to validate responses now. Tarpits serve plausible-looking Markov garbage. If you're not checking, it's already in your database.

Full writeup: https://foura.ai/blog/web-scraping-tarpits-collateral-damage