ScrapeGraphAI: Web scraping using LLM and direct graph logic
scrapegraph-doc.onrender.com
scrapegraph-doc.onrender.com
Ad blockers have had something very close to this for some time, without any sparkly AI buttons.
I’m sure someone would be working on a subscription based model using corporate models in the backend, but it’s something that could easily be implemented with a very small model.
the statistics are not in its favour
The complicated ops of scraping is running headless browsers, IP ranges, bot bypass, filling captchas, observability and updating selectors, etc. There are a ton of SaaS services that do that part for you.
Apify's Website Content Crawler[0] does a decent job of this for most websites in my experience. It allows you to "extract" content via different built-in methods (e.g. Extractus [1]).
We currently use this at Magic Loops[2] and it works _most_ of the time.
The long-tail is difficult though, and it's not uncommon for users to back out to raw HTML, and then have our tool write some custom logic to parse the content they want from the scraped results (fun fact: before GPT-4 Turbo, the HTML page was often too large for the context window... and sometimes it still is!).
Would love a dedicated tool for this. I know the folks at Reworkd[3] are working on something similar, but not sure how much is public yet.
[0] https://apify.com/apify/website-content-crawler
More interesting issue is being able to parse data from the whole page content stack which includes XHRs and their triggers. In this case LLM driver would control an indistinguishable web browser to perform all steps to retrieve the data as a full package. Though this is still a low value proposition as the models would get fumbled by harder tasks and easier tasks can be performed by a human being in couple of hours.
LLM use in web scraping is still purely educational and assistive as the biggest problem in scraping is not scraping itself but scraper scaling and blocking which is becoming extremely common.
I get it may be interesting for small tasks combined with a browser extension but for real scrapping just seems to be overkill and expensive.
Instead of writing a bunch of selectors that break often, imagine just being able to write a paragraph telling the LLM to fetch the top 10 headlines and their links on a news site. Or to fetch the images, titles, and prices off a store front?
It abstracts away a lot of manual fragile work.
Today, would you build a scraper with current LLMs that randomly hallucinate? I wouldn't.
The idea of a LLM powered scraper adapting the selectors every time the website owner updates it, it's pretty cool.
Do you have any workflow tools etc. to find hallucinations, I've got a project in backlog to build that kind of thing and would be interested in how you sort through bad and good results.
This is pretty much what we're building at Skyvern. The only problem is that inference cost is still a little bit too high for scraping, but we expect that to change in the next year
I don't believe scraping is such a solved problem that you can slap AI and some cute vector spiders on it and claim that everything works.
Something like https://github.com/Alir3z4/html2text.
I'm sure there are other (better?) options as well.
I haven't tried this library but I do use an LLM based scraper in addition to more traditional ones.