I get it may be interesting for small tasks combined with a browser extension but for real scrapping just seems to be overkill and expensive.
I get it may be interesting for small tasks combined with a browser extension but for real scrapping just seems to be overkill and expensive.
Instead of writing a bunch of selectors that break often, imagine just being able to write a paragraph telling the LLM to fetch the top 10 headlines and their links on a news site. Or to fetch the images, titles, and prices off a store front?
It abstracts away a lot of manual fragile work.
Today, would you build a scraper with current LLMs that randomly hallucinate? I wouldn't.
The idea of a LLM powered scraper adapting the selectors every time the website owner updates it, it's pretty cool.
Do you have any workflow tools etc. to find hallucinations, I've got a project in backlog to build that kind of thing and would be interested in how you sort through bad and good results.
This is pretty much what we're building at Skyvern. The only problem is that inference cost is still a little bit too high for scraping, but we expect that to change in the next year