An example use case: given a url, figure out content type (article vs product page for example), if product page, automatically extract all product details and specs without manually mapping xpath or css lookup paths
An example use case: given a url, figure out content type (article vs product page for example), if product page, automatically extract all product details and specs without manually mapping xpath or css lookup paths
There was a show HN here a couple of days ago that used chatgpt to create a scraper based on a webpage. Hook the 2 together and you are basically there!
The mozilla/readability library is a good first step though.
Edit: I'm thinking of something like LlamaIndex
Yes, I could try to figure out which bits of the page need to be sent along with the prompt, but that's hard in the general case. Squeezing a bit more out of the prompt window by stripping out unnecessary boilerplate is easy by comparison. (Multi-page articles are another headache).