> "By scraping raw text with AI, FetchFox lets you circumvent anti-scraping measures on sites like LinkedIn and Facebook. Even the the complicated HTML structures are possible to parse with FetchFox."
> "By scraping raw text with AI, FetchFox lets you circumvent anti-scraping measures on sites like LinkedIn and Facebook. Even the the complicated HTML structures are possible to parse with FetchFox."
The trickier part is “everything else” to make the extension work.
1. Remove all <style> and <svg > tags. These rarely add value, and can dramatically increase token counts.
2. For the “crawl” step, I exclusively pull out <a> tags and only look at those. The “extract” step looks at full HTML
3. For now, it only looks at the first 50k text characters, and the first 120k HTML characters. This is to stay within token limits.
The last part will be what I focus on improving in the next version.
They keep throwing it in my url bar. I refuse to click (big warning it sends to google's servers)
The benefit of this approach is it's very simple and easy, but the downside is it sends a lot of unnecessary tokens to the LLM. That drives up the cost, slows things down, and hurts accuracy.
I'm working on a few improvements now to improve this.