Probably not helpful, but I built a setup like that for myself years ago. I had a local proxy, originally written with Python Twisted, but later ported to Go, and set my browser to use that, so all requests went through it. Every URL that the proxy saw, it logged and for text/plain or text/html, it also took a copy of the request body and posted that to a web service that I was running. The web service was a Django app with a simple model that would track the URL, timestamp, and some other basic metadata. It would also save a gzipped copy of the request body to S3, and dump it into a SOLR instance. That gave me full-text search over the content of every site I visited and a backup copy of the text in case the original ever went offline. It was incredibly useful.
That was back in the day though, before HTTPS was common (outside ecommerce sites, which I didn't care about indexing), and before so many sites were SPAs that got their content via JS APIs. As more sites went to HTTPS, I realized that I'd have to re-write my proxy to MITM certificates if I wanted it to keep working and that wasn't really something I wanted to mess with. The project was useful, but not useful enough that I was willing to dive into the world of writing a browser plugin that could scrape directly from the DOM, so I eventually abandoned it.