I’m doing something similar: I print the website to PDF using the reader mode. I also copy the raw text into a markdown file. Both files end up in my Zettelkasten. I have a little search engine on top of that. For programming tasks, I often need to retrieve information that I searched before.
Recently, this has also become more useful as Google shows only ads on the first pages.
Nice workflow - I wonder what you're doing with websites that can't be displayed in reader mode or websites that renders awefully when printing to pdf with clipping text at the borders. Imo, this breaks the basically good idea of annotating and archiving web sources. Obviously theres no solution out there at the moment, or did you find one?
I'm trying not to optimize this too much. It’s unclear if I really need something in the future given that we have something like ChatGPT already. Also I don’t want to be an information hoarder.
In short, if something doesn’t work, I just manually paste a snippet.
The PDF doesn’t need to look nice. It’s more that I can get an impression of how the website looked.