[0] https://www.everycrsreport.com/reports/RL33810.html
[1] https://en.wikipedia.org/wiki/Fair_use#Text_and_data_mining
If you are afraid of being sued, you certainly should not do something like this.
On the upside, there is probably not much provable damage a publisher could ask to be repaired. So they probably are only liable to pay for the legal expenses of whoever sues them.
We mainly surface historical content, which receives additional traffic from DeepHN, instead of taking value away from the website.
That said, meaningful fulltext search will need data. But both Bing and Google crawl pages.
It's not even immediately clear that displaying a copy of a page is materially different from caching (which http do in many, many layers).
Now, if présent the text as your own - that might be a problem.
Ed: see also: archive.org.
Alternatively, base your site out of a country which gives you the freedom to do things like this. Build it in China and you'll be applauded instead of sued. Copyright suits for petty things like this would laughed off when it's something actually useful.
Founders there have no choice but to sell or get copied by the platform that you're building on. It's baba or tenscent.
I agree, some anti-monopoly regulation is in order in China, and I'm not a fan of the Alibaba/Tencent monopolies on the ecosystem. However, I think there are other ways to go about fostering and encouraging small businesses and protecting creators and their intentions than American-style IP law, which allows you to sit on an invention or work and hinder society from having access to the fruits of science, which I'm vehemently against. IP isn't even real property, IMHO.