Trieve is a very impressive achievement, you've managed to slam together every AI buzzword into a semi usable product! It's written in Rust, which makes the entire proposal even more perfect for HN :). Love it.
Some questions (and thanks for the detailed "about" page, it answered several of my initial questions!) ~
Will you re-index and keep updating the system to improve the quality of results, or what is the plan? It'd be awesome to have something more nuanced than Algolia which stays updated in near real time, like Algolia.
How easy or challenging is it to bootstrap / re-index? Is it possible to ingest new data with partial updates to the existing indices, or is a full indexing from zero always required?
Are GPUs strictly necessary? Is it possible to use only CPUs if indexing speed isn't a great concern?
Does it really require a terabyte of working memory to index and serve all of the data for HN? (4x 256GB / 128 CPUs is mentioned in your ops details) This is a lot of resources! Like A LOT!
Have you thought considered indexing other high-quality data sources? For example:
* Lobste.rs (I think you can email them requesting a DB dump, they want disclosures about the intended purpose)
* Slashdot (debatable quality, but goes back 27 years which could be interesting)
* Review sites: Chipsandcheese, TomsHardware, Anandtech, HardOCP
* Lwn, Phoronix
* I'm surely missing other good ones, the discussion in The End of Anandtech article from today mentions a bunch of interesting sources: https://news.ycombinator.com/item?id=41399872
I wonder if getting some of these data sources through CommonCrawl or archive.org would reduce the crawl+parse annoyance?
At some point I want to put together an HN-Awesome-Search page which covers all the custom search indexes HN folks have made over the years.
Thank you!