(the reasonable way is embedding search, which runs much faster with some precomputation, but you still have to store things)
That whole thing can be simplified to: compute and store embeddings for docs, compute embeddings for query, find most similar docs.
Is this because you want it to continuously watch for live data that could match your need?
If I go through my current tasks and see, that for some task I need a set of documents, emails, .., why cant I just prompt the system to get it in 30-ish minutes. But as someone already stated Apple Intelligence is supposed to fill this gap.
Many of us have ongoing problems pending for years - for just "a week", "where do I sign".
It really depends on the task.
>The idea was that he could graft queries in this that he did not expect to finish quickly but which he could let run for hours or days and how freeing it was to do more advanced research this way.
Just run the biggest model you can find out of swap and wait a long time for it to finish.
You'll obviously see more focus on smaller models, because most people aren't willing to wait weeks for their slop, and also don't have server GPU clusters to run huge models.
This kills the SSD