----
What model size/particular fine-tuning are you using, and how have you observed it to perform for the usecase? I've only started playing with Llama 2 at 7B and 13B sizes, and I feel they're awfully RAM heavy for consumer machines, though I'm really excited by this possibility.
How is the search implemented? Is it just an embedding and vector DB, plus some additional metadata filtering (the date commands)?