Danswer
https://github.com/danswer-ai/danswer
Khoj
So I don't really buy it and I have yet to see it work better than any rdbms search index.
Tell me I am wrong, I would like to see a local model based on my own docs being able to answer me quality answers based on quality prompts.
Basically if you have a database of three emails and ask when Biff wanted to meet for lunch, a RAG system would select the most relevant email based on any kind of search - embeddings are most fashionable, and create a prompt like
"""Given this document: <your email>, answer the question "When does Biff want to meet for lunch?"""
Sibling comment from discordance has a more accurate description of RAG. There's a longer description from Nvidia here: https://blogs.nvidia.com/blog/what-is-retrieval-augmented-ge...
> Given the prompt “When did the first mammal appear on Earth?” for instance, RAG might surface documents for “Mammal,” “History of Earth,” and “Evolution of Mammals.” These supporting documents are then concatenated as context with the original input and fed to the [...] model
Finding the relevant context to put in the prompt is a search problem, nearest neighbour search on embeddings is one basic way to do it but the singular focus on "vector databases" is a bit of hype phenomenon IMO - a real world product should factor a lot more than just pure textual content into the relevancy score. Or is your personal AI assistant going to treat emails from yesterday as equally relevant as emails from a year ago?
1. First you create embeddings from your documents
2. Store that in a vector db
3. Ask what the user wants and do a search in the vector db (cosine similarity etc)
4. Feed the relevant search results to your LLM and do the usual LLM stuff with the returned embeddings and chunks of the documents
Would you define RAG only as 'prompt optimisation that involves embeddings'?
[0]: https://openai.com/chatgpt/pricing#:~:text=8K-,32K,-32K
[1]: https://platform.openai.com/docs/models/gpt-4-and-gpt-4-turb...
This is Gabe, the founder of Zenfetch. Thanks for sharing. We're putting together an export option where you can download all your saved data as a CSV and should get that out by end of week.
I want the ability to search all my downloaded files and organize them based on context within. Have it create a category table, and allow me to "put all pics of my cat in this folder, and upload them to a gallery on imgur."
We've been thinking of this as a "subscription" to the creator's folder. Similar to how you might subscribe to a Spotify playlist
nice.
Use all the resources you want if you save me brainpower
Help me plan for upcoming meetings whereby if I put something in calendar, it will build a little dossier for the event, and include relevant info based on the type of event or meeting, mostly scheduling reminders or prompting you with updates or changes to the event etc.
I'm not sure I can imagine a scenario in production where Google would, or should, allow API access to individual gmail accounts. What's that for? So you can read all your employees' mail without running your own email server?
> You will no longer use a password for access (with the exception of app passwords)
I'm not seeing anywhere that I'd need to pay money to use OAuth via an app like Thunderbird or another email client. That app would either need to support using OAuth to let the user auth and get credentials, or use an app password.
I manage both gmail and protonmail via thunderbird - where I have better search and sort using IMAP.