Run and create custom ChatGPT-like bots with OpenChat
github.com
github.com
A few thoughts:
* allow for custom endpoint URLs, this way people can use open source LLMs with a fake openAI API backend like basaran[2] or llama-api-server[3]
* look into better embedding methods for info-retrieval like InstructorEmbeddings or Document Summary Index
* Don't use a single embedding per content item, use multiple to increase retrieval quality
1 https://github.com/underlines/awesome-marketing-datascience/...
Can you share some specific examples of what you mean by this? How would you process specific info types (eg: news article, or web page, or product catalogue data) this way, and how would you handle retrieval that makes the quality "better"?
*Edit: Thanks for all replies so far - yes I am aware about splitting or chunking the data, but interested in a good write-up of techniques and pros/cons of each with examples. Eg: Chunking sentences vs. paragraphs, providing context around the embedding result, asking GPT to generate questions to chunks and embedding that instead, combining interaction data (eg: purchases or clicks after search queries) with actual content data before embedding, embedding attributes around data, and so on.
The other often neglected approach is to use an LLM to derive new content and use the embedding of this as well.
E.g. ask the LLM “give me a list of questions that can be answered by the following passage”
You then use embeddings of the generated questions instead of embeddings of the original content.
It wasn't perfect, though. I'm tempted to try the avenue of having an LLM generate questions for each passage and then use those embeddings, but it sounds a bit expensive to set up given the length of the books.
By any chance, have you done any work with indexing code using embeddings? I'd like to do something similar there, but there's no obvious notion of "sentence", especially across languages.
Probably the closest analogue is just lines of code, but breaking lines on newlines might break an expression in the middle removing meaning from both halves.
I was planning on trying indexing overlapping groups of lines but haven't had time yet.
Is this feasible? Thx.
- https://github.com/microsoft/guidance
- https://github.com/NVIDIA/NeMo-Guardrails/
- https://github.com/r2d4/rellm
- https://shreyar.github.io/guardrails/
- https://lmql.ai/
- https://github.com/jbrukh/gpt-jargonLook into https://github.com/NVIDIA/NeMo-Guardrails and specifically to your question there are "topical rails" to ensure the conversation stays on a set of topics you greenlighted.
Also takes care of jailbreaks and allows custom conversation flow templates.
Because it doesn't run "ChatGPT-like chatbots" which implies a non-OpenAI model with similar results, it just runs OpenAI tech in a wrapper?
If I understand how all of these OpenAI dependent apps work, none of them actually have the LLM and are doing any kind of heavy processing. AFAIK, they’re all packaging your data, submitting it to OpenAI on every request and then repackaging the output. There’s no real indexing, no real tangible thing, you have to start from scratch every time. So it’s likely going to be very expensive and super slow.
Or am I wrong and I’ve missed something here?
I think the most common design pattern nowadays goes like this:
1. Chunk all your data (e.g. per paragraph of content)
2. Generate an embedding for each chunk
3. Index embeddings in a vector database
4. When a query comes in, find chunks relevant to the query (based on embeddings similarity) and ONLY send the relevant chunks + query to a LLM to formulate the answer
Quickly glancing through the repository from this post, I can see that it also follows this pattern. It uses OpenAI's embedding API for 2. and Pinecone DB for 3.
I don't think it is as much the context window size because you would chunk your data anyways. I think the counter argument is either that finetuning is limited by the risk of overfitting and catastrophic forgetting or cost prohibitive. I think it is more of the former. Am I on the right track with this arguments?
Another point to consider is probably the vector DB contains an exact version of your data you get that as a result whereas the model will only be able answer vaguely or by paraphrasing.
[1] https://platform.openai.com/docs/guides/fine-tuning/prepare-...
Yes, that's what everyone says and it makes total sense to me. I'm looking for (technical, but not too technical) arguments why it is not possible. There I'm not so much interested in the "grounded in the facts of my website" point but more in the similar "take the data from my large private knowledge base into consideration" point.
In other words I don't want to restrict the knowledge the model has or the answers it gives. I want to add a considerable amount of my own knowledge. This seems not to be possible without training from scratch. The question is "Why?"
I feel lately - GPT-4 is superb in performance, but locked up. Using a weaker model feels better because I can just spin up a server and run it on my own. Recent Twitter/Reddit changes remind that relying on others can be a bad thing.
>providing PDF files, websites, and soon, integrations with platforms like Notion, Confluence, and Office 365.
Means that anything you feed this ChatBot, gets turned into data that's uploaded to OpenAI. So if you're using an internal Confluence, consider all that data public now. We've already seen intranet pages show up on ChatGPT/OpenAI in the past.
https://openai.com/policies/terms-of-use
Use of Content to Improve Services. We do not use Content that you provide to or receive from our API (“API Content”) to develop or improve our Services. We may use Content from Services other than our API (“Non-API Content”) to help develop and improve our Services. You can read more here about how Non-API Content may be used to improve model performance. If you do not want your Non-API Content used to improve Services, you can opt out by filling out this form. Please note that in some cases this may limit the ability of our Services to better address your specific use case.https://github.com/openchatai/OpenChat/blob/main/llm-server/...
What's the best discussion forum to exchange ideas, experiences and to collaborate on using and customizing (local) LLMs for (indy) gaming and other cool projects?
How is this possible? Do they do fancy tricks at inference?
I think the framing where the statement could be considered true is if you assume "memory = persistent storage", which is IMO not what most people will do.
Note: I'm not trying to imply the authors are being intentionally misleading.