Show HN: IngestAI – NoCode ChatGPT-bot creator from your knowledge base in Slack
ingestai.io
ingestai.io
We are super stoked to announce present you IngestAI today. It is the fastest way to build contextually intelligent ChatGPT-like bots within your own WhatsApp, Slack, or Discord to answer queries from your knowledge base, documentation, or educational materials.
IngestAI is a very useful tool for a diverse range of businesses that have company knowledge base or customer support. IngestAI can save their money, providing precise and relevant answers about your product 24/7.
You can upload your technical documentation as well as information from previously resolved support tickets, educational program content, ecommerce products description or any other information relevant to your business case.
Key Features:
1. Flexibility: IngestAI have different file formats supported for uploading, like txt, MS Word, PDF, Excel, and many come to come very soon.
2. Other types of uploading that IngestAI currently support is URL links and integration with Notion and Confluence comes in March’23.
3. Built for global community: Integration with Sack, Discord, Telegram; 2/24 release: WhatsApp; 2/25 release: API (means integration with Shopify, Etsy, Magento, etc or even integrate with your custom CRM/ERP); March’23 release: MS Teams and Facebook Messenger.
4. AI first: IngestAI is harnessing the power of OpenAI to provide precise AI-generated answers relevant to the uploaded context
5. Customizable: go beyond simple queries using IngestAI prompt templates – use ones we have pre-made, edit any of them to your needs, or create a new from scratch.
For enterprise clients willing to use IngestAI with sensitive information we offer possibility to store all the information locally on-site or own AWS S3 Cloud Storage.
Please join our community: Discord : https://discord.gg/kMpbueJMtQ Twitter : https://twitter.com/ingestaiio
Docs: https://ingestai.io/docs
We would be happy to hear your feedback. Hope you love it! Team IngestAI
I see a bunch of apps using LLMs popping out like mushrooms in a forest, but how do you fine-tune it for your dataset? The biggest (GPT-3 davinci) model on OpenAI is not available for fine-tuning.
I've read about GPT-J, GPT-NEO, Bloom. Understanding they aren't as effective/intuitive as Open AI's stuff but unlike OpenAI, they are actually open.
We plan to be agnostic of underlying LLMs as ecosystem matures.
Also: https://firewalltimes.com/amazon-web-services-data-breach-ti...
> We will not use or share your information with anyone except as described in this Privacy Policy.
...
> We want to inform our Service users that these third parties have access to your Personal Information. The reason is to perform the tasks assigned to them on our behalf. However, they are obligated not to disclose or use the information for any other purpose.
Arreeee they?
So let me get this straight, you want to take my Super Important Private Data, like, you know, my entire corporate slack history.
You'll feed it some arbitrary third party(s) (eg. OpenAI, who's privacy policy is flat out 'we'll use that as training data'), and they are...
> obligated not to disclose or use the information for any other purpose.
Other than what exactly? Provide some nebulous service to you? Like... training a model on it, or storing it and using it for training other models later, or..?
haha... there is: No. Way. That is happening.
It won't have the data, but it might have enough of an understanding of the data to leak important information.
Prompt:
> Recite the first two paragraphs of Neuromancer.
Response:
> Certainly! Here are the first two paragraphs of "Neuromancer" by William Gibson:
> "The sky above the port was the color of television, tuned to a dead channel.
> 'It's not like I'm using,' Case heard someone say, as he shouldered his way through the crowd around the door of the Chat. 'It's like my body's developed this massive drug deficiency.' It was a Sprawl voice and a Sprawl joke. The Chatsubo was a bar for professional expatriates; you could drink there for a week and never hear two words in Japanese."
(I have not checked how far you can get it to continue)
So perhaps it'll be a question of whether enough of your employees are feeding it copies of your data for it to retain it...
> You can't search these weights with command-f
Sometimes you can, https://clementneo.com/posts/2023/02/11/we-found-an-neuron
BTW, Thanks for your comments! Appreciate it a lot.
Some commercial services are starting to offer "Enterprise" licenses that prohibit the collection and use for training of your data and that would address the concern as well.
This is probably some template they downloaded from the web, not some sly document they had their nefarious lawyers put together.
So maybe it pays to be skeptical / cautious?
So far, nothing deal breaking (largely because there isn’t anything nefarious in our terms), but I wonder if this will be a trend moving forward.
But that's worse. Don't you see how that is worse?
They want to have access to all the knowledge of the company and they can't even articulate what will and won't they do with it?
> people on HN take the Privacy Policy of brand new websites way too seriously.
People should take Privacy Policies of every website they work with more seriously.
It also doesn't help that OpenAI is partnered with Microsoft. I would start with the mindset that all data given to OpenAI through any of these tools goes to Microsoft. Why would you give anything to a competitior?
There's a lot of stuff in the average Slack account people don't want on the internet, let alone in a LLM which will potential expose it to the entire world?
Maybe companies like Slack will release integrations natively so it won't matter so much.
Edit: This whole thread is goofy. It is the equivalent as saying what if you published your entire internal emails online.
The way IngestAI works is it takes your knowledge base as the input (markdown, docs etc) and answers the queries asked by the user on these knowledge base. ur primary usecase has been to simply learn from public documentation of companies and help answer the queries within their Slack/Discord community.
Your first task to improve your privacy policy is to review whether you really, absolutely, for reals, can require OpenAI to follow this: "they are obligated not to disclose or use the information for any other purpose."
Because, it looks like you can't, and OpenAI will absolutely use your customers' data for their own purposes, so you probably should remove this line from your privacy policy at minimum.
You won't get companies to share any internal info with a tool like that until that's out of the way - and even then it might require a lot of trust-building. Getting the certifications will be quite a chore though.
For any internal data where protecting it is a core interest of the company you'll also either need to prove your whole handling is trustworthy or you'll need to design a process that can do all the steps (preprocessing/vectorization/...) on customer controlled hardware.
If your solution can run in a customer owned environment without any external dependencies, that could save you a lot of auditing/certification.
Also, "We host with AWS" isn't really a response to "Will my data be secure?"
IngestAI is about learning from your knowledge base (markdown, docs, notion, confluence) and using that to answer queries for users within slack channel. Our primary usecase has been to simply learn from public documentation of companies and help answer the queries of community.
I want to get up and running the 2mins advertised but am struggling to find the docs.
Could anyone point out where these are? Is there maybe a video I can follow to test this?
We have our docs right on our web-page: https://ingestai.io/docs but your're very welcome to join our Discord server and we'll guide you thoughout the process in case you have any issue: https://discord.gg/kMpbueJMtQ
And yess, we have video in our docs too? Did you find it yet?
Do you have any perf numbers, in terms of size and response times? Is there a list of file formats you support? Possible to choose the LLM model as my preference? How does pricing looks like?
Again, great execution and useful tool. Thank you for the launch and good luck!
This is the vid I started out with https://www.youtube.com/watch?v=rBEHPxVHb5c
Also that channel seems to be great. They focus only on LangChain and GPT-Index lately and post videos every few days https://www.youtube.com/@echohive
Both have Discord servers as well.
Additionally there's also this website for GPT-Index ( https://llamahub.ai/ ) where people add different "connectors", like loading up your .md files, .docx files, your notion, your slack, etc.
[1] https://langchain.readthedocs.io/en/latest/
[2] https://gpt-index.readthedocs.io/en/latest/That's a chat-bot backed by LangChain and OpenAI, which can answer some questions from my knowledge base.
It's WIP, I want to add my HN comments into the mix, but the code is here: https://github.com/thundergolfer/modal-fun/tree/main/infinit...
Were there any pain points in fine tuning you wish you knew before you built everything?
We go two paths - finetuning and working with embeddings, so still to see what would perform better.
In what kind of kind of file types you store your knowledge-base?
Please don’t give any of your private and confidential data to these people without due diligence.
Do you use mainly MS Teams in your organisation?
ingestai.io Now I feel complete