HNHacker News
TopNewBestAskShowJobs

Vasyl_R

58 karma · joined February 23, 2023

submissionscomments
Vasyl_R··on Show HN: Open-Source AI Embedding Pre-Processing Editor
We are thrilled to open source Embedditor.AI

Embedditor is the open-source MS Word equivalent for embeddings pre-processing, that helps you get the most out of your vector search, while saving up to 30% on embedding and storage costs.

This solution is inspired by the experiences of over 30,000 IngestAI users. Our insights revealed a common bottleneck in AI and LLM-related applications, one that goes beyond LLM hallucinations or token limits, which are far easier to resolve. The prevailing issue lies in the GIGO (garbage in, garbage out) principle.

With no one-size-fits-all approach to chunking and embedding, certain models excel with individual sentences, while others thrive on chunks of 250 to 500 tokens. Blindly splitting chunks by the quantity of characters or tokens, and embedding content without normalization and with up to 30% of redundant noise (such as punctuations, stop-words, and low-relevance frequent terms) often leads to suboptimal vector search results and low-performing LLM-related applications using semantic or generative search. The issue was consisting in trying to enhance vector search using existing technologies, which proved to be as challenging for our users, as creating an outstanding document using a basic .txt format.

We decided to address the root problem, so we developed Embedditor - the Microsoft Word equivalent for embedding pre-processing, enabling with no background in data science or technical skills to improve performance of their vector search capabilities while saving up to 40% on embedding and storage. We've made Embedditor open-source and accessible to all because we genuinely believe that by improving vector search performance and boosting cost-efficiency simultaneously, Embedditor may have significant impact on current NLP and LLM industry.

>>>FEATURES

>>>Rich editor GUI

->Join and split one or multiple chunks with a few clicks;

->Edit embedding metadata and tokens;

->Exclude words, sentences, or even parts of chunks from embedding;

->Select the parts of chunk you want to be embedded;

->Add additional information to your mebeddings, like url links or images;

->Get a nice looking HTML-markup for your AI search results;

->Save your pre-processed embedding files in .veml or .jason formats;

>>>PRE-PROCESSING AUTOMATION

->Filteer out from vectorization most of the 'noise', like punctuations or stop-words;

->Remove from embedidng unsignificant, requently used words with TF-IDF algorithm;

->Normalize your embedding tokens before vectorization;

GitHub: https://github.com/embedditor/embedditor

We hope you love it, and we would love to hear your feedback.

Vasyl_R··on Show HN: Veml – Vector Embedding Markup Language
hello Paul, thanks a lot for your feedback. I really appreciate that and we'll take it into consideration for our next steps. We saw people struggling with vector search that retrieves half of the relevant paragraph, just because it was chunked base on the qty of tokens. So our first step is to give users (I'm not talking about people that know Python, NLTK, and LangChain) can pre-process their embeddings, adding there images, and making cleansing, removing at least punctuations and stop-words, with a few clicks. But you're totally right - now we have to think not about a single document pre-processing but about embedding large set of documents.

Really appreciate your time and hope to have your star or see you among our watchers.

Vasyl_R··on Show HN: Veml – Vector Embedding Markup Language
apologies for that. Let me check with my co-founder. It will be there. Would be great to know your thoughts about our GUI for editing and embeddings + joining and splitting chunks, as well as filtering out punctuations and stop-words with one word. You can have a look on it in the /embedditor repo or in our web embedditor. ai.
Vasyl_R··on Show HN: Veml – Vector Embedding Markup Language
sure, it's in our docs. And you can edit it with a GUI available in the same repository but /embedditor .
Vasyl_R··on Show HN: Veml – Vector Embedding Markup Language
We’d like to present you the .veml [Vector Embedding Markup Language] that has the potential to greatly enhance the editing, use, and sharing of vector embeddings in various applications.

Adoption of VEML brings many benefits, like:

1. Standardization: VEML provides a standardized format for pre-processing and editing vector embeddings.

2. Interoperability: It ensures better interoperability among different applications and systems that utilize vector embeddings.

3. Extensibility: Just like XML, VEML has the potential to be extensible, allowing users to add new tags and attributes to represent additional properties or metadata associated with the vector embeddings.

5. Machine Readability: A well-defined markup language would also be easily parseable by ML, ensuring efficient processing and manipulation of vector embeddings by various software applications.

We'd love to hear your thoughts.

Vasyl_R··on Launch HN: Sidekick (YC W23) – AI support bot for developers
Congrats on this launch!

We at IngestAI[1] are in similar space albeit with different focus - we answer from all internal documentation in Slack/Discord/Whatsapp. We did a show hn just a few weeks ago[2] & has been wild since then - we have close to 10k+ users and counting.

If you would like build slack/discord/whatsapp/telegram bots answer your customer questions from internal docs - please try us

[1] : https://ingestai.io/

[2] : https://news.ycombinator.com/item?id=34909921

Vasyl_R··on Show HN: IngestAI – NoCode ChatGPT-bot creator from your knowledge base in Slack
That is a good point that you have brought up.

We plan to be agnostic of underlying LLMs as ecosystem matures.

Vasyl_R··on Show HN: IngestAI – NoCode ChatGPT-bot creator from your knowledge base in Slack
a good point. And what about companies that have on-premise storage?
Vasyl_R··on Show HN: IngestAI – NoCode ChatGPT-bot creator from your knowledge base in Slack
thanks for sharing. Do you think if all would be on the client's local server or cloud there would be still some, even rare, occurrence of that?
Vasyl_R··on Show HN: IngestAI – NoCode ChatGPT-bot creator from your knowledge base in Slack
yes, we can say it's something similar) Do you mean the maximum size our app supports? Currently we support files in different formats but with 10Mb size limit.
Vasyl_R··on Show HN: IngestAI – NoCode ChatGPT-bot creator from your knowledge base in Slack
would you be able to advise a good startup or service that provides good governance of Data Privacy management for startups? We would like to learn more and get this point as good as possible.. But you're right in some degree- we're builders, and we need help of professionals with data governance.
Vasyl_R··on Show HN: IngestAI – NoCode ChatGPT-bot creator from your knowledge base in Slack
wow. that's amazing! Do you means he's using our tool or something similar?
Vasyl_R··on Show HN: IngestAI – NoCode ChatGPT-bot creator from your knowledge base in Slack
Sorry, I could not get what is the meaning of 12 ? From your knowledge base we LangChain it to get the right context and ask answer. How do you envisage a solution of this sort work ? Happy to learn and make it better.
Vasyl_R··on Show HN: IngestAI – NoCode ChatGPT-bot creator from your knowledge base in Slack
we don't store conversations at all for the moment. I mentioned that in previous comments too. That's something some people ask, but we don't have 'conversation memory' for the moment))
Vasyl_R··on Show HN: IngestAI – NoCode ChatGPT-bot creator from your knowledge base in Slack
well. depends of how biased one is, I believe. No?
Vasyl_R··on Show HN: IngestAI – NoCode ChatGPT-bot creator from your knowledge base in Slack
thanks! We tried to do is as simple as possible from the users perspective. You just upload your knowledge base, create a Slack / Discord or other bot and start using it. And sure, we use OpenAI too. Yes, your're right - we use Davince model and then use use LangChain and other approaches that we're exploring to compare performance of different approaches. Have you heard about LangChain?
Vasyl_R··on Show HN: IngestAI – NoCode ChatGPT-bot creator from your knowledge base in Slack
thanks for your comment! Yep, we'd like to define it during our customer discovery, after we gather a lot of feedbacks, crystalize our understanding of the features users like and ask us to add. Maybe you'd be so kind to say how much this kind of products should cost from your POV ?
Vasyl_R··on Show HN: IngestAI – NoCode ChatGPT-bot creator from your knowledge base in Slack
We're at the moment using OpenAI and aim is to provide a ChatGPT like chat experience. And having ChatGPT api access would make it ideal though.
Vasyl_R··on Show HN: IngestAI – NoCode ChatGPT-bot creator from your knowledge base in Slack
please see my comment in this threads. We 're ONLY learning from your knowledge base and not from your slack chat history.
Vasyl_R··on Show HN: IngestAI – NoCode ChatGPT-bot creator from your knowledge base in Slack
hey, we are not learning anything from your slack history or channels.

The way IngestAI works is it takes your knowledge base as the input (markdown, docs etc) and answers the queries asked by the user on these knowledge base. ur primary usecase has been to simply learn from public documentation of companies and help answer the queries within their Slack/Discord community.

Vasyl_R··on Show HN: IngestAI – NoCode ChatGPT-bot creator from your knowledge base in Slack
hey, to clarify we are not learning from your Slack history or slack channels at all.

IngestAI is about learning from your knowledge base (markdown, docs, notion, confluence) and using that to answer queries for users within slack channel. Our primary usecase has been to simply learn from public documentation of companies and help answer the queries of community.

Vasyl_R··on Show HN: IngestAI – NoCode ChatGPT-bot creator from your knowledge base in Slack
hey, we are not learning from slack history. The Slack bot takes your knowledge base as the input (markdown, docs etc) and answers the queries asked by the user.
Vasyl_R··on Show HN: IngestAI – NoCode ChatGPT-bot creator from your knowledge base in Slack
We will work to fix this
Vasyl_R··on Show HN: IngestAI – NoCode ChatGPT-bot creator from your knowledge base in Slack
I thought I found some captcha on our web-site.. omg)) It was a good one)) Thanks))
Vasyl_R··on Show HN: IngestAI – NoCode ChatGPT-bot creator from your knowledge base in Slack
ahahah! Sorry! That's a good one!!!!! Amazing))) Love that! :)))
Vasyl_R··on Show HN: IngestAI – NoCode ChatGPT-bot creator from your knowledge base in Slack
sure, BERT is to come very soon and I think it will bring even wider adoption of AI. And yes, Data Privacy will become even more buzzed topic. Next few years will be of security engineers and of data privacy lawyers, don't you think?
Vasyl_R··on Show HN: IngestAI – NoCode ChatGPT-bot creator from your knowledge base in Slack
thanks for your comment. If believe if Data Privacy is almost the only topic discussed here, that's good. Means, people like the idea but at the same time people are concerned about personal / sensitive data. And to be honest- I hope that would be our biggest challenge, as I believe it's a bit easier to get a good, solid Data Privacy policy compliant with all legal and moral requirements than to build a product. At least that's what I hope))
Vasyl_R··on Show HN: IngestAI – NoCode ChatGPT-bot creator from your knowledge base in Slack
thank you so much for your support! Hope there're more people that think like you and are a bit 'empathic' about new services.. We improve with every single client, with every single request, and issue we face. So yes- Data Privacy is something very important but it's not only about IngestAI, I believe. Most of the SaaS solutions use API and are on AWS.. So, again, it's a wide topic, a buzz topic now.. Thanks again)
Vasyl_R··on Show HN: IngestAI – NoCode ChatGPT-bot creator from your knowledge base in Slack
thanks for your comment and support. That's what I'm trying basically to say in every third comment here. Generally, data privacy is a buzz topic since some time and with OpenAI it became even more viral.. It's not about IngestAI or any other SaaS, it's much wider topic and totally agree (thanks!).
Vasyl_R··on Show HN: IngestAI – NoCode ChatGPT-bot creator from your knowledge base in Slack
thanks for your comment! Are the data privacy concerns the main SaaS challenge from your POV ?
Page 1 of 2Next →