HNHacker News
TopNewBestAskShowJobs

kacperlukawski

104 karma · joined May 2, 2022

submissionscomments
kacperlukawski··on Ask HN: Is building a calm, non-gamified learning app a mistake?
Although it's in a different area, I wanted to mention https://calmcode.io/ as an excellent example of a calm learning platform.

There is a whole movement around enshittification, and I see potential in this kind of app, even though it still seems to be a niche.

kacperlukawski··on Show HN: TalkNotes – A site that turns your ideas into tasks
Is there a free trial available? Even 24 hours should be enough to say if I like it, but currently I have to pay from the day one. Or did I miss it?
kacperlukawski··on Show HN: I made an extension that turns Google Sheets into Google Slides
I'm always a bit worried if an extension gets permission to do anything it wants with all my files, including deleting them. Is there a way to restrict it and allow it to modify only the files it created?
kacperlukawski··on Don't use cosine similarity carelessly
The problem is to scale that properly. If you have millions of documents, that won't scale that well. You are not going to prompt the LLM millions of times, aren't you?

Embedding models usually have fewer parameters than the LLMs, and once we index the documents, their retrieval is also pretty fast. Using LLM as a judge makes sense, but only on a limited scale.

kacperlukawski··on Instant Video Search
Interesting! Does it work based on speech or transcriptions?
kacperlukawski··on Automatically Detecting Under-Trained Tokens in Large Language Models
Why is that an issue? Training the tokenizer seems much more straightforward than training the model as it is based on the statistics of the input data. I guess it may take a while for massive datasets, but is calculating the frequencies impossible to be done on a bigger scale?
kacperlukawski··on Automatically Detecting Under-Trained Tokens in Large Language Models
Are there any specific reasons for using BPE, not Unigram, in LLMs? I've been trying to understand the impact of the tokenization algorithm, and Unigram was reported to be a better alternative (e.g., Byte Pair Encoding is Suboptimal for Language Model Pretraining: https://arxiv.org/abs/2004.03720). I understand that the unigram training process should eliminate under-trained tokens if trained on the same data as the LLM itself.
kacperlukawski··on OpenAI announces GPT-4.5 Turbo
Yeah, it seems like it got published too early.
kacperlukawski··on Qdrant 1.7.0
I'm unsure if there is any comparison of LanceDB and Qdrant available out there, but there shouldn't be any issues with Python 3.12 and qdrant-client compatibility. Windows is also not a problem, as the typical local setup is usually based on Docker. Are there any specific features you are interested in?
kacperlukawski··on Qdrant 1.7.0
If you will be the only app user, then the Python SDK's local mode might be suitable. However, in the long run, when you decide to publish the app, you rather have to switch to an on-premise or cloud environment. Using Qdrant from the very beginning might be a good idea, as the interfaces are kept the same, and the switch is seamless.

Local mode: https://github.com/qdrant/qdrant-client#local-mode

kacperlukawski··on Show HN: PromptTools – open-source tools for evaluating LLMs and vector DBs
Qdrant here! We're already working on that :D
kacperlukawski··on Serverless Semantic Search, Free tier only
How would you host sentence-transformers model for free? You need it to vectorize each query so that has to be hosted somewhere. Is there any way to do it for free?
kacperlukawski··on Serverless Semantic Search, Free tier only
If you need semantic search locally then it's fine, but serving an embedding model might be still challenging. And if you want to expose it, your laptop might be not enough.
kacperlukawski··on Introduction to vector similarity search (2022)
This is also an interesting piece of how to do it completely for free: https://news.ycombinator.com/item?id=36693239
kacperlukawski··on Serverless Semantic Search, Free tier only
It's a bit easier in Python if you use tools like https://www.serverless.com/. I'm not sure if Rust has something similar yet.
kacperlukawski··on MdBook – A command line tool to create books with Markdown
It's Hugo, with a custom styling. https://gohugo.io/documentation/
kacperlukawski··on Ask HN: Best tools to create LLM prodcuts?
Definitely Qdrant is a great option. It's an open source vector DB, so you can experiment locally without any cost, but if you prefer managed solution Qdrant Cloud is an option: https://cloud.qdrant.io/
kacperlukawski··on PrivateGPT
Chroma doesn't seem to be a real DB, it's rather a wrapper around tools like hnswlib, DuckDB or Clickhouse. Qdrant is way more mature - it has its own HNSW implementation with some tweaks to incorporate filtering directly during the vector search phase, supports horizontal and vertical scaling, as well as provides its own managed cloud offering.

In general, Qdrant is a real DB, not a library and that's a huge difference.

kacperlukawski··on The most cost-effective Vector Databases
Did you run the clients in the same regions as the servers? That may impact the results.
kacperlukawski··on Do you need a vector database?
There are some other options in between as well. FAISS is a library, so not suited well for production usage unless a single machine is enough. The variety is wider than SaaS vs library. Tools such as Qdrant or Weaviate are Open Source. And Qdrant might be launched without spinning a server, as long as you use a Python client. So you can actually start locally with in-memory mode, run your dev environment on-premise and then prod on Qdrant Cloud, if you prefer.
kacperlukawski··on Do you need a vector database?
I'd love to hear more about your thoughts on the complexity that cannot be, in your opinion, captured by the vector DB. I probably didn't get your point.

Disclaimer: I work for Qdrant, and we believe a database should be just a database. I remember attempting to move logic to the database layer and coupling neural encoders into the vector database sounds the same.

kacperlukawski··on Show HN: Semantic Search on AWS Docs
I'm still wondering why OpenSearch and ES have those limits for the dimensionality of the embeddings while the vector databases such as Qdrant do not.
kacperlukawski··on Ask HN: How long before databases get vector similarity search in them?
It's a different paradigm. I wouldn't expect the relational databases to start competing with the proper vector databases, as vectors do not fit SQL tables. Each problem should be solved with a specialized tool to achieve the best performance possible. We might be seeing some new plugins to relational DBs, but in reality, vector search comes with a bunch of other issues we need to solve. For example, memory usage. Here is how we made it at Qdrant: https://qdrant.tech/articles/scalar-quantization/
kacperlukawski··on Introducing Agents in Haystack: Make LLMs resolve complex tasks
That's a great news! Especially since we implemented the integration between Haystack and Qdrant: https://github.com/qdrant/qdrant-haystack/
kacperlukawski··on Ask HN: Best way to “donate” dev hours to charity?
I'd say doing some sort of training might be the most beneficial thing you can bring to the NGOs. There are also organizations such as Omdena (https://omdena.com/). They solve different challenges and would probably appreciate some support, even some sort of mentorship. We supported them that way at Qdrant, and organized a semantic search workshop, while one of their local chapters was implementing a chatbot.
kacperlukawski··on Launch HN: Metal (YC W23) – Embeddings as a Service
What are your plans for providing some additional metadata except for embeddings? Semantic search often requires additional filtering, as vectors are not all we need. At Qdrant we have a unique mechanism for incorporating metadata filters into HNSW, so they might be applied during vector search phase (no pre- or post-filtering required): https://qdrant.tech/documentation/indexing/#filtrable-index
kacperlukawski··on ChatGPT Opened a New Era in Search. Microsoft Could Ruin It
Yeah, that's correct. We also believe that's a good direction for improving search results quality and combining different methods of search: https://qdrant.tech/articles/hybrid-search/
kacperlukawski··on Qdrant Scalar Quantization: up to 2x faster, 4x less memory
Qdrant 1.1.0 introduced the scalar quantization mechanism. It means up to 4x lower memory footprint and up to 2x performance increase with little to no impact on the search precision.

Disclaimer: I work for Qdrant

kacperlukawski··on Vector database built for scalable similarity search
I'm affiliated with Qdrant and was quite surprised to see us listed in your comparisons with some false statements. If you claim to be the only database with billion-scale vector support, it would be great to make your benchmarks public, as we did: https://qdrant.tech/benchmarks/

Btw, Milvus is described in your comparisons as "a fully open source and independent project", while Weaviate and Qdrant, in contrary, are "maintained by a single commercial company offering a cloud version". Why then the suggested way in the Milvus Quick Start on github is to use Zilliz Cloud?

kacperlukawski··on ChatGPT Plugins
Yeah, you need to enable the plugins you want. I'm just saying you can enable all the ones that make sense for you, and you don't have to switch between them.
Page 1 of 2Next →