HNHacker News
TopNewBestAskShowJobs

jeadie

123 karma · joined October 30, 2022

jeadie.xyz
submissionscomments
jeadie··on Distributed DuckDB Instance
This is exactly what we found. Ingest rates were tough. We partitioned and ran over multiple duckdb instances too (and wrangled the complexity).

We ending up building a Sqlite + vortex file alternative for our use case: https://spice.ai/blog/introducing-spice-cayenne-data-acceler...

jeadie··on Distributed DuckDB Instance
You might find https://github.com/apache/datafusion and https://github.com/datafusion-contrib/datafusion-federation of interest
jeadie··on Vector database that can index 1B vectors in 48M
We’re building vector indexes into Datafusion for search (starting with S3 vectors).

Open source at https://github.com/spiceai/spiceai

jeadie··on Airport for DuckDB
This is one of the ideas behind using DuckDB in github.com/spiceai/spiceai
jeadie··on Show HN: TextQuery – Query CSV, JSON, XLSX Files with SQL
There’s also https://github.com/spiceai/spiceai
jeadie··on Pinecone integrates AI inferencing with vector database
This is a common feature now. If anything, for being so early to vector databases, Pinecone was rather late to integrating embeddings.

Timescale most recently added it but, yes a bunch of others: Weaviate, Spice AI, Marqo, etc.

jeadie··on Pg_parquet: An extension to connect Postgres and parquet
Why not just federate Postgres and parquet files? That way the query planner can push down as much of the query and reduce how much data has to move about?
jeadie··on Pg_lakehouse: Query Any Data Lake from Postgres
This looks functionally similar as using http://github.com/spiceai/spiceai with a postgreSQL data accelerator.
jeadie··on Ask HN: Who is hiring? (April 2024)
Spice AI | Senior Software Engineer | GMT+10 (e.g. Australia) through GMT-7 (e.g. Seatle/SF/LA) | Remote | Full Time

Spice AI provides building blocks for data and AI-driven applications by composing real-time and historical time-series data, high-performance SQL query, machine learning training and inferencing, in a single, interconnected AI backend-as-a-service.

We just launched github.com/spiceai/spiceai, a unified SQL query interface and portable runtime to locally materialize, accelerate, and query data tables sourced from any database, data warehouse, or data lake.

We're hiring experienced software engineers, ideally with Rust and/or Golang production experience. We're focused on large data and distributed systems, experience in these is important too. More details: https://spice.ai/careers#section-open-positions

jeadie··on Show HN: Spice.ai – materialize, accelerate, and query SQL data from any source
And yes, Iceberg is very high up on our list
jeadie··on Show HN: Spice.ai – materialize, accelerate, and query SQL data from any source
Yes! It can connect to FlightSQL compatible servers (see https://docs.spiceai.org/data-connectors/flightsql ) and its also a FlightSQL compatible server
jeadie··on Show HN: Yes, another vector embeddings API
Have you seen github.com/marqo-ai/marqo? It does all this wrapping, and you don't even need to pay for OpenAI or pinecone
jeadie··on GGML – AI at the Edge
I'm very glad that this has some added funding. I am building a serverless API on the cloudflare edge network using GGML as the backbone --> tryinfima.com
jeadie··on Weaviate – Open-Source AI Native Vector Database
"AI Native" catching on
jeadie··on PrivateGPT
I've tried both Chroma and Qdrant. I don't think Chroma lacks that much. Definitely newer, but is also a great product. I think cloud support coming Q3 2023
jeadie··on Ask HN: Seeking a Vector Database for ClickHouse Users – Suggestions Appreciated
(Not affiliated with hyperDB)
jeadie··on Ask HN: Seeking a Vector Database for ClickHouse Users – Suggestions Appreciated
I've been using https://github.com/jdagdelen/hyperDB and it's been really easy to use. I think Clickhouse support is on the short-term roadmap.
jeadie··on After All Is Said and Indexed – Unlocking Information in Recorded Speech
Most people, like me, who end up needing to use vector DBs, are wanting to use LLMs on a specific, often private dataset/use case. Typically one starts with something like unstructured JSON data, then need to pick and manage LLMs to create embeddings, then store these and the original JSON data in a vectorDB. Then the application is some variety of CRUD operations + searching over both the original data and the embeddings.

Chroma, Pinecone, I guess FAISS/HNSWlib/etc only handle vector operations. Really what I'd want, which Marqo does, is handle everything end to end.

jeadie··on After All Is Said and Indexed – Unlocking Information in Recorded Speech
Not a dumb question at all! Essentially what can do Marqo, and this blog shows, is that there is alot of logic and work to do what you said (i.e. pass raw data into LLM, get embeddings, store in vector DB, then query both embeddings and original data).
jeadie··on After All Is Said and Indexed – Unlocking Information in Recorded Speech
Its a great tool. Unlike vectorDBs alone, Marqo helps the full process that alot of people end up wanting to use vectorDBs for (e.g. have structured data, use LLMs to create embeddings, and perform search/CRUD on embeddings + original data).
jeadie··on After All Is Said and Indexed – Unlocking Information in Recorded Speech
Being able to handle and ask questions of audio data is a pretty big field. https://www.assemblyai.com/, for example, is a company entirely dedicated to audio intelligence. They have some great example use cases on their page.
jeadie··on Do you need a vector database?
This is generally very context/use case specific. In general, if a document is a `Dict[str, Any]`, then you either have to have one (or multiple) vector(s) per field, unless you want to combine vectors across fields (it's not self-evident how you'd best do that). In saying that, specific reason's to do this (or why I've done it in the past).

1. Chunking long text fields in documents so as to get a better semantic vector for them (also you can only fit so much into an LLM). 2. Differently to 1. chunking long text fields (or even chunking images, audio, etc), is one way to perform highlighting. It helps to answer the question, for example, for a given document what about it was the reason it was returned? You can then point to the area in the image/text/audio that was most relevant. 3. You may want to run different LLMs on different fields (perhaps a separate multi-modal LLM vs a standard text LLM), or like another comment said have different transforms/representations of the same field.

Perhaps 100 vectors is non-standard, but definitely not unseen.

jeadie··on Do you need a vector database?
I'm skeptical about some vector databases these days, but your article misses a few import points when it comes to LLMs.

1. To use LLMs effectively, you often need to generate and store more than 1 vector per document. 10 million vectors may only be 100,000 documents. This may still be enough for alot of small problems. 2. Pgvector currently has great limitations on recall/latency because underlying its ANN its using IVF (I'm currently working on adding HNSW-IVF and HNSW support to PGVector). In some cases, even elasticsearch can have issues with scale (the problem comes from the constraint of one ANN index per index segment, and immutability). 3. Pre-calculate seems like the wrong word to describe HNSW graph construction.

I think a point you miss that is important to consider for LLM + vector DBs is the fact that so much of the complexity of these uses cases cannot be captured by the vector DB (e.g. pinecone, chroma, qdrant, etc). I think there are some more end to end systems, at least in search, attempting to solve this (e.g Marqo, maybe Weaviate). Overall, I like the article. It makes a worthwhile claim and counterpoint to all the vector DB hype.

jeadie··on After All Is Said and Indexed – Unlocking Information in Recorded Speech
A really interesting blog post I found using LLMs for audio search which I think is a pretty nifty/new idea.

I've found it cumbersome using some of the new vector DBs (chroma, faiss, etc) to make end to end systems, but with Marqo it doesn't seem too hard.

jeadie··on I haven't heard much about ChatGPT and plugins. Is it meh or too good?
I think so, but checkout a bunch of awesome resources, and make up your mind https://github.com/Jeadie/awesome-chatgpt-plugins
jeadie··on Faiss: A library for efficient similarity search
Maybe https://github.com/hora-search/hora but I've never used it
jeadie··on Faiss: A library for efficient similarity search
I forgot about https://github.com/qdrant/qdrant. It's a DB not a library so again may not be an exact answer for what you're looking for
jeadie··on Faiss: A library for efficient similarity search
Although there is some work going on right now to add support for the type of algorithms in pgvector to alot it to scale better (and also to have better recall/speed tradeoffs).
jeadie··on Faiss: A library for efficient similarity search
I know rust has beings to FAISS (see https://github.com/Enet4/faiss-rs), I don't know if there's anything that would be considered comparable. Alot of work has gone into FAISS
jeadie··on Faiss: A library for efficient similarity search
A big difficulty in using vector DBs in production for things like embeddings or LLMs it that there is alot that goes into converting and processing raw input into a vector form (think chunking, formatting, encoding, inference, metadata, etc). DBs like pinecone just don't handle any of that and therefore you have to build out large systems to do it yourself.

There are some platforms and open source tools that handle it end to end. https://github.com/marqo-ai/marqo is one, for example that is both open source and has a cloud offering.

Page 1 of 3Next →