HNHacker News
TopNewBestAskShowJobs

chuckhend

848 karma · joined February 7, 2023

https://github.com/ChuckHend
submissionscomments
chuckhend··on Pg_vectorize: Vector search and RAG on Postgres
There is a RAG example here https://github.com/tembo-io/pg_vectorize?tab=readme-ov-file#...

You can provide your own prompts by adding them to the `vectorize.prompts` table. There's an API for this in the works. It is poorly documented at the moment.

chuckhend··on Pg_vectorize: Vector search and RAG on Postgres
There is no chunking built into the postgres extension yet, but we are working on it.

It does check the context length of the request against the limits of the chat model before sending the request, and optionally allows you to auto-trim the least relevant documents out of the request so that it fits the model's context window. IMO its worth spending time getting chunks prepared, sized, tuned for your use case though. There are some good conversations above discussing methods around this, such as using a summarization model to create the chunks.

chuckhend··on Pg_vectorize: Vector search and RAG on Postgres
pg_vectorize is a wrapper around pgvector. In addition to what pgvector provides, vectorize provides hooks into many methods to generate your embeddings, implements several methods for keeping embeddings updated as your data grows or changes, etc. It also handles the transformation of your search query for you.

For example, it creates the index for you, create cron job to keep embeddings updated (or triggers if thats what you prefer), handles inserts/upserts as new data hits the table or existing data is updated. When you search for "products for mobile electronic devices", that needs to be transformed to embeddings, then the vector similarity search needs to happen -- this is what the project abstracts.

chuckhend··on Pg_vectorize: Vector search and RAG on Postgres
RAG can cost a lot of money if not done thoughtfully. Most embedding and chat completion model providers charge by the token (think number of words in the request). You'll pay to have the data in the database transformed into embeddings, that is mostly a one-time fixed cost. Then every time there is a search query in RAG, that question needs to be transformed. The chat completion model (like ChatGPT 4), will charge for the number of tokens in the request + number of tokens in the response.

Self-hosting can be a big advantage for cost control, but it can be complicated too. Tembo.io's managed service provides privately hosted embedding models, but does not have hosted chat completion models yet.

chuckhend··on AI commented the entire Spring Boot codebase
A very kind response by someone from the project team
chuckhend··on Postgres as queue
some notes about pgmq, https://github.com/tembo-io/pgmq, that is on this list. It is built as an extension in Postgres, which makes it compatible with all languages that have a Postgres driver.

There's no 'magic' to it, it uses existing Postgres features so all the performance and consistency guarantees of Postgres are to be expected. Easily gets to 10k+ concurrent reads and writes even on smaller sized Postgres instances, which is more than most applications need.

chuckhend··on Transforming Postgres into a Fast OLAP Database
Clickbench is really nice because it is so easy to compare and contribute benchmarks. Is there anything out there like Clickbench, but for the TPC-DS and TPC-H benchmarks?
chuckhend··on Show HN: Visualize your Python code execution in VSCode
This is very cool. Great that it is integrated w/ vs code as an extension!
chuckhend··on Show HN: LLM Benchmarks Leaderboard with 60 model and API host combinations
I like that it shows price!

MTEB is also worth a look: https://huggingface.co/spaces/mteb/leaderboard

chuckhend··on Compiling Rust is testing
Compare this with a language like python -- you'll end up writing a lot of tests which do nothing but assert types of python objects. Compared to rust, the compiler is 'proving' that your code is set up with the right types. Switching from python to Rust, there's tests you don't have to write and you get immediate validation on that class of problems them via the compiler.
chuckhend··on PostgreSQL is enough
pgmq (which is linked on this gist) provides an api to this functionality. It can be 0 seconds, or 10 years if you want. It's not a row lock in, which can be expensive. In pgmq, its build into the design of the visibility timeout. FOR UPDATE SKIP LOCKED is there to ensure that only a single consumer gets any message, and then the visibility timeout lets consumer determine how long it should continue to remain unavailable to other consumers.
chuckhend··on PostgreSQL is enough
You get exactly once when you consume with pgmq and run your queue operations inside transactions in your postgres database. I can't think of an easy way to get some equivalent on SQS without building something like an outbox.
chuckhend··on Relativistic Spaceship
We'd still be limited by however many Gs human body can handle right?
chuckhend··on We built our customer data warehouse all on Postgres
There are definitely ways to cleanly make Postgres scale for analytics. We didn't discuss in this blog, but we will be writing about them in the future. For example, check out what the folks at ParadeDB are doing. https://github.com/paradedb/paradedb. Neon is doing an awesome job separating compute from storage. Supabase contributed foreign data wrappers make it super easy to read from S3 into Postgres. Lots of great work going out there :)
chuckhend··on OLMo: Accelerating the Science of Language Models [pdf]
It looks like this team gave us everything we need to reproduce their models, the actual artifacts needed to reproduce it. As far as I can tell, they share the data and every step along the way to final model...not just describing what they did.
chuckhend··on We built our customer data warehouse all on Postgres
There will be another blog from us at some point about running the data warehouse at scale. We're already working on integrating with s3 storage, and distributed compute is in the roadmap. Both possible today with open source extensions, and our friends in industry are already doing it.
chuckhend··on We built our customer data warehouse all on Postgres
That looks like an awesome tool. I am going to try it out. Gave you a star!
chuckhend··on We built our customer data warehouse all on Postgres
Check out what the folks at https://docs.epsio.io/ built. I'm not aware of any other projects doing this, but I think it's a great idea.
chuckhend··on OLMo: Accelerating the Science of Language Models [pdf]
There's some more commentary on their open-ness in this blog too https://www.interconnects.ai/p/olmo
chuckhend··on OLMo: Accelerating the Science of Language Models [pdf]
The training datasets are also available, which sets them apart a bit IMO.

https://huggingface.co/datasets/allenai/dolma

chuckhend··on Antarctic fungi survive Martian conditions on the International Space Station
My bad. It's still a very interesting read!
chuckhend··on Rust is replacing C as the backend to Python
That's how it starts!
chuckhend··on Tembo Operator: a Rust-Based Kubernetes Operator for Postgres
Postgres + containers + package management for extensions + stacks

For example re: extensions, you can add postgres extensions to your spec and the operator will handle getting the extensions installing extensions from pgt.dev into your Postgres instance.

chuckhend··on Benchmarking Postgres Vector Search Approaches: Pgvector vs. Lantern
Really great to see how the different config parameters (m, ef_construction, ef_search) impact latency and recall.
chuckhend··on An overview of distributed Postgres architectures
> Guideline: the durability and availability benefits of network-attached storage usually outweigh the performance downsides, but it’s worth keeping in mind that PostgreSQL can be much faster.

I think it often goes overlooked just how slow network attached block storage is though, and some organizations get very surprised when moving from an on-prem data center to cloud.

chuckhend··on Zero-ETL for Postgres: Live-query cloud APIs with 100 open source FDWs
These look awesome! Would love to see these get published at https://pgt.dev/
chuckhend··on Yeeting over 30k messages per second on Postgres with Tembo MQ
Thank you! We built PGMQ because we needed a queue to sit in between our managed service's control-plane and data-plane. We made it an extension so that it became a feature of the database, which meant we could use it with any programming language. IMO, this is a benefit of Postgres extensions that often gets overlooked...any language with a Postgres driver can use Postgres extensions.
chuckhend··on High-Scale AI/ML Feature Serving at Low Cost with Caching
I haven't used Tecton too much outside of POC, but I have worked with Feast quite a bit. Anyone able to compare/contrast Redis as an online store w/ Feast to Redis cache on Tecton? Are they the same thing?
chuckhend··on Show HN: I Wrote a Book on PostgreSQL for Rails
Congrats on the release!
chuckhend··on The Documentation System
Agree. When docs go on a tangent, or try to cover all these in a single doc, I think it's easy for the reader to get lost.

Can you recommend any other useful resources?

← PreviousPage 2 of 3Next →