And yes, Andrew Kane (et al) are the people to thank for pgvector.
We (Tiger Data) developed pgvectorscale and pg_textsearch (and timescaledb, and some others)
1,573 karma · joined October 16, 2010
ajay (at) tigerdata (dot) com
And yes, Andrew Kane (et al) are the people to thank for pgvector.
We (Tiger Data) developed pgvectorscale and pg_textsearch (and timescaledb, and some others)
Feels like a net negative for the HN community.
Why not let your audience decide what it wants to read?
I say this as a long time HN reader, who feels like the community has become grumpier over the years. Which I feel like is a shame. But maybe that's just me.
please ask your RDS rep to support it
we (tiger data) are also happy to help push that along if we can help
We just launched a bunch around “Postgres for Agents” [0]:
forkable databases, an MCP server for Postgres (with semantic + full-text search over the PG docs), a new BM25 text search extension (pg_textsearch), pgvectorscale updates, and a free tier.
We think we're still building great things, and our customers seem to agree.
Usage is at an all-time high, revenue is at an all-time high, and we’re having more fun than ever.
Hopefully we’ll win you back soon.
ClickHouse was fast but required a lot of extra pieces for it to work:
Writing data to Clickhouse
Your service must generate logs in a clear format, using Cap'n Proto or Protocol Buffers. Logs should be written to a socket for logfwdr to transport to PDX, then to a Kafka topic. Use a Concept:Inserter to read from Kafka, batching data to achieve a write rate of less than one batch per second.
Oh. That’s a lot. Including ClickHouse and the WARP client, we’re looking at five boxes to be added to the system diagram.
So it became clear that ClickHouse is a sports car and to get value out of it we had to bring it to a race track, shift into high gear, and drive it at top speed. But we didn’t need a race car — we needed a daily driver for short trips to a grocery store. For our initial launch, we didn’t need millions of inserts per second. We needed something easy to set up, reliable, familiar, and good enough to get us to market. A colleague suggested we just use PostgreSQL, quoting “it can be cranked up” to handle the load we were expecting. So, we took the leap!
PostgreSQL with TimescaleDB did the job. Why overcomplicate things?I’ve learned (sometimes the hard way!) that every design choice comes with real trade-offs. There’s no magic database architecture that optimizes every dimension (e.g., scalability, performance, ease-of-use) simultaneously.
Social media often pushes us into oversimplified "winner vs. loser" narratives, but this hides the actual complexity of building great infrastructure.
Recognizing and respecting these differences makes us smarter engineers, better community members, and frankly, just more enjoyable people to chat with.
PS Thank you for helping me add a new book to my list :-)
"ClickBench evaluates databases using a single table of clickstream data, representative of workloads like web analytics, BI, and log aggregation. It also favors full-table large scans and large-scale aggregations on denormalized data.
Real-time analytics inside applications is different and needs a new benchmark." [0]
This is why we published RTABench. [1]
We believe that it is more representative of real-time analytical workloads.
[0] https://www.tigerdata.com/blog/benchmarking-databases-for-re...
"The future is already here, it's just not very evenly distributed" - William Gibson
Also TigerBeetle is an insect, not a fast cat.
https://www.timescale.com/blog/how-we-scaled-postgresql-to-3...
(Post is a year old, IIRC the database is over one petabyte now)
And yes you are correct, pgvectorscale scales pgvector for embeddings, and pgai includes dev experience niceties for AI (eg automatic embedding management).
Would love to hear any suggestions on how we could make this less confusing. :-)
So we would recommend using both from the start. There is no cost (technical or financial) for doing so.
There is discussion about getting this added to AWS RDS (as well as other PostgreSQL providers), but too early to share anything.