Using Postgres for Everything
timescale.com
timescale.com
This idea, if true (and I believe it is) tells us how lucky we are compared than the past programmers. Advanced databases and languages are so advanced that they can do almost everything, even if they are not good at doing everything. And hardware and RAM is so generous today, that will mask a lot of potential issues unless you are at a big scale.
https://thmsmlr.com/cheap-infra
I'm not joking either.
High availability is not difficult to achieve with a single box, it's just depending on luck
I've had boxes run with 100% availability for multiple years
In the event some real failure happened, I had a clone of the AMI and could easily dump it onto new hardware - after I knew about it of course.
All that said - that situation wasn't my Job - and if Im being paid to have uptime, Id better cover my ass and have a hot failover on standby.
Since Im speaking AWS here - of course I used EBS for the virtual filesystem such that a machine crash didn't lose my SSD based filesystem. As I'm sure many here know, you can run a database on an EBS backed filesystem (and I have to assume even mount a swap file there too). I thought it was crack-smoking a when someone suggested that could easily be done.
"But if its that simple - why is everyone running all these N tier setups and living with all this unnecessary complexity?!?!"
It’s all bullshit. DevOps is really so complex that a medium-sized company needs an entire team?
Earlier today I was trying to set up some elaborate Postgres database backup and couldn’t get it to work so I just put the command in the cron tab and called it a day. Took 1 minute.
Depends on your goals, but most people probably don't need it.
> I just put the command in the cron tab and called it a day
Obviously this will lead to data loss if any data is written and then a corruption event or issue occurs in-between cron events. The road to over-complicated infrastructure is paved with mistakes and outages. The home-lab server running on a pi in my closet can use cron jobs and scripts, but at work we maintain higher standards because our data loss would cost more than a few extra days of engineer time.
Everything should be as simple as possible, but let's not pretend everyone is a chump because their engineering tolerances don't allow for data loss.
My 6th gen i5 (4 cores) can render ~15k page views per second with database queries, or several times that serving cached pages from nginx. 150 requests/second wouldn't be noticeable if I were using it for gaming at the same time.
A few days ago I asked our ops team, what's our rps? ~500. Out of curiosity, I ran a few benchmarks to stress test my monolithic pet project in Go with a single DB for everything and it easily handled 15k/s on my home computer, too, without any special optimizations. I understand its load is totally different from our production server, and my benchmark is most likely flawed, but the order of magnitude of difference is interesting.
I was told the production cluster also has additional 500 rps just for communication between the microservices. I suspect there's just too much overhead in their setup: communication between services, serialization of data between different DBs, etc. If it was a single modular monolith with a single DB, I suspect it could be faster and easier to maintain... Another issue is that they also sell it as an on-premise solution. And with the complexity growing, it now often fails to run on clients' infra.
In a ZIRP/startups environment this kind of BS was rewarded by the market and thus "engineers" would take any opportunity to build up such unproductive over-engineering skills to polish up their resume for the next opportunity.
> And with the complexity growing, it now often fails to run on clients' infra.
then I saw this and... yeah completely different beast
At my last startup (with ~3 technical employees), I was talking to the CTO about our cloud spend and planning for properly scaling infra after we got our first customer. We weren't "a SaaS website", but (hand-waving) we did complex physics modeling and analytics on data streams. The (physicist) CTO had heard of k8s and cloud stuff and wanted me to investigate how we should scale our cloud. He was convinced it'd be expensive to do "all those physics calculation with 64 bit numbers".
I showed him our entire cloud - gateway, database(s), physics modeling, metrics collection, log aggregation, etc - running at 1000x estimated 1-customer load from a MacBook. I offered to set him up a personal raspberry pi to stress test before a k8s cluster.
We ended up with a sensible single medium EC2 instance running everything, with some extra stuff for fail-over. AFAIK the only change made after I left was using a cloud-vendor DB.
Also batching database queries is important for high throughput.
ElectricSQL, SQLedge, pglite: https://news.ycombinator.com/item?id=38690588
pgreplay, pgkit: https://news.ycombinator.com/item?id=37959945
pg_timeseries: https://news.ycombinator.com/item?id=40417347
pgvector, postgresml: https://news.ycombinator.com/item?id=37812219
cmu ottertune, pg_auto_tune, postgresqltuner,
timescaledb-tune: https://github.com/timescale/timescaledb-tune
awesome-postgres: https://github.com/dhamaniasad/awesome-postgres
My theory is that every new database tech starts and a new domain specific db, then over time gets merged into the generic engines to serve the 99%.
I feel they were expecting a more complex solution with kafka queues.
He! Interviews don't go well for saying "just use postgres"
When PGlite (or SQLite) is used with Electric (https://electric-sql.com/), we already provide good reactive primitives. But we hope to improve this so that where possible the full queries don't have to be re-run.
There are several fantastic options that remove the need to manually write resolvers, set up subscriptions, solve N+1, reinvent the wheel, etc.
Plug notice: this is my company.
Part of me is thinking about htmx or something from alpine and wondering if it’s enough. Shoutout to livewire too.
Shaving complexity is a good goal, but is partially negated by adding it again closer to the core.
Anyway, maybe for most (in unique installations, or companies, not amount of servers dedicated) the workload won’t be heavy enough to deserve the extra complexity of handling many different database servers.
There are probably other things in your stack than just the database.
Uhhhh in theory... maybe... not really. Switching PG out from under our relatively toy-ish SaaS product would be a huge undertaking.
We actually have 100s of customers who use Timescale for their production time-series _and_ OLTP workloads :-)
- web server
- relational db
- a queue system for async processing
So, yeah two of such components can be handled by postgres (although I’d go for rabbit mq for async processing).
-- Ian Malcolm
James did put a lot of thought into this post. I saw multiple iterations of it before it was published. I think calling it a shitpost is being unkind to him.
I think we just find that some developers don't like reading long posts, so he kept this one short :-)
But if you want something with more length/depth, here is another one we recently published:
I didn't want to go down that path for this article because I didn't want it to end up as a sales pitch. The article you describe needs to be written, but it's a follow-up in my mind (and much more technical, for a slightly different audience).
One of the biggest concerns about any DB is speed of query execution. Today a decent app/site can generate terabytes of machine data that you need to analyze. There is a massive difference in usability of a query that takes < 5 secs vs even 5 minutes. So, some benchmarks of speed vs major use cases would have been ideal.
I understand, general purpose DBs can't be the best, but I can compromise for 80%. If it's 50% or less, then I really have to question that choice.
(You could still use PostgreSQL, just not the same PostgreSQL servers that are running your user-facing features.)
But also, point taken. The article this comment thread describes also needs to be written.
Orioledb, pg_analytics, hydra are getting there, but nothing out of the box postgres yet.
Compression also nearly non-existent except for TOAST.
I don't know if columnar engines will make it to core Postgres (I doubt it), but I do see the work of OrioleDB making it to core someday. We've joined an "alliance", alongside the core OrioleDB folks and some people from Timescale and Yandex, to push for it.
I have also some dream that one day we will see something like rocksdb/lsm backing postgres.
https://www.postgresql.org/message-id/315b7ce8-9d62-3817-0a9...