Heroku Postgres is now based on AWS Aurora
blog.heroku.com
blog.heroku.com
It manifested the way that the Aurora instances would use up their available (meagre) memory, then start thrashing, taking everything down. Apparently the instances did not have access to any temporary local storage. There was no way to fix that, and it took some time to understand. After having read all the little material I could find on Aurora, my personal conclusion is that Aurora is perhaps best thought of as a big hack. I think it's likely there are more gotchas like that.
We moved the database back to a simple VM on SSD, and Postgres handled everything just fine.
Luckily since November 2023 it also has r6gd/r6id classes with local NVMEs for temp files. [3] This should in theory solve this problem but I haven't tried it yet.
[1] https://docs.aws.amazon.com/AmazonRDS/latest/AuroraUserGuide...
[2] https://www.reddit.com/r/aws/s/sIhBQhsG80
[3] https://aws.amazon.com/about-aws/whats-new/2023/11/amazon-au...
When hitting one of these with a write you end up with massive delays. The 7 seconds and below tends to be from us-west-2 and the higher numbers are from our Japanese users.
Our OPS team has struggled to figure out why the delays happen. There’s some code fixes we could probably do (i.e always write to the writer) but as team lead for the development side the deadline is too close and I don’t want to rewrite core parts of the app to split reads and writes. They engaged AWS support so I’m hoping something is just misconfigured or maybe this just isn’t the use case for Aurora.
It sounds like you might be using global aurora with write forwarding? That’s pretty new and not something I have experience sorry. AFAIU though it’s a whole different thing under the hood.
Yes I believe this is what they chose. Honestly I’m going to leave it up to them and aws support. I have other fish to fry to get the functionality finished.
It has a couple of quirks, but on balance it feels like the future - the next evolution of what traditional rdbms are capable of.
But if you don’t have scale or resiliency needs it probably doesn’t matter to you.
Example: in normal MySQL, “RENAME TABLE x TO old_x, new_x TO x;” allows for atomically swapping out a table.
But since we moved to Aurora MySQL, we very occasionally get stuff land in the bug tracker with “table x does not exist”, suggesting this is not atomic in Aurora.
Is this documented anywhere? Not that I’ve been able to find. I’m fine with there being subtle differences, especially considering the crazy stuff they’re doing with the storage layer, but if you’re gonna sell it as “MySQL compatible” then please at least tell me the exceptions.
However, Aurora isn't cheap and is at least ~80% of our monthly AWS bill. I wonder how it is cheaper than Heroku's previous offerings? Is it Aurora Serverless v2 or something like that to reduce cost? Aurora billing is largely around IOPS, and Heroku's pricing doesn't seem to reflect that.
https://docs.aws.amazon.com/AmazonRDS/latest/AuroraUserGuide...
Just wondering, why is that?
Xata is (like Heroku) based on Aurora, but offers database branching and has different pricing. That should be ideal for lightly-used test instances, because you only pay for storage, and 15GB are included in the free tier.
Why don't you run your own Postgres?
It's not hard - why pay such a premium for the Amazon version?
It's not as difficult as a PhD (I assume; I only got as far as an MS), but based on what I've witnessed, it's up there in complexity. There are dozens of knobs to turn – not as many as MySQL/InnoDB to be fair, but still a lot – things you have to know before they matter, etc.
> People do it because there's limited time and no business advantage to operate postgres clusters. Use the time on what your business actually does.
I've seen this argument countless times for SaaS anything. I don't think it's accurate for a database. Hear me out.
For most companies, the DB is the heart. Everything is recorded there, nearly every service's app needs it (whether it's a monolith or micro service-oriented DBs), and it's critically important to the company's survival. Worse, the same skills necessary for operating your own DB generally overlap heavily with optimally running a DB, by which I mean if you're good at things like DB backup automation, chances are you're also good at query optimization, schema design, etc.
It's that latter part that seems to be missing from many engineering orgs. "Just use Postgres," people say; "just add a JSONB column and figure out the schema later," but later never comes. If your business uses a DB, then you do not have the luxury of running one poorly. Spend a few days learning SQL, it's an easy language to pick up. Then spend a few days going through the docs for your DB, and try the concepts out in a test instance. Your investment will be rewarded.
All I'm saying is it has nothing to do with difficulty. In my job for example we self-hosted HBase which is a beast compared to postgres, implemented custom backups etc, all because there was no good vendor for it. Postgres is much simpler and we always just used RDS and then switched to Aurora for the higher disk limits when it was launched. If there's a good enough vendor, you're just stroking your ego re-implementing these things when you could move on to the actual thing the business wants to release.
I've also seen senior engineering leads "proving" self hosting "saves money" but then 2 companies working on the same type of problem in the same industry with a similar feature set, on one side we had 5 people maintaining what on the other company it took 6 teams of 4-8 people. So it depends if you'd like to have a lot of your labor focused on cutting costs or increasing revenue. And they never include the cost of communicating with extra 5 teams and the increased complexity and slowness to release things this creates, while also being harder to keep databases with current versions, more flimsy backup processes, etc.
Ps: we got rid of hbase, do yourself a favor and stay away
Common sense isn't so common. I've met a handful of devs across many separate companies who care at all how the DB works, what normalization is, and will read the docs.
> It doesn't mean you should spend time implementing your own custom backup process with ability to go back to a specific point in time, configurable in 1 minute.
If by implement you mean write your own software, no, of course not. Tooling already exists to handle this problem. Off the top of my head, EDB Barman [0] and Percona XtraBackup [1] can both do live backups with streaming so you can backup to a specific transaction if desired, or a given point in time.
Or, if you happen to have people comfortable running ZFS, just snapshot the entire volume and ship those off with `zfs send/recv`. As a bonus, you'll also get way more performance and storage out of a given volume size and hardware thanks to being able to safely disable `full_page_writes` / `doublewrite_buffer`, and native filesystem compression, respectively.
> If there's a good enough vendor, you're just stroking your ego re-implementing these things when you could move on to the actual thing the business wants to release.
Focusing purely on releasing product features, and ignoring infrastructure is how you get a product that falls apart. Ignoring the cost of infrastructure due to outsourcing everything is how you get a skyrocketing cloud bill, with an employee base that is fundamentally unable to fix problems since "it's someone else's problem."
> Ps: we got rid of hbase, do yourself a favor and stay away
HBase and Postgres are not the same thing at all. If you need the former you'll know it. If people convince management that they do need it when they don't, then yeah, that's gonna be a shitty time. The same is true of teams who are convinced they need Kafka when they really just need a queue.
My overall belief, which has been proven correct at every company I've worked at, is that understanding Linux fundamentals and system administration remains an incredibly valuable skill. Time and time again, people who lack those skills have broken things that were managed by a vendor, and then were hopelessly stuck on how to recover. But hey, the teams had higher velocity (to ship products with poor performance).
[0]: https://pgbarman.org
[1]: https://www.percona.com/mysql/software/percona-xtrabackup
However, at large scales cloud won’t make sense anymore. They do have a markup and eventually what you’re paying in markup could instead buy you a few full time employees.
Yes, many times, which is why I've developed this opinion.
> However, at large scales cloud won’t make sense anymore. They do have a markup and eventually what you’re paying in markup could instead buy you a few full time employees.
The issue is once you've finally realized this stuff matters, and have hired a DB team, I can practically guarantee that your schema is a horror show, your queries are hellish, and your product teams have neither the time nor inclination to unwind any of it. Your DB{A,RE}s are going to spend months in hell as they are suddenly made the scapegoats for every performance problem, and are powerless to fix anything, since their proposals require downtime, too much engineering effort, or both.
Hence my statement. Learn enough about this stuff so that when you do hire in specialists, the problems are more manageable.
All of the troubles you described sound like bad management. I’m sorry if you’ve had to go through that. DBAs that are setting up a replacement are going to need time to do that right and expectations need to be set that this is a tricky problem.
I'm migrating a 1tb database to it right now because I'm paying too much for iops even on the regular rds postgres.
I'm also quite sure this is what heroku must be using or they would be out of business because of the pricing mismatch.
Product Storage Max Connection Monthly Pricing
Essential-0 1 GB 20 $5
Essential-1 10 GB 20 $9
Essential-2 32 GB 40 $20
The pricing looks quite competitive, although I'm not sure what the prior rates were.10 years ago I spent 10x+ per month for 32GB (RAM) Heroku Postgres instances, IIRC they were around $400/mo, maybe even more.
Aren't you comparing RAM vs Storage there? The pricing chart here says nothing about RAM.
We're expanding the Aurora-backed offerings to include larger dedicated DBs in the relatively near future as well.
Gail Frederick our CTO talked a bit more about it at high-level during re:Invent 2023: https://www.youtube.com/watch?v=fZLcv7rwj7Y&t=1955s
A VPS (hetzner, etc) + managed postgres DB (supabase / AWS / etc) or a local one might more more than enough these days.
If you wanted to spend money on something else, circleci and others help manage ci/cd as well.
Despite interesting competition, my feeling is that the Heroku of 2024 remains... Heroku.
I feel this way even though -- depending on how you segment -- the list of "interesting" competitors is quite long at this point: Render, Railway, Northflank, Fly.io, Vercel, DO App Platform, etc.
It's crazy how the ergonomic still just aren't there.
(And: bugs. I'm also surprised by the kinds of issues I run into on some of those sites in my list -- problems that, even if not show-stopping, feel like revealing indicators of quality.)
Any specifics/examples? I find it hard to imagine those "big name" companies/platforms you just mentioned don't have entire teams dedicated to hyper-optimizing experience.
We were getting close to one of the big jumps on the standard pricing of Heroku Postgres, and we would have had to basically double our monthly cost to lift the max data we could store from 1.5TB to 2.0TB. On Crunchy Data, that additional disk space will be like 1% more rather than 100% more.
While investigating Crunchy, I ran some benchmarks, and I found Crunchy Bridge Postgres to be running 3X faster than Heroku Postgres.
Heroku seems to be working on some interesting new things, but I feel burned by the subpar performance and lack of basically any new features over many years. I don't know if the new Aurora-based database will be faster than Crunchy, but the benchmarks they're talking about sound like they're finally about to catch them. But we also have better features on Crunchy, too, such as logical replication. Logical replication is still not available on Heroku.
The experience for deploying apps and having add-ons is still pretty easy, but we'll see how that improves. HTTP2 support is still in beta.
I settled up on AWS ECS :)
My main issue with Heroku was that they have not changed anything in _years_. No support for gRPC, no IPv6, and simple VPC peering costs $1200 a month.
They just shipped HTTP/2 terminated at their router [0], and have it on their roadmap [1] to support HTTP/2 all the way through. But it seems like it's at minimum a few months off.
(As for VPC peering: the moment you need that, it sorta feels like Heroku is no longer the right place to be, even ignoring the costs.)
[0] https://blog.heroku.com/heroku-http2-public-beta [1] https://github.com/orgs/heroku/projects/130
I also moved my app hosting to NorthFlank from Heroku and have been really happy with that as well. It’s got the features I always wanted on Heroku (simple things like grouping different types of instances together into projects really helps) plus again excellent responsive support.
Would strongly recommend them to anyone looking to move off Heroku.
I got to talk to someone who was intelligent about Postgres, who answered various questions I had, who offered a few pieces of insight that I wouldn't have thought, etc.
Compared to every single support interaction I've ever had with Heroku for 10 years, and this was light years more friendly, informative, and productive.
I am so happy we're switching. Way to go Crunchy!
That was very surprising to me. Most businesses that are willing to pay 2x for an HA database are probably NOT likely to be ok with that kind of data loss risk.
(AWS and GCP's HA database offerings use synchronous replication.)
But I spent several hours fighting with a DNS change, trying to host my marketing website as Cloudflare pages site from my root domain (with DNS managed by Cloudflare), and then wildcard subdomains routed to a Render server. I couldn't get it to work no matter what configs I tried. My root domain marketing site is proxied through Cloudflare and I was trying to get the wildcard subdomains as DNS-only, and I suspect this was the problem but idk. In other words, the Cloudflare pages marketing site is https://bookhead.net and I wanted my customer's subdomains to route to the Render server like https://forlornbooks.bookhead.net/ (I still get the error since I haven't finished my migration to Heroku). The subdomains worked with no problem until I tried to setup a separate marketing site at the root domain.
Also, I had a hard time setting up SSH with a containerized server. It was a weird DX that was a bit confusing to document so I can remember later. Can only imagine how confusing it might be if I ever have teammates. The Render CLI looks promising, though.
These are only the most recent issues. Seems like y'all improved the headaches I ran into the time I tried.
That said it does feel a bit like a ghost town, I'm always happy to hear when someone is doing something over there.
Weird for Heroku to ignore this huge efficiency opportunity.
Aurora is one of the few real innovations in the database space recognized by SIGMOD: https://sigmod.org/sigmod-awards/citations/2019-sigmod-syste...
It provides a lot of benefits to the user and also a ton more to the service provider. Specifically you don’t overprovision storage or compute. Plus at least theoretically you can provide invite IO throughput at the storage level.
There has been a couple more iterations on the design since. Microsoft separated transaction log from storage: https://www.microsoft.com/en-us/research/uploads/prod/2019/0...
Neon added an object store and branches so you can integrate backups and add a Time Machine.
PolarDB separated memory from compute - this makes serverless compute more nimble and unties memory and CPU.
EDIT: and yes, it's not cheap
Never recommending them to anyone anymore
Currently our Aurora instances are in private beta. If you're interested in trying it out, drop me an email: richard@xata.io
I am not a fan of Aurora. I don’t get the appeal at all. I’ve tried MySQL and Postgres varieties; it’s just expensive for no reason.
Aurora splits out the compute and storage layers; that's its secret sauce. At an extremely basic level, this is no different from, for example, using a Ceph block device as your DB's volume. However, AWS has also rewritten the DB storage code (both MySQL/InnoDB and Postgres). InnoDB has a doublewrite buffer, redo log, and undo log. Postgres has a WAL. Aurora replaces all of this [1] with something they call a hot log. Writes enter an in-memory queue, and are then durably committed to the hot log, before other asynchronous actions take place. Once 4/6 storage nodes (which are split across 3 AZs) have ACK'd hot log commit, the write is considered persisted. This is all well and good, but now you've added additional inter-process latency and network latency to the performance overhead.
Additionally, the storage scaling I mentioned brings with it its own performance implications. If you're doing a lot of writes, you'll encounter periodic performance hits as the Aurora engine allocates new chunks of storage.
Finally, even for reads, I do not believe their stated benchmarks. I say this because I have done my own testing with both MySQL and Postgres, and in every case, RDS matched or beat (usually the latter) Aurora's performance. These tests were fairly rigorous, with carefully tuned instances, identical workloads, realistic schema and queries, etc. For cases where pages have to be read from disk, I understand the reason – the additional network latency of the Aurora storage engine seems to be higher than that of EBS. I do not understand why a fully-cached read should take longer, though.
As a further test, I threw in my quite ancient Dell servers (circa 2012) for the same tests. The DB backing disk was on NVMe over Ceph via Mellanox, so theoretical speeds _should_ be somewhat similar to EBS, albeit of course with less latency since everything is in a single rack. My ancient hardware blew Aurora out of the water every single time, and beat or matched RDS (using the latest Intel instance type) almost every time.
[0]: Arguably, it's also better at globally distributed DB clusters with loose consistency requirements, because it supports write forwarding. A read replica in ap-southeast-1 can accept writes from apps running there, forward them to the primary in us-east-1, and your app can operate as though the write has been durably committed even though the packets haven't even finished making it across the ocean yet. If and only if your app can deal with this loosened consistency, you can dramatically improve performance for distant regions.
[1]: https://d1.awsstatic.com/events/reinvent/2019/REPEAT_Amazon_...
Anyway, that's also ignoring the features that Aurora offers, which is why people pay more for it. The ability to have multi-AZ deployments and auto-scaling of (what can be cross-region) read replicas make it very resilient and it's dead simple to operate what would normally be considered advanced features of a DB cluster.
If you just need a managed Postgres or MySQL traditional single instance and none of those extra features, then obviously you would not need to pay the premium for Aurora. RDS exists for that reason.
RDS Multi-AZ Cluster gives you much of the advantages of Aurora, but with higher performance and more tuning capabilities, though you are limited to 3 nodes. Tbf 3 nodes is almost certainly enough for most companies. A few hundred thousand QPS would be easily handled by that.
Re: cross-region read replicas, eh… if you’ve somehow managed to ensure that every single aspect of your app is capable of withstanding the loss of an entire region – including us-east-1, since most of the control plane functions are there – then sure, maybe. But do you need it? If an entire AWS region drops out, half of the internet goes with it, and you can just blame that. I doubt the small possibility of higher uptime is worth the literal doubling in monthly costs.
To put it simply, over the past two days, I attempted to deploy a full-stack assignment on AWS services. The front end was written in React, using Vite as the framework. For such Single-Page Apps (SPAs), I personally prefer using specialized services like Netlify or Cloudflare Pages for deployment, as these services offer very robust CI/CD services, allowing for one-click deployment and automatic updates, saving a lot of hassle.
Initially, I planned to manually deploy on AWS using the S3 + CloudFront model (since it was just a one-time assignment), but later I discovered that AWS has a service very similar to Netlify called Amplify, which also offers CI/CD one-click deployment services. Amplify goes even further by including user directory services, allowing for one-click registration and login via related components.
It sounds great, but you only realize how problematic it is after using it. After some research, my initial deployment method was to upload the code to GitHub and then click the deploy button in the Amplify interface. This is also the deployment method I use most often with Netlify.
However, I later found something wrong. The key issue was that applications deployed this way using Amplify couldn't directly use Amplify's UI components to access Cognito user directory services. After much searching, I found that Amplify has an Amplify CLI initialization command to create a new CI/CD project in the Amplify service, which also deploys additional resources like Cognito.
It seemed feasible, so I did it. Then I found some issues. The initial "issues" were just on the AWS account management level: after deploying the project via Amplify CLI, my AWS account quickly filled up with a bunch of "things"—the reason "things" is in quotes is that Amplify created a lot of fragmented resources, including but not limited to CloudFormation, IAM roles, etc., even creating two Cognito identity pools for me—it's hard not to call them "junk." Moreover, most of these resources have names that are impossible for humans to remember or distinguish, and there are no explanations or grouping features to tell you what these things are for.
If it were just like this, it wouldn't seem to impact the development process, right? The biggest problem is that the local debugging and production environment apparently don't use the same configuration files, and when I was cleaning up the automatically created resources in my AWS account earlier, I somehow deleted the roles calling the Cognito user pool in the production environment, causing the production environment to be unable to access the two user pools created by Amplify, constantly throwing 400 errors.
After several rounds of "deploy-delete-redeploy-redelete," I decided to start over and look for the related documentation again. Later, I found that Amplify has a set of documentation outside of AWS's own documentation system, and this documentation recommends a deployment method: clicking the deploy button on the GUI webpage—yes, you heard it right, the same deployment method I used initially.
So, how do you deploy additional components/services like Cognito this way? Amplify's answer is configuration files. As long as you create a folder for configuration files in the root directory of your project and write the corresponding configuration files in it, the cloud will automatically create the resources you need in AWS once it reads them.
It sounds reasonable, right? Then you go to find the part about configuration files in the documentation... What's going on? Why can't I find anything in the search box in the documentation? There's not even a sample configuration file! Algolia indexing service can't be this bad, right?
Searching for "defineAuth" in the Amplify official documentation returns mostly irrelevant information.
Is my search method incorrect? I entered keywords like "site .amplify.aws defineAuth" in the Kagi.com search engine but couldn't find any examples or explanations of configuration file items. At this point, I'm completely convinced that the Amplify documentation is garbage. Fortunately, the API documentation of the Amplify framework is quite good, at least reducing my urge to buy a ticket to the US and blow up Amazon's headquarters while guessing the configuration file items...
Also, Amplify has a UI that is completely different and more modern than other AWS services. The discrepancy is still a minor issue; the main problem is that if you create a project using the (slightly outdated) Amplify CLI, and then try to configure the back-end services like Cognito it deployed on the webpage, you'll enter an old interface. That is, once you click in, you see a slightly ugly but familiar interface, yet it feels completely disconnected from the previous Amplify interface...
So now I understand why I hadn't heard of Amplify before—it's really hard to use. Complete integration is indeed an advantage, but even being born with a silver spoon doesn't excuse Amplify's messiness, simply throwing everything together and telling users "it just works." Users look at it, wondering what on earth all these things are, and then you hand them a manual that looks fancy but has zero information. Users, flipping through this tome with no useful information, can only throw this pile of stuff into the historical junk heap behind them in frustration.
Some rights reserved Except where otherwise noted, content on this page is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International license.
ENJOY
READ THIS Short story: "The Cry of Accelerated Demise" in a place accelerating towards demise.
Says who? My experience is the opposite - tending towards too much reliance on the main providers because of the credits
And once those credits run out we are planning to expand our owned training hardware. Currently we just have 3x L40S but would expand to 32x L40S. I’m excited to now be a sys admin in addition to a full stack web dev.