Running a database on EC2? Your clock could be slowing you down
heapanalytics.com
heapanalytics.com
It makes some sense with Sql Server and Oracle in a few cases because of licensing but hosting your own Postgres instance on AWS is the worse of both worlds -- you're paying more than with a cheaper VPS and you have to do all of the maintenance yourself and not taking advantage of all of things that AWS provides -- point in time restores, easy cross region read replicas, faster disk io (Aurora), etc.
Maybe RDS is already configured correctly. Maybe not. But at a certain scale not having your hands tied becomes more valuable than having everything taken care of for you.
Note that these are educated guesses from their statement: "Heap’s data is stored in a Postgres cluster running on i3 instances in EC2. These are machines with large amounts of NVMe storage—they’re rated at up to 3 million IOPS, perfect for high transaction volume database use cases"
1. Aurora charges per-request ($0.20/million). Given that analytics comes with tons of events and that they wanted servers that have up to 3 million IOPS, it can get pricey fast.
2. RDS has database instances that have SSDs that provide "up to 40,000 IOPS" per instance in their provisioned case, which is probably not enough.
If you are architecting everything on AWS trying to avoid “lock-in” you’re going to move slower and pay more - the worse of both worlds.
Neither costs a fixed, single amount. Both are ranges that depend on the competence of the practitioner. For example, it's cheaper to use a reserved EC2 instance for 1 year than just on demand. Similarly, the default cost for DIY is much higher than the lower bound of the range.
If DIY includes "enterprise" hardware, software, and/or support contracts (which arguably contradicts the "Y" in DIY), then the top end of that cost range is easily above the top end of the AWS cost range. In this way, those people are correct.
However, if we limit it to truly doing it yourself [1] and with commodity hardware, software, and methods, that's no longer true.
More importantly, as the GP mentioned, "If you're going for cost savings", it's the bottom end of the cost ranges that are key, and, even at modest scale, AWS is way above DIY. In this way, those people are incorrect.
[1] There's a limit, of course. For me, that limit is commodity versus custom. It's safe to outsource server manufacturing (all the way up through assembly, initial burn-in, and even rack/stack), since, it's easy enough to replace with a different vendor. It's not safe to outsource specifying what goes into those servers or final "smoke" testing. Sometimes, even for something that's safe to outsource, like remote-hands replacement of failed parts, it may not be worth the price.
Also, you can’t replicate out of RDS. I like to know where my data is and how to bring it back online during a disaster.
What’s more likely. That your one data center has a disaster or a globally redundant AWS infrastructure.
That's not to say that it's impossible to do. Facebook ably achieves it. It's just that the level of expertise that would be required across so many services is significant. AWS has hundreds of services, each of which would need to be hiring highly skilled DBAs to handle the sharding etc. etc. etc necessary to scale. It's easier to just point people at DynamoDB where they've effectively handled all those needs for you, you just have to put a bit more logic in your application side, which also has the neat property of scaling horizontally more easily.
I do this presently because I have some custom stuff I do for MySQL which needs it's own EC2 instance because RDS doesn't support it.
My gripes with the performance issues still stand though. I have queries that take 30-40 seconds on RDS that complete in milliseconds on an EC2 instance which is much smaller.
But who are we kidding... it's impossible to resist and its why AWS rakes in cash.
That said, not being able to use vDSOs for time querying APIs isn't just a problem for databases, it's potentially a significant problem for asynchronous software, like Node.js, that do userspace event scheduling bookkeeping. Typically every iteration of the event loop performs at least one time query, but depending on how the software is coded often there might be one or more time queries per event processed per event loop iteration.
Remember (not you but the people who made your comment dead) AWS aswell as every other provider want to lock you in.
I say this as someone who uses AWS. But their are things that still tick me off with the platform. As an example They don't make it easy to RDNS a light sail instance verses an EC2 instance so you want to RDNS to help with that outgoing email server you want to set up its easier to pay for a micro ec2 then use lightsail for the same purpose (which comes with included bandwidth and its an outgoing email server so its not like CPU is a major issue) or use SNS, but its more beneficial to AWS for you to use ec2 or sns even if its not to you because of your end of month bill.
Anyways my point is its possible to resist the AWS lock in with a bit of forward thinking as long as you code for the possibility that you might want to swap providers.
It's how the mainframe and Windows ecosystems worked. Kudos to AWS for figuring out how to capture the exploding market for Linux- and OSS-dependent stacks.
But they provide management, point in time recovery, and other nice things. All done for you automatically. If you move you have to take that on yourself (if you don’t move to another provider that does it too).
Citus has been a huge part of our scaling story from RDS (maxed out instances and did a ton of tuning) to Citus Cloud. I can only imagine that Heap has a ton more problems than we've experienced.
Their is nothing special you do to support Postgres RDS.
https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/CHAP_...
They solve it with a large cluster of Postgres server running fast disks with partitioning handled by the Citus extension, along with several low-level tweaks.
RDS does not support these scales or any clustering. Aurora does not have parallel processing. Redshift would be fast but is very expensive and does not have the same level of Postgres features.
There is Citus Cloud so you can get close to Heap's setup with Citus maintaining it all, but that gets pricey too.
(I work at AWS but not RDS/Aurora)
[1] https://aws.amazon.com/blogs/aws/new-parallel-query-for-amazon-aurora/The post is specifically about our Citus cluster, which stores the analytics data for all our customers. Most of the reasons we do this have been given by other folks in the replies:
* RDS doesn't support the Citus extension
* data is stored on ZFS for filesystem compression
* we get significantly higher disk performance from these
instances' NVMe attached storage, which isn't available for
RDS1) price: you’ll easily spend $$$$$$ on RDS. If you host it on something equivalent with native SSD you’re looking at $800 a month with better performance
2) performance. It’s way faster and you can tune your indices, create views and make it fast and efficient in a predictable way.
3) if you want to, you can easily migrate to a different service. We did just that two months ago, from google cloud to AWS. It gives us vendor independence.
Reasons for moving:
1) Google had a mean bug that dropped long standing connections within their own network, and they blamed us despite following all the guides (linger timeouts etc). So had to move anyways.
2) The biggest reason was that we were given free credits by AWS.
Cost - Our primary data store has >1 Petabyte of raw data stored across dozens of Postgres instances. The amount of data we store is at the point where RDS is too expensive for us. The cost of an instance on RDS is more than twice the cost on EC2. For example, an on-demand r4.8xl on EC2 instance costs $2.13 an hour, while an RDS r4.8xl costs $4.80 an hour.
Performance - The only kind of disk available on RDS is EBS. EBS is slow compared to the NVMe the i3s provide. We used to use r3s with EBS and got a major speedup when we switched to i3s. As a side note, the cost of an i3 is also less than the cost of an r3 with an equivalent amount of EBS.
Configuration - By using EC2 we can configure our machines in ways we wouldn't be able to if we used RDS. For example, we run ZFS on our EC2 instances which compresses our data by 2x. By compressing our data, we get a major cost saving and a major performance boost at the same time! There isn't an easy way to compress your data if you use RDS.
Introspection - There are times where we've needed to debug performance problems with Postgres and EXPLAIN ANALYZE won't suffice. A good example is we used flame graphs to see what Postgres was using CPU for. We made a small change that resulted in a 10x improvement to ingestion throughput. If you are curious, I wrote a blog post on this investigation: https://heapanalytics.com/blog/engineering/basic-performance...
You can also run get bare metal I3 instances by launching the "i3.metal" instance type. You don't need to wait for the Nitro hypervisor, you can go with no hypervisor at all.
Since you work at Amazon, do you have a sense of big of a difference there is in performance between i3 and i3.metal for database workloads like Postgres?
As for stability, have been two major sources of instability with ZFS:
The first issue was with the default value of arc_shrink_shift. By default, ZFS will evict ~1% of ARC, the in memory file cache, to disk at a time. Our machines have several hundred gigs of ARC, so ZFS was evicting several gigs of data to disk at a time. This was causing our machines to frequently become unresponsive for several seconds.
The other issue is for some reason ZFS will lock up for long periods of time if we delete several hundred gigs of data. We haven't been able to identify a root cause of the problem. So far we've worked around this problem by adding a sleep in between data deletions.
Other than these problems, ZFS has worked pretty well for us.
How do you manage this?
Also, how frequently do i3 instances fail?
Over the course of a month, we usually have about one machine fail.
Running costs: Plain EC2 DB is cheaper than RDS. RDS is instance costs plus RDS tax.
Configuration and maintenance cost: RDS obviously cheaper
* Performance As other replies have detailed plain EC2 with NVMe etc could be significantly faster.
* Convenience RDS wins most of the time.
* Risk RDS wins unless heavy investment
So it depends on your requirement. If high performance with massive DBs is important, then the costs of managing your own DBs may make sense.If a low throughput, less risky DB then your own DB may make sense.
If a normal business use DB, outsourcing the maintenance and risk to RDS may make sense.
At the time, we were trying to benchmark disk I/O for new platforms, but we found that things were underperforming compared to the specifications for the hardware. We figured out that fio was reading the clock before/after each I/O (which isn't really necessary unless you really care about latency measurement) and just by reading the clock we were rate limiting our I/O throughput. By switching to "clocksource=tsc" in our fio config, we managed to get the performance behavior we expected.
can you put this into roughly quantitative terms? How much of a performance hit did you remove this way?
I remember with 8 disks that should have been able to do 60K 4K IOPS each (early SSD models), we were capping out at 90K IOPS with all disks in parallel at a queue depth of 32 while reading from CLOCK_MONOTONIC. When we switched to TSC I think we ended up getting around 320K IOPS. Still not perfect, but we were also capped by the particular HBA we chose (which didn't have multiqueue support).
If you're interested in clocks on Linux, you might also find this article useful (shameless plug): http://btorpey.github.io/blog/2014/02/18/clock-sources-in-li...
Huh? That’s definitely not true now, and I don’t think it ever was. Linux uses LFENCE or MFENCE, depending on CPU.
https://software.intel.com/sites/default/files/managed/39/c5...
Is it a good idea for a production database to depend on a feature not being used when the vendor hasn't said that they don't or won't use it? They may very well live-migrate when convenient, but just don't expose that functionality to customers since they don't want customers demanding it.
One way in the case of a database could be a second EC2 instance configured as a read replica in a different AZ.
But yeah, you need to be ready for your disks to go away no matter where they are: ephemeral, EBS, physical, whatever.
Said more directly: no, it is not ephemeral. It is local storage that is tied to the life cycle of the instance.
The "ephemeral" term is a legacy. Unfortunately it is part of the EC2 API for the block device mapping [1] of the "classic" instance store interfaces on the Xen platform. I don't know exactly when we stopped using "ephemeral" in our documentation, but I think it was with the introduction of EBS around 2008.
The "ephemeral" term confuses a lot of customers, and that's why we stopped using it. Data written to local storage is not transient, fleeting, or short lived. By 2010 we had transitioned to using "instance storage" in the documentation [2], which included a big note about how the data remains if an instance reboots for any reason (planned or unplanned).
Still, there is a misconception that data on local instance store volumes (both the more "classic" HDD or SSD volumes that are virtualized by Xen, as well as the new generation of local NVMe storage) could vanish due to this vestigial term that lingers in the API. Many customers, as well as services like Amazon Aurora [3], build highly durable and available systems on local instance storage.
[1] https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/block-de...
[2] https://web.archive.org/web/20111113011016fw_/http://docs.am...
[3] https://www.allthingsdistributed.com/files/p1041-verbitski.p...