We saved millions in SSD costs by upgrading our filesystem
heap.io
heap.io
- Total storage usage reduced by ~21% (for our dataset, this is on the order of petabytes)
- Average write operation duration decreased by 50% on our fullest machines
- No observable query performance effects
It seems like the better compression ratio and resulting reduced IO more than makes up for increased CPU compared to lz4. I wish they had mentioned the actual effect on CPU.
Compare to the recent thread "The LZ4 introduced in PostgreSQL 14 provides faster compression" [0] where the loudest voices were saying that zstd would not work due to increased CPU. This is a different layer (filesystem compression vs db compression), but this article represents an interesting data point in the conversation.
Obviously you need a LOT of CPU to throw at xzip if you want to use it.
zstd is very much more optimized for compression at speeds comparable to traditional gzip.
I use xzip primarily for things that will get compressed to long term storage and the time to create the archive isn't a really important factor.
in this test: https://sysdfree.wordpress.com/2020/01/04/293/
zstd level 19 wins on time vs. xz levels 5 through 9, but the xz ultimate compressed file size is definitely smaller.
Perhaps better than stepping to a different compression algorithm, zstd has multiple levels of compression that might be used at different times. The advantage there is that the same decompression algorithm works for all.
One might reasonably hope that decompression tables may be shared amongst multiple of the 64k raw blocks, to further squeeze usage.
Worth noting that part of the reason this is relatively low impact for our read queries is that the hot portion of our dataset is usually in Postgres page cache where the data is already decompressed (we see a 95-98% cache hit rate under normal conditions). We've noticed the impact more for operations that involve large scans - in particular, backups and index builds have become more expensive.
For backups in particular, are ZFS snapshots alone not suitable to serve as a backup? Is there something else that the pg backup process does that is not covered by a "dumb" snapshot?
That being said, wal-g has worked well enough for us that we haven't put a ton of time into investigating alternatives yet, so I can't say for sure whether snapshots would be a better option.
Also, auto-expiry of no-longer-needed WAL segments (that we use due to our reliance on async hot standbys) along with previous backups is pretty great.
And we haven't even started taking advantage of pgBackRest's ability to do incremental restore — i.e. to converge a dataset already on disk, that may have fallen out of sync, with as few updates as possible. We're thinking we could use this to allow data science use-cases that would involve writing to a replica, by promoting the replica, allowing the writes, and then converging it back to an up-to-date replica after the fact.
One tiny example: I prefer to work with databases using CLI interfaces (mysql and psql).
psql CLI is a tool which is pleasant to use, has no bugs in the interface and it even gets improvements from time to time.
mysql CLI is awful to use (e.g. doesn't display long lines properly, has difficulties with history editing, etc) and looks like there wasn't a single improvement since 1996 (I'm sure there were, I just never felt the effect of such improvements).
https://github.com/hanslub42/rlwrap
Note that I've never tried it myself with the mysql/mariadb CLI, but I have used it with other tools, and it's brilliant.
This is probably not a fair comparison.
On the existing lz4 machines, the zfs pools are already badly fragmented, making it difficult to find regular blocks: https://utcc.utoronto.ca/~cks/space/blog/solaris/ZFSGangBloc...
On the new zstd machines, the zfs pools are still pristine, after being restored from backup.
If we want to isolate the effects of migrating from lz4 to zstd, we also need some new lz4 machines for comparison.
The "special" device is a full-blown vdev that you add to a pool (it is not a cache) - typically a 3 or 4-way mirror of SSDs.
Now all metadata read and write happen at SSD speeds.
We wrote about this in the rsync.net Technical Notes for Q3 of this year[2].
I know what this kind of SSD based vdev does to typical mixed file performance but I'm not sure how metadata-heavy a postgres implementation is ...
[1] Yes, they really are called that.
[2] https://www.rsync.net/resources/notes/2021-q3-rsync.net_tech...
It's all about the latency.
There's another reason why you don't want to go beyond 80% utilization, and that's because the block allocator will switch behavior to a more involved search, which can take a lot more time.
Thus allocating new blocks can get really slow once you get past 80%.
So turns out it's a bit more involved than what's been commonly told as a straight up 80% == bad scenario. ZFS by default divides[1] each vdev (RAIDZ or mirror set) into ~200 allocation regions called metaslabs[2].
When allocating from a metaslab[3] it will check if the free space in that metaslab is below the threshold defined by metaslab_df_free_pct. It seems the threshold was changed to 4% free space at some point[4].
If the free space is above the limit it will use the fast first-fit search, if not it will use the expensive best-fit search.
However, as noted that threshold is per metaslab. So if the pool is fragmented, even though the overall free space in the pool is above the 4% threshold, there might be metaslabs with less than that free, which will lead to the expensive best-fit search.
So it's not a hard limit, but it should start to be noticeable above 80%.
[1]: https://www.delphix.com/blog/delphix-engineering/openzfs-cod...
[2]: http://dtrace.org/blogs/ahl/2012/11/08/zfs-trivia-metaslabs/
[3]: https://github.com/openzfs/zfs/blob/master/module/zfs/metasl... (note metaslab_df_free_pct)
[4]: https://www.truenas.com/community/threads/zfs-tweak-for-firs...
When it hits that threshold it becomes immediately noticeable - I/O performance will basically fall right down. At my old job we noticed it around 93% on Solaris 11.x, and maybe 85% on Solaris 10 off the top of my head, on two different production systems. Best practice was considered to keep pools on any ZFS system (OpenZFS or Solaris) below 80%.
If you have say a couple of 4TB disks in a mirror vdev, get it say 98% full, then expand adding new vdev with two 8TB disks mirrored, then you still have to pay the price whenever ZFS allocates from the first vdev.
ZFS tries to spread the load between vdevs in a weighted manner, which means most of the writes would go to the new vdev. But if it allocates from the first vdev you'll pay a latency price.
This means the latency in this case can be quite unpredictable, and not something that'll go away until enough blocks are freed from the first vdev.
> Heap is the only tool that automatically captures all user interactions on your site, from the moment of installation forward. A single snippet grabs every click, swipe, tap, pageview, and fill — forever.
Revolting... at least my content blocker had their collection script already blocked.
Not trying to make a case for large scale data collection without purpose - but if clickstreams are going somewhere I would rather it no be to a company selling me ads.
> The maximum amount of time that Analytics will retain Google-signals data is 26 months, regardless of your settings.
First party can be longer, but is usually shorter than this.
Disclosure: I'm a Googler. Logs of my product are retained for 14 months and until recently it was 1 month, but it turned out we need to debug cases older than that.
Google and Facebook are arguably better because they at least are somewhat incentivized to keep your data to themselves.
Even though that's like arguing about which STD is better instead of using condoms.
Let's do some research before crying foul. Google doesn't have to share your data with outside people directly, they are more than enough of a customer for that data themselves.
I did read the terms and it was anything but clear on that, it only says that data generated is owned by the customer but also allows them to create "anonymized" versions of the data that they do own.
Even if the terms didn't give them a free pass to share the data (and that they don't sell it anyway like many do) that data is only a company policy change from being sold... The probability of that happening quickly approaches 1 with data brokers sitting on their doorstep with bags of money.
> Does aws share your data with 3rd parties?
Assuming that you're talking about the data stored on rented servers - and not their customer and usage data which they very likely do share with third parties - they have very obvious incentives to keep that data readable only by the customer.
Not that I would be all that surprised to learn that some cloud provider had been mining their customers disk storage for data...
> Google doesn't have to share your data with outside people directly, they are more than enough of a customer for that data themselves
That's what I meant with that they are somewhat incentivized to keep it to themselves. That's marginally less disgusting than having it being sold in bulk to companies that specialize in aggregating, de-anonymizing and re-selling data.
Fundamentally it is guaranteed that sharing your data with a third party for free is gives you less recourse/expectation of privacy than a company that is getting paid to do so (like AWS S3 in my example).
> Revolting... at least my content blocker had their collection script already blocked.
Inspired to add a content blocker by this, thank you. Which do you use?It would seem to me if you can keep the utilization of the disk under 80%, and support TRIM (which lets the SSD know which pages can be erased), you should be able to get really high performance out of them with a Copy on Write file system.
ZFS is Copy-on-write (One byte written in a block requires the whole block be rewritten, with the old one scheduled for reclamation).
The underlying SSD wear levelling algorithm is Copy on Write (Writing a single byte involves writing the data from the page to a new block, and then erasing the old one sometime later)
That means a tiny 1 byte modification to a postgres record involves creating many new unnecessary copies of a lot of data...
I imagine that if the three layers could be combined, dramatic performance benefits could happen, since written data might go down by an order of magnitude at least....
The same thing happens with all the layers of virtual-machines
It provides VMs with drivers for purely virtual “virtio” devices (storage, network, etc) with no effort or overhead put into mimicking the mechanics of a real sas/sata/scsi device or whatever.
The result is less complexity and a much more direct data-path so things should hopefully not only be more performant, but also more stable.
the original block will be removed from the mapping entirely (not referred to by a copy)
That means any write of data must be made to a new, freshly erased, location.
The old version of the data, sitting in the middle of a large eraseblock is no longer used, but cannot be reused until everything else in the eraseblock is either unused or copied elsewhere.
In the worst case (of a nearly full SSD), it means that for every 1 byte write into a 4k page, a full 1Mbyte of data needs to be copied. Typical cases are better (when the drive has plenty of spare space, so can delay the reclamation of the eraseblock for as long as possible, hoping that other things in the block are invalidated, reducing the amount that needs to be copied).
TL;DR: It's complex, but a lot of copying is involved with most writes...
AFAIK а modified record is copied to the same page if there is enough space (which you can tune).
https://engineering.fb.com/2016/08/31/core-data/myrocks-a-sp...
COW filesystem means you can make (virtual) copies of files /blocks without writing (duplicating) them on the media. They only write metadata for bare copies and delay duplicating data until a virtual duplicate is altered.
Actual writes are written into a free block, then the old block is marked clear. The old block is not copied and then written over. In my understanding that's not what COW characterizes - its refering to how copying data is almost free (only costs metadata changes in COW filesystems) until copies are altered (written to)
In Postgres, Transactions work on a "snapshot" of the data that existed at one point in time. That snapshot is logically a copy of the data, but in reality uses copy-on-write of records to avoid having to make a copy of the entire database at the start of any transaction.
In ZFS, it works as described.
In SSD's, operating system 'write' commands are treated as transactions - ie. certain ordering semantics must be preserved in case of a power failure. Since performance is improved by having extra parallelism and not doing the actual operations in the order they are presented by the OS, a copy-on-write model is used to ensure that an incomplete transaction can be rolled back. This isn't supposed to be user-visible, but occasionally in a badly broken SSD, you hear users complaining of 'it works fine, but then when I reboot my computer everything I did is undone'! Well that's because no transactions are committing...
I guess arguably this comes from the lack of availability of ZNS drives, especially ones that aren't outcompeted by a Samsung 980 series drive on workloads that don't exceed the 980's warranty's 600 drive writes over 5 years. And depending on what you're doing, you may even use such drives for the IO-heavy load and retire them after 90% of their TBW to something like fileserver duty.
A zoned-like workload that is append-only within large chunks and uses a Deallocate (Trim) or Write Zeros command to clear chunks should perform well even on standard SSDs. I don't think the exact zone size has much impact, as long as it's comfortably above the erase block size.
https://www.usenix.org/conference/inflow14/workshop-program/...
Using a COW filesystem adds at least some amount of usage, since instead of modifying in place, you'd write a new block and only trim the old block sometime after the new block is committed; but if you don't have snapshots and you have zfs autotrim (and it trims all your old blocks), the commit interval is short (5 seconds by default?), so I wouldn't expect a big difference in effective free space here.
Taking to the next level - Batching your I/O in software is how you can start saying things like "Transactions per disk I/O", not just "Fewer I/O for those transactions which now fit in fewer blocks due to compression". Batching doesn't have to mean "nightly processing". It can mean "all requests which occurred over the last 100uS". From a user's perspective, this can effectively still be a real-time RPC experience. For systems with very heavy load, this sort of micro-batching can add many orders of magnitude improvement in throughput. Also bear in mind that the more transactions you have available to compress each time means you get better odds when dealing with entropy.
I have personally developed software that can insert 2-5x the stated write IOPS figure with these sorts of tactics. On modern NVMe devices this can mean you start tickling 8 figure transactions per second if the size of each request is very modest.
Sounds like you solved very interesting problems.
Edit: reading other comments, it looks like their block size is only 64KiB, so this isn't the case, so I don't know the answer. I can only think that perhaps it is an issue because ZFS doesn't deallocate the freed blocks quickly enough and they are making changes fast, meaning a significant amount of disk space is used by blocks that are about to be TRIMed but haven't been yet.
Same for internet connection, you can buy transit, no need to become your own ISP. So you don't need people who deal with peering agreements, etc.
For electricity you can make a support deal with a local electrician company. You don't need a guy who can build and maintain a custom power supply unit.
It does help however to have someone with basic sysadmin and network skills. But if you don't have that, you will sooner or later screw up your AWS infrastructure too.
Maybe if you're wiring a closet in your office, but no colo facility is letting you within ten feet of their power infrastructure. The best you're getting is a racked-mounted UPS.
This is far more fiscally irresponsible for most companies than just using AWS.
We can argue whether it is billion dollars or few millions , but there are 1000s companies spending 50million + or even 500 million a year on compute globally.
For a lot of these companies it would make economic sense to keep their own DCs , however hiring that talent and retaining is very very hard, especially if you need even 1-100th of what AWS offers in their IaaS services, this is what will cost a company a lot more than cloud.
It is really shortage in devops/sysads that drives cloud . The talent shortage is just getting worse, newer devops engineers who know only k8s blink seeing a bare metal VM let alone know how to do proper cable management.
No, it's not. I've have never once met a company that went cloud because they couldn't find someone to rack and configure infrastructure. There's considerably more to building and running a data center than just hiring staff.
This thread is full of comments making it sound that running your own hardware is black magic and cloud is the only way to deploy at scale. Yes it is hard, but not so hard that spending $500 Million/year [1] on cloud is cheaper than running your own DC.
It is even harder for companies that already had investments in DC operations to justify.
Leveraging public clouds for workloads that need flexible scaling or use cloud native services which are not cost effective / feasible to develop in-house it always makes sense to shift, however the lift and shift for everything we have been seeing has no justification for companies already running their own infra unless developers and system admins are demanding cloud, either they want cloud in their resumes or don't know how anymore to write traditional applications.
Many enterprise apps even today are not cloud native or leverage anything beyond basic compute and storage IaaS services and do not have high load variations and not even set to auto-scale, they never had technical or economic reasons to be moved to the cloud if you already had sunk costs running a DC.
[1] We can argue what that number is all day , when including everything you need to do run a DC it is cheaper than on the Cloud. There is a definitely a number, as at very top end of usage companies like Apple or Facebook etc are not sitting on top of AWS and it would never make sense for them at their scale, somewhere in between that point does exist.
The rest of your comment makes it clear you're entirely speculating, like hiring being an obstacle to spending tens of millions of dollars to build a DC.
I am not entirely speculating, I have been part of the recruitment industry for the last decade, but you have already decided to discard statements from a random stranger on the internet (totally reasonable!) so no point in arguing this further.
If you have a petabyte of chatlogs, sure, you have 24x7 obligations to millions of people. If you have a petabyte of astronomy data, you have like 3 research scientists using it.
Not necessarily. Hardware requires people to physically replace failed drives and otherwise do on-site maintenance.
In the unlikely event that an AWS volume fails, I can (and have) automation to fix that. While everyone sleeps.
This is the premise of colocation (as opposed to building your own server room). A colo is a secure building with round the clock staff. Hardware vendors offer rapid on-site parts replacements and can gain access via the on-site staff, and the colo has services to perform on-site work like "remote hands" as well.
> In the unlikely event that an AWS volume fails, I can (and have) automation to fix that. While everyone sleeps.
Fault tolerant architectures can be be deployed on colocated hardware too.
There may exist some colo where I can get a server(or storage, or network cards or anything else) added in minutes over an API call but I haven't heard of any. That's usually found on the VPS side.
> Fault tolerant architectures can be be deployed on colocated hardware too.
They can, usually requiring that you specifically setup redundancies and the like. Which is something that you already have for many cloud offerings. Your automation and redundancies sit on top of the vendor's existing redundancies.
For instance, the EBS volume I mention. It is not a disk. Its not even just an array. It's a far more sophisticated abstraction. If there are issues, it can automatically fetch blocks from your snapshots(if the blocks are unmodified, something they also keep track of). Not happy with spinning disks and want a SSD? No need to place a service order to your colo provider, just send an API call and this will be automatically migrated to SSDs without your applications ever noticing the difference (other than the response time) and with zero downtime. Your software could even do this if it notices that the workloads require it.
If an AWS datacenter goes up in flames the systems I manage will still function (and will self-heal, assuming they even get affected, which for big zones they might not be). I don't have to talk to anyone. I can be sleeping and this will still happen.
It's a completely different level of abstraction.
If you want to compare a big cloud provider with either your own datacenter or colocation facility, there's a big disparity in scale. At a minimum, you would have to compare with several interconnected datacenters or colos. You still don't get the abstraction layer.
It's all missing the point though - I was pointing out that software doesn't necessarily need to have 24x7 staff, as the parent poster was pointing out, even for exceptional (but predictable) issues. Sure, you need someone on-call to handle completely unexpected events, but I don't think that was the point being made.
To be fair, you can't get hardware added in AWS via API call either. What you can do is spin up instances/storage/etc via API call, as long as that spare capacity hardware is already set up, available and ready to be allocated to you. Which you can also do on on-prem hardware.
If you're saying that your utilization is so peaky or unpredictable that you end up needing an order of magnitude more, or fewer, resources available day to day, then you are absolutely correct that provisioning so much spare capacity on-prem would be prohibitive. This is an use-case where AWS excels.
But if your utilization doesn't have dramatic peaks and growth is mostly predictable, then it becomes practical to provision for it on-prem and it'll be a lot cheaper.
You're right, hypergrowth (a subset of unpredictable capacity needs) is a perfect use case for AWS or similar.
The vast majority of companies, for the majority of their corporate life, aren't in these categories (spiky, unpredictably) phase though, so for those, AWS is a large cost premium.
I say this as someone who built, manages and operates datacenters and colo spaces.
If the problem is that your group of six-figure salary people only know how to put data into AWS, or other cloud services, and not design/engineer/maintain your own bare metal infrastructure as well, then that would definitely be a limitation.
For reference, a few petabytes of data is not actually that many systems these days, if you have something like a bunch of 72-drive supermicros or equivalent with 14-16TB drives in them. Set up properly this can be administered by one FTE (of course with additional staffing/tech resources for when that FTE is on vacation/unavailable, and appropriate training for other persons who might have admin on the setup).
my very rough calculation here says that a 36-drive ZFS RAIDZ2 composed of 16TB drives is something like 492TB (447TiB) usable storage capacity.
so five such arrays would be 2460TB.
compared to the monthly AWS bill for 2.0 to 2.5TB of data you could probably afford to entirely duplicate the whole setup in a twin identical set of hardware at a geographically diverse off-site location.
It's not about whether or not the engineers can make the colocated setup work.
It's that you're going to pay a lot of hidden costs with a colocated setup. Engineers can't set up, maintain, and do on-call for the colocated setup without subtracting from their primary working hours.
Each additional engineer you have to hire to help with the colocated setup is $200-400K fully loaded out of your company's budget. If you have to hire 3 additional engineers to fill out your colocated on-call schedule and help set up and maintain the system, that's easily an extra $1 million per year on your budget. Cloud is expensive, but $1 million goes a long way.
It's easy to look at a potential AWS bill and a potential colocation and hardware bill and declare colocation the winner, but then you still have to set up and maintain it all as well as constantly train everyone on it.
With AWS, you can hire engineers with AWS experience and they'll understand the big picture of how to work with things on day 1. With a custom setup, you're at the whims of whichever employees set up the system because they know it best.
Colocated systems tend to work very well at first when the original engineers who set it up are all still at the company and it hasn't run long enough to start encountering rare failure modes. They quickly become a nightmare when your engineering staff turns over multiple times and nobody can remember who knows how to do what on the colocated system or if the documentation is up to date or not.
Everything above really sounds like it's just regurgitating AWS sales person talking points.
Sounds like a systemic management / CTO-level problem to me if a company isn't willing to put in place the hiring practices and compensation, documentation systems and operational procedures to deal with that sort of concern.
If your core engineering staff is turning over multiple times for arbitrary reasons you have other problems to deal with.
> Engineers can't set up, maintain, and do on-call for the colocated setup without subtracting from their primary working hours.
If a company can't hire datacenter techs to install hardware, cables, and swap hardware as smart remote hands, maintaining as little as a couple of 45RU cabinets of gear, you also have other management/systemic problems to deal with. I'm looking at this from the point of view of a facilities based bare metal ISP that owns/runs all of its own hardware, and can tell you it's not rocket science.
People leave for all sorts of reasons: Moving for family reasons, becoming stay-at-home parents, moving for a spouse's job, retiring, starting their own companies, or even just getting bored and wanting to do something different. Or it could be as simple as getting promoted to a different role.
It's unrealistic to make engineering decisions with the assumption that the same engineers will be around and stuck on the same project forever.
Like the OP said: Every hour they have to spend working on the colocation setup is an hour they aren't spending on your company's competitive advantage, so you have to hire more engineers (and more managers) to compensate.
> If a company can't hire datacenter techs...
How many techs do you think you need for reasonable on-call coverage? 3? 6? Add a manager in the mix because you need someone to manage them.
The costs add up quickly.
It's weird to see people championing colocation as a cost saver and then pivoting to arguments that you just need to hire more engineers and techs and manage them.
Employees are expensive. One of the primary benefits of cloud is that you don't have to hire and manage all of these employees to do all of these things at the colo.
Humans are incredibly expensive and notoriously unreliable when compared against “machines”; or in this case an API.
It’s usually worth paying 2-3x the cost to have someone else manage something for you with a given SLA, because that’s what it will end up being when you decide to bring it in house when you take into account the time and effort needed as well.
A “good” & “reliable” Systems Engineering team, that can offer 24/7 support will take around a year to hire and setup, and they need roughly the same amount of time to transition you off AWS in to your system. They probably need closer to 3-5yrs to give the same level of documentation, API’s, tooling, processes, UI’s and training that AWS already provides.
Let’s call it 5 years to get to the level of AWS when you started the transition.
A decent team of 5-7, including engineers + PO/PM + UX and so forth, is at least $1.5M/yr. That’s $7.5M over 5 years, not including your new hardware and networking costs. Let’s call it $10M. You’re also 5 years behind AWS now, and over that transition you’re still paying AWS, and your development speed has halved as you wait for your new team to build or transition infrastructure.
You can trade cost and quality for speed and have everything ready in 2-3 years by setting up a few teams. Add HR support, more contractors etc etc. you’re looking at a $10M+ outlay again, regardless.
Or you can keep paying AWS $5M/yr, renegotiate fees often and literally not worry about that headache and focus on your product.
> They probably need closer to 3-5yrs to give the same level of documentation, API’s, tooling, processes, UI’s and training that AWS already provides.
You don't need to build an internal AWS to manage your own servers.
I can’t remember the last time I took a call outside of office hours, and even in hours it’s very rare. There’s enough resilience built in that any issues can wait until morning
The last major outage was in 2017, before we had a third member of the team. I was on the other side of the world installing a new system, the other was on leave. We had a network issue, OSPF melted and knocked out some services, we were down for about half an hour as I rebooted the core switch pair remotely.
(We’ve since redesigned so that doesn’t happen)
We get paid nowhere near six figures either.
Sure you can be ridiculous, I remember one team I worked on that employed a full time unix contractor (on 3 times the staff wage) to look after 6 servers and deploy a tar ball every few months. I replaced him with a small shell script. Another was a DBA looking after a small oracle database (oracle - which of course is that generations “just use amazon”)
And then basic other things like having remote smart hands ready to go, and common failure items like fans, power supplies, fan trays, hard drives pre-positioned and ready to swap in. With MOPs for swapping them. Stocks of basic things like fiber patch cables, commonly used transceivers, copper patch cables, stored in every cage.
Or we have seen all of this and that's exactly why we don't want it.
Building a company is hard enough. Adding the overhead of developing, maintaining, recruiting for, and staffing our own datacenter is madness when I can click a few buttons and get the same thing from a cloud provider without hiring anyone extra to manage the datacenter.
No one is denying that a proper data center management system can exist. We all know it can exist.
The issue is that it's a huge distraction with a lot of potential pitfalls. Your network infrastructure with Cradlepoint LTE radios in the colocation cages sounds great after it works, gets set up, stays documented, and all the bugs have been ironed out. But that's a lot of hidden work that could have been allocated to launching the product faster.
If the use case is somebody developing a software product that is a totally other scenario.
One big pain that no one seems to mention are the mundane details that come with running your own DC. Even in our fairly small operation we had half of a person dedicated to maintaining the warranties on the hardware, sourcing replacement hard drives for 10yo servers, disposing of retired hardware, purchasing new hardware, etc.
These things take a huge amount of focus and don't contribute to your product at all. The "cloud" has removed all of this busywork and I wouldn't go back to the old way for almost any amount of money.
“Building a company is hard enough, adding the overhead of actually building a company is madness”
> But that's a lot of hidden work that could have been allocated to launching the product faster.
The topic of this thread isn't about a brand new start-up with no resources trying to build their product as much as possible, it's about a company that spends millions in cloud invoices.
Building a company using AWS, totally legit. Scaling this usage up to millions of dollars because you don't want to manage the hassle, fair enough it's your money, I don't mind if all your profit goes to Amazon really…
People move to the cloud to escape their company’s IT process… there might be some unicorn company out there that does infrastructure “right” but I’ve yet to work there.
* https://liveramp.com/developers/blog/google-cloud-platform-g... * https://liveramp.com/developers/blog/migrating-a-big-data-en...
One big point was challenges of maintaining multiple colocation sites, with cross replication, for disaster recovery. Since Hadoop triple replicates all data within one DC, this requires 6 times the disk storage capacity of data size for dual DCs. In contrast, cloud object storage pricing includes replication within a region with very high availability such that storing once in cloud storage may be acceptable. Further, you also need double the compute, with one of the DCs always standing by should the other fail.
*) https://hadoop.apache.org/docs/r2.8.0/hadoop-project-dist/ha...
Probably some learning pains but at least i don't have to "lose" half my storage because the sizes aren't paired with anything.
For a growing company I would always weigh the organizational overhead of moving on-prem to the actual cloud costs. If your goal is in the next year to grow revenue from 50MM to 100MM, then a company wide engineering effort to move to bespoke platform to save 5-10MM just isn't a good use of time (keep in mind you are trying to hire and release new features).
It's not that it's impossible; it's that trying to both scale up your business and provide infrastructure to yourself is just unnecessarily increasing the difficulty on your businesses execution.
As an example, imagine if the founders and engineers of Backblaze thought like that.
It all depends on the size of the investment and how you need to run it. I built a "new" environment on a company premises due to some compliance requirements that would be cost-prohibitive in AWS or GCP. The gear was procured through a leasing vehicle, and the hardware vendor had an SLA for delivering compute and storage. HPE happened to win the bid.
There is very little difference operationally. From a costing perspective, it's about 40% less than an AWS solution. But in fairness, the customer had an existing investment in a facility - you'd reduce the savings if you had to lease appropriate space in a colo. There are some differences in terms of headcount, but those staff aren't in NYC/SFO/BOS, so they are very cheap -- senior level engineers for $80-120k, fully loaded.
Startups do stupid shit like buy supermicro computers and cobbling together hardware that gets them into trouble when the mad scientist moves on to a new gig. Makes sense when you're drowning in VC money and need to hire people, but doesn't make sense in most other ways. You avoid that by doing competitive procurements and paying marginally more for HPE/Dell/Lenovo/etc.
If you think that standards based x86-64 hardware running Linux and ZFS, or FreeBSD and ZFS is something that is super unreliable and requires a "mad scientist", then yes, you are definitely in HPE and Dell's target market.
Home built web frameworks (which apparently aren’t “bloated” and “slow”), piles of bash scripts because they never heard of Salt (or whatever is the latest config management tool)…
Almost always they think they are “saving money” by doing what they do, rarely do they ever consider the opportunity costs to rolling the entire software stack from the hardware to web stack on their own.
Good times.
btw, we also wrote a HA cluster software for sun solaris in year 2000...
> Home built web frameworks (which apparently aren’t “bloated” and “slow”)
Well, let me show you something: https://code.djangoproject.com/ticket/31624
This is a performance regression of Django's ORM for deleting rows from a table, a query that should be about as simple as it gets:
delete from tbl;
And yet it's orders of magnitude slower because it somehow generates a bloated query. So yeah, let me tell you, the frameworks that people are using off the shelf–there's a pretty good chance that they're 'bloated' and 'slow'.> piles of bash scripts because they never heard of Salt (or whatever is the latest config management tool)…
Well, that's kind of the point, isn't it? If you're trying to get a project off the ground as quickly as possible, you kinda don't care about 'whatever is the latest config management tool', you will use what is already in your toolbox that will get the job done.
And AWS systems do not have to face the same problem? Is it easier for a new staff to come in and modify a millions worth of AWS system than a colo system?
Last thing I remember, I was
Running for the door
I had to find the passage back
To the place I was before
"Relax, " said the night man,
"We are programmed to receive.
You can check-out any time you like,
But you can never leave! "
2PB is really not that much stuff these days. It's less than two cabinets of equipment and that amount of space (80RU or so) includes routers, switches, AC power distribution, OOB, etc.
* to get performant access to a storage cluster is non-trivial, there are many different variables in place which must be correctly tuned to get good performance. Network topology, high quality NICs and switches, a tuned Linux kernel, client side caching settings, network packet sizes, file system block sizes, erasure coding settings, etc.
* your solution mentions nothing of backups, offsite failovers, or disaster recovery plans.
* your solution mentions nothing of a physical datacenter: fire suppression, battery backups, hvac, power supplies, backup generators, server racks, sound isolation, workspaces for hardware maintenance, network cable routing.
* If you have multiple geolocations you need to have dark fiber or ip transit between locations with multiple ISPs to have high speed connections between sites without downtime.
In addition to the raw costs you have to factor in the lead time of building a qualified infrastructure team, building out the requirements, provisioning hardware and datacenter space, setting everything up, and then tuning everything. With infinite money this is probably still a 2 year lead time at minimum.
I do agree that running multi petabyte workloads in AWS is probably not optimal, but when you are a startup in the growth stage it is probably better worth your time throwing VC money at AWS and building out your product. Eventually, the business should probably migrate to self managed infrastructure once the right product fit has been found and the business is looking to streamline.
You can generally buy whatever you are renting from AWS for 1 to 3 months of an AWS bill.
The only thing I don’t get from colo is a bunch of other customers thrashing the cache on my CPUs
Databases are not web servers there’s no possible way to not run a database on smaller / fewer instances when running at non-peak times. Instant scaling is the only possible advantage AWS could bring. However with the prices they charge it’s simpler and cheaper to just buy/rent your own hardware. Especially if you have to pay egress fees. (bandwidth is really the biggest ripoff)
I don’t like lock-in, but the prevailing view has always been in favour of lock-in, be it IBM mainframes, oracle databases, windows servers etc, and if you swing that way aws has tempting offers.
Oh and databases do scale. Say you want to run end of quarter financials that require a lot of processing for a day, you bring up tons of read replicas and away you go
Also, this is generally why you run financials overnight. If your hardware is serving transactions during the day it can easily run your quarterlies at night.
nginx is far easier to maintain than AWS load balancers which is what load balancers their load balancers are. The best part about nginx configs? They are cloud agnostic and will work on everything from a Raspberry Pi to a 128 core EPYC server.
I'll tell you something about RDS reliabilty, your monthly maintenance window brings your DB down far more often than a single unreplicated server ever fails. EBS (like the entire thing) has failed more times in the last year on us-east-1 than my colo RAID.
The selling point of AWS is that if you pick AWS and it fails you can say, well the richest guy in the world can't figure this stuff out so it must be impossible, when in reality high school kids could make a more reliable system. If you pick AWS you have the unreliability of the base software / hardware of their systems plus whatever the AWS engineers fuck up. At this point it's pretty clear that they can't even keep a SAN working.
What is the amount of work required every time some team wants to spin up new stuff? What is the turnaround time between when they file the ticket and the work being completed?
You put your config in there and then you can copy them anywhere, even between AWS accounts. They’re even cloud agnostic so when you move to GCP they still work.
Files work great with fucking shell scripts, entirely compatible, it’s rumored that shell scripts beneath the hood are also files. https://github.com/brandonhilkert/fucking_shell_scripts
You can use FSS on docker, kubernetes, bare metal (otherwise known as computers), Windows, Mac, Amiga, VMS etc.
Plan9 which was made by the creators of Go works exceptionally well with files. Some people even say that AWS converts your LB config object to a config file for nginx when it’s provisioned.
If you want to get really crazy you can put all your files in one dir like /etc and then you can use this program grep to search all your configs at once. You can even use things like perl -pi to provide a programmatic interface to editing all your configs at once.
You wanna push 1 evening a bit late to deliver something valuable for the project? Sorry, no can do.
I don't even factor in horribly expensive migration projects that brought actually 0 added business value for the type of apps we use. We still have to keep our Network, Windows and Unix admins, various App support personnel etc., there is plenty of work for them with AWS. Not 1 single IT guy was made redundant.
No cost savings, in contrary.
In other words your assessment would only be true if AWS had a 0% or a negative margin.
The public cloud and managed services work great for the most common use cases, but go outside those and you start having to engineer around limitations.
If you have a sizable footprint in any given dimension you're trading one complexity for another.
One possible solution for this if you want to do it as bare metal you control, is leased equipment (even with $1 buyout at end of term), which can be accounted for differently than purchasing it up front.
And yes, reservations make a massive difference economically.
Maybe it's worth a couple of million to not have to deal with the risk, and just keeping status quo.
Usually, those types of judgements are based on thinking of AWS as a "dumb datacenter" such as a bunch of harddrives or just bare cpu.
AWS is more cost-effective if you use high-level AWS services instead of just storing files in the cloud. In this case, it looks like Heap is also using AWS Redshift and probably a bunch of other services in the AWS portfolio. A similar comment I made previously: https://news.ycombinator.com/item?id=28288352
So for self-hosting hardware, Heap would not only build up the petabytes of diskspace, they also have to replicate Redshift functionality and the entire AWS services portfolio they're using. If you use enough AWS _services_, it becomes cheaper than self-hosting because you don't have to reinvent the wheel.
Perhaps that's more of a cautionary tale for new projects than a justification for the expense though.
It's not that migrating out isn't possible, it's that Amazon is providing "Engineering/SiteOps Departments as a Service" at a price that's hard to compete with in house.
From the GP's post - I was trying to say I read this as "the collection of services AWS provides would require several in-house engineering teams to compete with" not "vendor lock-in".
A single service is relatively easy to replicate if it is core to your business, but an entire on-demand datacenter w/ abstractions like Time Series databases, pub/sub services, etc. isn't as trivial to do yourself. Many of these services require teams of engineers to manage at scale, and engineers that _understand_ them well.
The on-demand service catalog of the cloud provides significant value. It's more than "pub/sub" or "blob storage" as a service. It's an entire engineering organization and data center as a service w/ pre-built architectures for you to start using today.
You can find case studies for both positions:
- migrate to AWS to save money: Netflix, Guardian newspaper [1]
- migrate away from AWS to save money: E.g. Dropbox [2]
A lot of companies (especially non-tech businesses) don't have the technical skills to run internal datacenters at the same competency as AWS. Thus, they don't want to be "locked in" to their own IT department that's slow and handicaps their business.
Dropbox, Facebook, and Walmart would among the very few that can competently run their own datacenters with advanced services like AWS.
[1] https://web.archive.org/web/20160319022029/https://www.compu...
[2] https://www.google.com/search?q=dropbox+migrates+off+aws+sav...
[0]: https://aws.amazon.com/solutions/case-studies/dropbox-s3/
The story described Dropbox moving "34 PB of analytics data (Hadoop)" to AWS.
My reading of Dropbox's Magic Pocket / Diskotech appears to be storage for customer raw data -- similar to BackBlaze type of raw storage.
It's 2 different use cases so it's not surprising Dropbox found AWS to be effective for analytics workloads. AWS has an extensive portfolio of software services to analyze data so Dropbox may have concluded paying AWS would cost less than reinventing the analytics pipeline in-house.
Mostly this only works when your utilization is low(ish). Once you have high load 24x7, the AWS profit margin will quickly overtake the self-hosted solution.
Considering the number of people here commenting about costs but have never managed a P&L, I don't think certainty is high on the list.
- Automation - this was noted by another commenter, but with AWS we can fully automate the instance replacement procedure using autoscaling groups. On hardware failure, the relevant database is removed from its autoscaling group and we automatically start restoring a fresh instance from latest backup. This would be much more difficult if we were to manage our own hardware.
- Flexibility - we have the ability to easily change instance classes via selling/buying reservations. Some of our biggest wins historically have come from AWS releasing new instance families - we've been able to swap out the hardware for our entire cluster over a week or so, for negligible cost (often saving money in the process due to the cost per unit of hardware decreasing on new instance classes). While we could leverage the same developments in a self-managed environment, it would be more difficult, and likely more expensive due to how capital-intensive self-hosted is.
Additionally - there's a ton of value in the integration of the AWS ecosystem. We use many AWS managed services, including heavy use of RDS, Kinesis, S3, and others. For a company with a relatively small engineering team managing a large infrastructure footprint, it hasn't made financial sense yet to invest in moving to self-hosted infrastructure.
Separately to your comment about LVM... the LVM snapshot requires that a separate part of the volumes be set aside to hold the snapshot data.
If the snapshot volume fills up with changes being made to volume that holds your data before the snapshot completes, then the snapshot will fail.
This does not occur with ZFS as you have noticed.
Also, with many customers also using AWS, much of the traffic may not even leave a datacenter, improving speed, reliability, and maybe even cost.
I'm not sure exactly what "This" refers to? Just wanted to note that a ZFS snapshot can be destroyed when the parent pool runs out of space too, but you don't need to allocate a volume.
I’m curious: The engineers that brought these significant cost savings to your company, did they receive a share of the money saved?
The engineers didn't do it in isolation - how much of a share goes to the office receptionist that answered the phones and kept visitors out of the way of the engineers? How much goes to the Finance department that kept the engineering paychecks coming while they did they work? How much goes to the salespeople who kept the deals flowing and money coming in that filled the disks in the first place... and so on and so on.
Once a company exits the "a few devs in a garage" stage, many people contribute to the company's success.
No need to allocate. Something like "our clever engineers managed to save us $2M per year, so we're giving everyone a $500 bonus this month" seams entirely reasonable.
At the same time, when I thought I had made a material contribution to the company's success that was above and beyond my normal role, I would ask for compensation of one kind or another. At times, I was told no, and at other times, my request was granted. I believe Wayne Gretzky said something like "You miss 100 percent of the shots you don't take."
My lived experience is that life is not fair. In the organizations where I was an employee, owner or executive, it has never been the case that everyone, in their "heart of hearts," really believed everything was entirely fair.
I think it's right and proper to recognize and work to correct injustices, to the extent possible, yet I think it's also true that people vary all over the map and on a perhaps uncountable number of dimensions, and trying to achieve complete fairness is simply not possible.
I will also say that my belief that life is not fair does not require us to go through life constantly angry and frustrated.
I could be wrong, and I don't wish to put words in your mouth. Interested in your thoughts, if you want to say more.
In my experience, doing things in the cloud is about as expensive per 12-18 months as buying the hardware up front is. That's super interesting for a fast growing startup that could go bust any minute and wants to spend every second of their time on growing, expanding and marketing.
But when you're spending so much on AWS you can save millions just by reducing filesystem overhead by 20%, it should have stopped making sense a while ago. $2 million should get you a team of 10 sysadmins and devops engineers. Sure automation would be more difficult, but you'd have the manpower to achieve it. Isn't that what running a business is about?
Flexibility, when you're growing quickly it's nice that you can provision new hardware instantly, but AWS is so expensive you could continuously over provision your hardware by 50% and still always be ahead of the AWS price curve. And as I said, you could fully swap out your hardware every 18 months and be at the same price basically. You could even hire a merchant to offload your old hardware and recuperate 50% of those costs.
And I'm not saying to throw AWS overboard altogether, that you have your core business outside of AWS's datacenter doesn't preclude you from buying into RDS, Kinesis, S3.
Is AWS just cutting you more financial slack than we're getting as a tiny company? Or am I underestimating the costs of getting that sysadmin team on board?
It seems like you are glossing over the other costs: staff to implement and manage, development time and maintenance for automation to re-implement everything that AWS includes, data center costs (not clear if you were thinking of hardware ownership only or data-center also).
I'm not saying you didn't think about those things just saying that they can't be ignored in these types of comparisons.
If it is your core business to provide that thing? Sometimes or even often worth it. Otherwise, often not.
My sense that in the last decade startups have not lost the skills to do Co-Location setups as they did in 2000s and think it is more complex than actually is. Co-Lo hardware management is hard yes, but if it not even worth doing 10+ million /year budgets we would never have SaaS companies pre-cloud at all.
Hypothetically if you are spending $50 Million+ / Year on the cloud a dedicated team of even 10 senior engineers to setup your co-location with your hardware to consider migration of your costliest and also least cloud native components would maybe cost $2-5M more. With attractive debt financing that is readily available these days you can easily amortize your purchase expenses over the 2-3 year hardware lifecycle and realize savings, there is not much justification not to also pursue this along with all other features you are also pursuing.
The cost is a very low investment compared to your costs with potential for very high saving ROI, so even if the chance of success is low you should give it a shot. i.e. If you can save say $10 Million on the $50 Million, your $2 Million investment needs only 20% probability of success to have expected value in the green.
[1] I don't think there should be only one core (the) business problem for a startup, there are always few critical problems startups have to solve for at any given stage.
Their core business problem is providing databases, and apparently they see leveraging the huge VM and storage pools available at AWS as a major advantage here (and I for one can’t blame them), over hardware spend absolute efficiency.
Being able to providing a couple hundred TB of extra SSD with a config file change (or return it and stop paying for it almost immediately), has real advantages over rolling it yourself, especially if you only have a 10 person ops team or the like.
Considering the apparent business model, I can see their point.
This project being discussed on the thread is likely a couple folks for a few months - low hanging fruit to save millions. What you’re referring to is a major business effort, if not doubling of headcount, for such a company with at best similar payoff. Running their own colos also means a lot of thinking, planning, and lifecycle management when it comes to equipment generations, upgrades, making sure you’ve got the right amount of spare capacity but not too much, etc.
Also, let’s not forget geo/availability zones.
Not saying co-located hardware is not always worth it - rather they seem to be aware of the trade offs, and are making a rational decision based on their business model.
Later, if they have switched from ‘rapid growth and adjustment’ to a more stable state where they can predict things more in advance, maybe they’ll switch. Maybe they won’t.
Like a large energy consumer, running on utility grid at a certain size in a certain area is often much better than rolling your own generation capacity. Sometimes it’s impossible or less cost effective. Sometimes it doesn’t make sense to even try to do the math, and just get hooked up to the grid.
Regarding being the core business, in this regard, that doesn't matter that much. You'll either pay amazon or hire your own team. In either case you'll be spending money in something that's not your "core" product. If you can replace amazon with a bespoke system for a fraction of the price and same resilience, why not?
Someone have said in this thread before, managing those things is not really rocket science. You can have a small, focused team who's able to manage a lot of resources and that can and often is cheaper than outsource. Obviously, it depends in a number of factors. Whether or not it's your core business doesn't seem a decisive one.
If something is a core part of the business, 1) it’s something they either already have a demonstrated level of competency in, or they wouldn’t be in business, and 2) efficiencies and improvements here should make them more money in a direct and measurable way, and 3) attempting to outsource it exposes them significantly to counterparty risk that can put them out of business, which is generally not considered a good thing.
In some ways it’s like a factory that uses a lot of power. Should they build their own power plant or use the utility. That depends on many factors. Using the utility is often the better choice and works out better, but not always.
If the quality and price of the power is a core part of what makes the company competitive, probably - and that is going to be a key factor in where the factory is located, when it operates, etc.
The devil is in the details, and I wouldn't say that it never makes sense to bring operations in-house, but your post didn't make a clear case from my point of view.
Or you could spend half the budget and in my opinion still be way ahead, but that depends on your execution and the talent pool that's available to you of course.
The fallacy is comparing hardware costs to services cost. The hardware is the cheap part.
When you run your own system, you have to develop the entire system up front and maintain it on the backend. The hardware is cheap by comparison to the salaries and development costs you pay.
> $2 million should get you a team of 10 sysadmins and devops engineers.
Probably double that once you add in fully-loaded costs as well as the compensation for ~2 managers to manage them.
Yes, you have to manage the hardware, and Y doesn’t automatically go to zero for year two, but the convenience of the cloud isn’t always cost effective. Don’t get caught up with the details. The 2 million figure doesn’t matter as much. It’s finding that inflection point and making the better business decision.
The development costs are OTC that are amortised during the life of the solution. Whereas in xAAS, they are MRC. The longer your tech refresh cycle, the cheaper it is. It's inherent to the pricing model.
Also when you get down to the cruz if it, these solutions (like openstack or vSphere), are software platform that provides similar features. There's not much development costs, it's just software licensing and PS.
In terms of operations, it's not like you can get rid of sysadmins, they just morphed into DevOps.
>Probably double that once you add in fully-loaded costs as well as the compensation for ~2 managers to manage them.
You might as well add all sorts of additional costs such as egress charges on exit and cloud consultancy.
Managing staff vs managing AWS ... I know what I'd choose (without really knowing the numbers)
With 8KB record size, there would be no read/write amplification caused by postgres. More importantly, ZFS fragmentation would be much less of an issue.
Of course, disks procurement can be a nightmare process, which is why AWS is printing money.
Heap using AWS just means they've not yet reached a point on that trajectory where the capital investment moves the needle enough to warrant it. That could be for any number of reasons.
I had to double check just in case, but petabyte is only 1000 terabyte. It may be big in terms of database, but rather small in absolute terms. You could fit a single Petabyte in a 1U server.
I doubt they pay listed price. And AWS is now mostly a Enterprise and Sales game. So once you ran other cost involved in managing, I would think you need to be multiple Rack scale before the cost break down better for your own hardware.
And that is excluding other benefits of sitting inside AWS ecosystem. The only thing I think AWS isn't so good at is the low cost, sub $1000 per month spending scenario. Where you are paying a lot more just for staying inside the ecosystem for things you may not be using. Those tends to flavour Linode or DO.
That seems a bit over the top. I see 18 TB drives available, but let's posit 20 TB drives, so you need 50 of them. I don't think you can fit 50 3.5" drives in a 1U space, even if there's no motherboard or power supply. 50+ drive storage chassis are generally 4U. I did see some 16 drive 1U servers though, so I'm pretty sure you could fit that much storage into 3U even though I also didn't see any 3U storage chassis.
[1]https://www.supermicro.com/newsroom/pressreleases/2017/press...
Scale makes IT cheaper, and AWS has scale. That means the actual total cost (not what is paid to Amazon, but what it actually costs to run) for something running on AWS will almost always be lower than a custom data center, due to the one-off reinvent-the-wheel work you'd have to do to run your own.
The only remaining question is, who keeps the savings. AWS can make a nice profit by selling their services at a higher price than it costs them to provide their part, but still cheaper than running your own data center. That means AWS will always be able to provide you the service cheaper than if you build your own.
Whether they're also willing to do that is another question, but it seems logical. At list prices, it's probably not worth running in AWS, but I highly doubt someone doing petabyte scale is paying list prices. Amazon has every motivation to provide a hefty discount to make the "own datacenter" approach unattractive.
Back in the university days we've built an information retrieval system. It ran on an IBM PC XT, with a 20MB HDD, which was pretty slow.
The heaviest queries to the information system involved full scans. They were too slow, slower than had been agreed with the customer.
So we installed a disk compression program, maybe Stacker or something similar. It ate some of the already slow 8088 CPU, and some of the scarce RAM. But crucially it compressed the data to about 50% of the original size.
This made the number of blocks to read, and, most importantly, to seek twice as low. The query speed increased twofold. We successfully completed the (tiny) software development contract.
E.g. https://drcoddwasright.blogspot.com: "In a time of SSD, multi-core/processor, two terabyte memory and Optane App Direct Mode machines, there is no reason not to build from BCNF data. Time to do what Dr. Codd demonstrated. Technology has finally caught up with the maths."
Who recently iterated on cstore_fdw to create columnar: https://www.citusdata.com/blog/2021/03/06/citus-10-columnar-...
But I don't think Heap's using columnar
> ... multi-petabyte cluster of Postgres instances... blocksize relatively high at 64 kb ...
The dataset should be the Postgresql "page size" which IIRC is 8KB, the reasoning for this is RMW cycles will read 64kb modify 8KB and write out the full 64KB amplifying writes 8 fold.
Also IIRC Postgresql will automatically use TOAST when needed?
And yes - Postgres will automatically TOAST oversized tuples and compress the relevant data (if you configure it to do so). This is much lower impact for us than filesystem level compression, as it doesn't affect the main relation heap space (or any indexes).
16k record size 2x amplification and still (?) allows compression w/ lz4
The latter does more compression (and therefore requires less storage and less IO), but is slower at decompression. Results show that query (read) performance did not actually change, whereas write operations needed only half as much time. Storage also saved ~20% space. So it was a win-win all around.
I love some of the articles, but in this case I definitely went the “I’m happy for you or sad it happened, but I ain’t about to read all of that” route.
It probably saved them from having to buy more storage this quarter, but it is a one time savings.
Old adage, but. If the data cost money to produce or is valuable invest in ECC and on-disk integrity checking.
Managing a few Pb I'll also emphasise updating your kernel to mainline of possible, especially if your on a centos like distro. These 2 things alone can potentially net huge performance gains.
I've ordered one for our research lab and it's not mission critical in any way and we have our sweet time to set it up. Is there anything more advanced than: default bare metal install with FS of choice, run benchmark, reinstall with different FS, rinse repeat? Or do you test with VMs?
The idea of treating db nodes less like pets seems great, there must be a few tricks to it tho..
A lot of the time I agree that titles can be quite clickbaity. But this one doesn't really seem to be... At least to me. The company upgraded their filesystem and it saved them a bunch of money. Title feels appropriate.
If the title were "Top ways to save millions on SSDs!" or similar, I'd wholeheartedly agree.
>Total storage usage reduced by ~21% (for our dataset, this is on the order of petabytes)
If 21% reduction is "millions" they should be spending in excess of 10 millions (each what? week/month/year?) provided that there is a linear correlation between storage usage and (failed and needing to be replaced?) SSD costs.
... relative to their previous ZFS configuration.
They didn't evaluate alternatives to ZFS, did they? They're still incurring copy-on-write FS overhead, and the compression is just helping reduce the pain there, no?
XFS, which we ran on for years before rolling out ZFS, does not compress.
LVM? Managing it sucks and the performance is bad (there's plenty of information on this).
I think Red Hat's Stratis was supposed to improve on some of this but it's not mainlined yet. For a while it sounded like Red Hat were going down the device-mapper functionality route of Stratis, VDO, dm-integrity and XFS, but I don't know where this is at.
Btrfs is probably the closest bet, but it's software RAID 5/6 has the known write hole bug.
Still waiting for bcachefs to get mainlined.
ext4 doesn't do CoW. XFS kinda does with its reflink functionality but it doesn't do filesystem-level snapshots.
Me too! As well as waiting for the successful launch of the James Webb Telescope...
Can you give some recent references to that claim? With thin LVM (introduced in rhel6-rhel7), performance is supposed to be OK.
Talk of rsync backups on live DB systems. zfs. On the fly disk encryprion.
All I can think of is, clearly these guys never worked with spinning disks and large datasets.
So much headroom to waste with SSDs, people are spoiled today.
It all depends upon how much headroom you have, the type of io activity, etc. I've operated systems under consistent 80% io load with massive datasets at the time, under spinning disks.
Running rsync would be madness on such a system. I know. I only did it once.
As well, ionice doesn't help with this degree of load.
They’re doing it correctly and would still be doing it correctly if they were on spinning disks. ZFS is designed exactly for these types of operations.
Over the years, I've run comparatively larger datasets, on significantly less hardware.
edit:
When I switched to SSDs for the first time, to give you a performance example, on some read queries I saw a 1000x to 10000x speed improvement.
This of course was on a read only secondary, long running reporting queries, no one runs queries of that nature on a primary.
SSD were an insane game changer.