Ahrefs saved $400m in 3 years by not going to the cloud
tech.ahrefs.com
tech.ahrefs.com
He's taken to calling AWS the "lifetime employment program for enterprise software developers."
But you're right that some of them may be more interested in complexity of used solutions, bigger budgets, and their personal job security than in company results and sustainability.
Cloud is amazing for highly variable loads. I also see it as a bit of a luxury service for ops types (like me) - I don't have to go to a DC and manage hardware or deal with a DCops crew, so that's nice, but you probably wouldn't buy luxury cars if you were starting a courier business.
Some can do it, some require you to basically image the entire machine onto a new one, some can't handle it at all.
First, changing out the CPU type to something that isn’t almost identical out from under a running VM is a mess and may require degrading the system by removing features from the starting CPU. Switching manufacturers at runtime (AMD vs Intel), while sort of possible in theory, is effectively a lost cause.
To add CPUs, you have to deal with the architecture’s nasty hot-add CPU mechanism, and the guest OS needs to be expecting those CPUs starting from when it boots, and that expectation isn’t free. (The latter is Linux’s num_possible_cpus vs num_present_cpus.)
Adding memory involves getting that memory into the kernel’s memory map in the right places. This is more complex than one would like. Removing memory is worse.
And all this happens, in the cloud data center, on essentially normal hardware. If you want to add RAM, the CPUs need to be on a system with more RAM, preferably on the same NUMA node. And vice versa for adding CPUs. If the tenant is paying for local storage, that needs to move, too.
The thing that really pissed me off is when a cloud provider is UNABLE to add more CPU and RAM on reboot. I even understand having to migrate the instance disks to another host that has CPU/RAM available.
Everyone's locked into a vendor at some level or another. Want to move off of Oracle? It's work. Want to move off of servers and into VMs? It's work. Do you want to move off of some Java library like Spring? It's work. Even moving from apache to nginx is work (assuming you're not just serving static files).
Your technology choices always lock you in. You can minimize the cost of moving with architecture, but there's only so much you can do.
The key isn't avoiding vendor lock-in, it's understanding what the tradeoffs are for that lock-in and making sure it's a decision instead of a by-product.
It's like writing your stuff in NEW_TECH_LANGUAGE. Sure it might give you better performance, but now you have a bunch of deployment and maintenance headaches, and your people will cost more and be harder to find. And does it even work/scale in production/HA? How do you monitor it? What kind of weird runtime things are you going to get abused by?
I do spend a fair amount of time and effort trying to keep my company nimble; we try to use open source software and generic cloud primitives so we can move with some ease if need to. As you say, it is all work.
I wouldn't be surprised if Oracle or any outsourcer did that...but that implies a breach of contract, not a transition after a contract is done.
IBM sells mainframes because they're impossible to replace due to the way they're used. Those big printers that print financials statements apparently don't talk to anything else (funny, I was just talking about those the other day). It's about I/O and throughput with those things. You'd think someone would have come up with a replacement, but presumably the market is so small that nobody cares.
Yaay savings!
Similarly, “easy” scalability doesn’t necessarily make up for COGS when scaling up. If I can flip the autoscale switch while I grow to $1bn ARR and pay $750M/yr to the cloud, I’m not doing nearly as well as if I hire ten FTEs and pay $150M in combined capex and opex over two years.
(And if my COGS is 75% of the list price, it’s hard to give discounts, pay for sales, etc)
Everyone looked at me like an alien speaking a different language, lol. Then was politely dismissed, even though I'm experienced in running hardware.
AWS, GCP, Azure, really managed to expertly pull off the greatest heist of all time. They're useful for sure, but somehow they've gotten everyone to drink this potent marketing kool-aid that if you're not using them then you're not serious about success. So cringe.
Saving $100k would not be worth the hassle/risk for most well funded startups.
Hardware these days is really good.
People have this illusion that servers are extremely hard to maintain because big companies constantly have to maintain their 1000s, and because clouds are incentivized to sell this lie.
It would still be a huge cost saving and trivial for them to contract someone to maintain the servers for like 1 hour a month with emergency on-call. Of which there are numerous such people who are not me.
I worked in an IT department as an intern. I thought this company should build custom PCs for all their workers. You can get a much faster CPU or much faster GPU for the same price! It's a no brainer!
Then I realized that if I wanted to build custom PCs for all the workers, I'd have to provide support for all of them. I'd have to be as good as CDW, who was supplying the company with PCs. This means next-day repairs and replacements. Drivers. Security. Firmware updates. For hundreds of PCs.
Needless to say, a lot of people under appreciate all the things that vendors like AWS or CDW provide. They think they can do it because they know how it's done. But they're ignoring a lot of other factors that can come back to bite you.
I'm sure you're very competent at managing server hardware. But if I'm a seed stage startup, I'd politely decline your offer.
It’s not either AWS or “bring your soldering iron”. Self-hosting K8s on managed hardware is still insanely cheaper than AWS.
GitOps, infrastructure as code, DevOps platforms, CNCF best practice setups are entirely doable without paying the Amazon tax. If your lead DevOps quits, you’re right, it’s easier to find someone who does AWS.
The last time I worked for a company that was invested in AWS, their AWS budget corresponded to five full-time senior developers in running costs.
It’s not much different than outsourcing DevOps, which I think is what some hackers frown upon, because that’s their skillset you buy in town at a premium.
> The last time I worked for a company that was invested in AWS, their AWS budget corresponded to five full-time senior developers in running costs.
This means nothing to me -- what is the TCO of the AWS deployment vs. the TCO of the non AWS deployment? How flexible are the different systems for the future? What are the reliability/uptime requirements and the marginal cost of outage? How replaceable are the dev ops engineers who are running your K8s/managed infra?
People moving to cloud isn't irrational. There's a reason that it's pervasive.
I can only speak of one PHP/TypeScript app with a poor performance profile: from $250/mo to $5.000/mo in running costs, not counting the migration. The team that made the app couldn’t maintain the Helm chart or the Terraform resources, adding to a DevOps team bottleneck that appears to limit how teams innovate (kind of how like creating new branches was expensive on CVS).
> How replaceable are the dev ops engineers who are running your K8s/managed infra?
I haven’t met a competent AWS DevOps engineer who wasn’t also competent to run the clusters; it’s just that they prefer not to.
> People moving to cloud isn't irrational. There's a reason that it's pervasive.
Yes, it’s made really solid. AWS makes sense to reduce operational risk at a large scale.
The price is not my main argument against using AWS, although you’ve gotta make good money to justify the expense.
My motivation is to own one’s own infrastructure. I like to work on things that are both important and contented enough that you can’t, as a principle, outsource the operation.
Power, air-conditioning, security etc isn't my problem. That's what we pay the colo for.
I go there about once a year to install new servers, and about once at some other point in the year to replace a failed disk. I've taken various developers who are interested each time, so if I'm out of town there are several people who can cover, plus the colo staff.
You can also pay for people to do this routine work.
And then there's still the option of renting managed servers, so that company handles all the hardware.
I've found there's generally developers who are into building custom gaming PCs or Raspberry Pis or whatever who are interested to see the servers and lend a hand installing new ones.
I do not want to be on call 24/7.
You can't get much more time efficient than data centre operators handling hardware.
Now there’s NVMe and there are quite a few form factors. And NVMe is generally a point-to-point link to the CPU, so you either need a motherboard with the right number and types of connectors or you need a PCIe/NVMe splitter or switch. And the latter don’t seem widely available on the non-sketchy-hardware market.
So the situation is fine for system integrators who make their own PCBs, but not amazing for DIY builders.
I mean, it's not like a colo doesn't have issues. At one point our colo was running some cable and we discovered that their cables to our rack were wiggly...after our stuff went offline intermittently. That was hard to track down.
And you need to set up the LOMs, the management network (hey, you don't want that on the public internet), VPN, serial cables (backup to network LOM), maybe a load balancer. And you have to hope that when you reboot the box it doesn't fail POST. And you have to remember to turn off power saving in the BIOS (how many people forget to do this?). And you have to set up the RAID 5/10.
And you have to do this sitting on the floor, because not a lot of places provide chairs in the racks. Oh, do you have a rack-mounted KVM/monitor/serial console or are you using your laptop with a serial cable to do all this?
I mean, it's not hard, it's just a lot of detail work. And I suspect that these skills are going away.
Oh, I forgot that you also have to manage the certs on the LOMs. Luckily browsers still allow you to bypass expired certs by default.
I wonder if you can still get into machines with the old SHA-1 stuff?
That's why serial consoles are better IMO.
As far as the crypto stuff, it's an issue. I fairly recently had to tweak some old 11th generation Dell Servers and tried out the DRACs, but Java really doesn't want you to use RC4 - I remember that being the dealbreaker. Nicely, someone put a conatiner up on Docker Hub that dealt with it https://hub.docker.com/r/domistyle/idrac6/
8 years ago you could already set up RAID or change BIOS settings using LOM, and KVMs were obsolete. The slightly janky Java remote access tools have been replaced by HTML 5 ones, and at least HP and Dell provide tools to script setting up a fleet of similar machines.
When I set up new servers, I photograph the sticker with the serial number, MAC address and LOM password, mostly so I can be sure exactly where each machine has been racked. Connect power, ethernet, LOM. Check the DHCP server log on the LOM network to see there are the expected number of new MAC addresses present, then go back to the office to finish the setup.
In exchange, you give up the usual cloud conveniences like easier budgeting, easier scale-up (and scale-down), lower up-front costs (which is a big deal for most early startups), better SLAs, being able to use cloud object storage without paying transfer fees, access to support, fewer things that can go wrong, etc.
Apparently I'm getting your share.
You can provision dedicated servers with Terraform.
They can still be cloudy in the way you deal with hardware maintenance and risk of failure, and how you authenticate with them and how you name them, and how you hook them up on VPNs. But they're physical units without the AWS bells and whistles or the AWS premium. Not everyone needs their own colocated rack, and manage their own UPS and network peering.
It's covered more elsewhere in this thread, and I appreciated learning about it.
I wouldn't say that renting managed servers is simple at scale.
But it is both simpler and less expensive when not at scale.
…unless they are not physical servers but vservers. Hetzner, for instance, has such a "cloud" offering where you can dynamically spin up new vservers via an API/Terraform.
They represent yet another point on the spectrum but are not "fully cloud" yet if we take "fully cloud" to mean that, beyond hardware, the software stack is partly managed by a third party, too.
Put differently, to me the spectrum looks more or less like this:
- On-prem self-hosting (own data center, own hardware etc.) - (Off-prem) non-managed hosting (data center provided by third party, but you're providing the server hardware and put it in one of their racks) - Managed hosting of physical servers (data center + server hardware provided by third party) - Managed hosting of vservers (Hetzner Cloud, AWS EC2, …) - "Full cloud": Software stack is partly managed by 3rd party (e.g. managed Postgres, managed Kafka, managed ElasticSearch, serverless, etc.)
I can recommend this article for those who are interested - https://www.cloudzero.com/blog/capex-vs-opex
https://www.itpro.co.uk/cloud/public-cloud/369521/singapore-...
Which part would actually saving them the money? - rearchitecting/simplifying, or changing where it runs? There would need to be an apples-to-apples comparison to be meaningful.
That said, I don't believe everything needs to be in the cloud - but maintaining data-centers are also very expensive, and need to be accounted for very differently.
Before it was running on a dedicated Xeon with the database available on a UNIX socket. After... where do I start.
The app was very database heavy and experienced timeouts due to too many database calls (IOPS). Not a problem prior to migration, because you don’t pay separately for UNIX sockets. Just the IOPS ended up costing several hundred bucks a month.
All environment variables were stored as secrets at $.40/mo. per variable. I know this is peanuts, but it is so illustrative of the pricing model.
It feels like those candy bags where each tiny piece of candy is individually wrapped in plastic.
I found $35k/mo I’d have spent differently.
In AWS you need to be careful what you use.
There are a variety of other expenses associated with an on-premise environment considered indirect expenses. These expenses are often referred to as “hidden” expenses due to how often they are overlooked rather than “hidden.” These include:
The real estate of the storage space used for the servers
Tools used for temperature control in the data center
The cost of set up, configuration, and ongoing upgrades
Staff salaries for administrators that maintain an on-premise data center
Networking infrastructure set up and ongoing maintenance
The cost of downtime while the team troubleshoots the issue
Productivity lost when the system experiences downtime
The cost of keeping the servers powered 24/7
Depreciation of the hardware and software
Time spent on disaster recovery
Administrative costs associated beyond IT staff such as HR, purchasing, financing, and other departments.
https://pmsquare.com/analytics-blog/2022/6/9/calculating-on-...
As for staff salaries, lol. It's not like AWS is self administering.
The figures in the article are for illustration purposes only and should be taken with a large pinch of salt. The author doesn't detail what hardware they are running, what EC2 instances he has selected for comparison and how comparable the storage statistics are. I would also love to hear the Finance Department's version of his calculations.
AWS is more expensive than self-hosting. However, it is not as skewed as the author claims. Otherwise, very few companies would be using the cloud.
It is hard to believe how high cloud costs are.
I however would like to address some of your points which I find absolutely ridiculous:
> Staff salaries for administrators that maintain an on-premise data center
Every cloud-native company I've been at had an entire team wrangling YAML files around. I'm not exactly sure why they had to do so (because everything was stable and as you say, the cloud is supposed to handle all maintenance/etc for you) but they did and cost a pretty penny.
The only case where I genuinely agree that the "cloud" saves money on administration is fully-managed PaaS providers.
> The cost of keeping the servers powered 24/7
A non-cloud-based, peak-capacity-sized deployment costs less to run than a cloud-based deployment scaled at minimum capacity. Servers are super fucking cheap nowadays. Hetzner will happily sell you a 16-core, 128GB of RAM, 4TB redundant NVME SSD machine with 20TB of included external bandwidth for ~150 bucks a month. AWS will cost more than that on bandwidth alone.
The irony of it is most of the cloud services are open-source technologies packaged up to be easy to use and administer from a web interface.
Maybe if someone came up with a set of tooling in AWS to use many lightsail VMs to provide more of the basic AWS services, it could be a way to get the best of both worlds.
But alas, nobody will ever go for it and I'm left to dream.
Performance. You'd also get way faster hardware too.
Enterprise-grade SSDs is what I meant by direct-attach storage - I was comparing against network storage which is what all cloud providers use (your EBS volume is accessed over the network internally, which incurs some latency, ultimately limits IOPS and destroys random access performance which can't be cached or read-ahead by the underlying hypervisor).
“multiple nodes only give you reliability, not performance” - this, actually, is not true. Even mysql can offer ways to tune and gain performance improvements.
Presumably they would also lack the durability guarantees of the normal block storage offered by the provider.
What about the specialized services from AWS provides leverage outside of the “cattle mindset”?
In my experience, AWS’s primary value and leverage is based on intermittent burst compute. It’s why you rarely get a hard answer on the Hz of each vCPU and have credits to the overuse or underuse of said instances. The cpu is then used to arbitrage value in the dedicated services like lambda and others - so your mindset and infra needs to match AWS’s incentives to get the same value, correct?
We use ~50 fairly high performance (RAM, CPU, disc) servers pretty much continuously, so the savings compared to AWS etc are considerable. Every couple of years there's a software upgrade required, and it would be convenient to have another 50 servers to use during the transition, just for a few days.
Compared to AWS prices, we can easily justify keeping some old servers around to help with this sort of thing. A managed server company (Hetzner etc) would be somewhere in-between; presumably able to rent us 50 extra machines for the duration, but (last time we checked) still more expensive than managing the servers ourselves.
This could be only the case is t* instances? I learned the hard way that CPU misses can add even 30% of request latency, especially when you often do I/O eg to external services like db/Redis/etc
I didn't notice such behaviors on AWS. but when thinking about it, I read somewhere that one reason they have dedicated CPUs is for partitioning CPU cache so it's not shared between VMs. so maybe with some tricks they could get some free compute without side effects?
I use Ansible at home as a replacement for the 1990s style of sysadmin where you manually managed your servers (which I did until the mid 2010s). Yet, I’m aware that Ansible is not declarative. It’s looks declarative, if you take your glasses off and don’t think too hard about it, but the key to understanding Ansible is really that the configuration is idempotent (not declarative).
Servers are cattle either way, it’s just that with Ansible, your servers will keep changes that you made months or years ago that have been deleted from your playbooks. That, to me, is what makes it old-fashioned.
I don’t think old-fashioned is bad. I’ve gotten into arguments where people have told me to manage my personal server(s) using K8s and I, politely as I could manage, told them to fuck off.
It sounds like you should make a business case. 90% savings is a big deal, unless it's the difference between $1k and $100.
Will you have redundant hardware running to failover to? What about disaster recovery, are you going to colo in more than 1 physical data center?
Who's going to be on-call: the same people who are on call now. HW failures don't need special on-call people.
Who is going to build/manage your monitoring/alerting stack. Same people as now. Your servers are monitored, right? Adding RAID into the mix is hardly difficult.
Will you hire a network engineer? No, colos will often offer to manage the network for you up to the switch and at that point it's just a matter of plugging the server in. You don't need any expertise to do that and remote hands can do it for you.
Redundant hardware? Only if you need it currently. HW is reliable these days and you'd going to have spare capacity anyway, so this really doesn't change much.
Disaster recovery? Take whatever answer you currently have in the cloud for multi-AZ.
In the cloud, not in the cloud, frankly it seems to me like their business would be far more efficient if they tuned what their crawler actually crawled.
My favorite design pattern is a for loop that tries once per item and completely fails to never try again, nor report it failed at all. /s
I think people see the notion of a network neutrality as an all or nothing thing. Network Neutrality between networks, yes, but within your racks you’re trying to balance a large equation instead of two or three smaller ones.
If they reclassified sketchy pages they wouldn’t have nearly as much of the sorts of trouble you mention.
E.g. this article is missing even basic stuff like the (prorated) salary costs of employees buying, installing and servicing the hardware.
You say that like you don't need a dedicated team to manage AWS.
Sometimes all people need is a big box in a data center. Management of the environment is already a sunk cost. The hardware is already expensed. They don't care about scalability/DR/etc.
From a hard dollar point of view staying colo makes sense. Plus politically they don't want to do it, which is really the only point of view that mattered.
If they needed to have actual DR then the numbers would be different.
I don't miss cycling around Manchester on a Bank Holiday weekend because I'd miscalculated how much network cabling I'd need for an upgrade.
I don't miss keeping a spreadsheet of storage so I knew when to order disks, negotiating with suppliers for cost of new disks because I was buying a slightly smaller bulk than AWS
I don't miss having to explain to folk in datacenter support that they could take the disks out of my failed server and put them in a new server if they had one available
I don't miss the day the single point of failure in the rack failed and everything was offline while I waited for a new doohicky to be shipped to me because it didn't make sense to keep spares of everything on hand
I don't miss trying to figure out if some new generation of server hardware would work for or would fit in my rack as manufacturers stopped making the kit we did use
I'm not going to say every workload should run in the cloud (cliche nod to StackOverflow) but it certainly isn't free to get all of the benefits
Really, the main things that cost actual money in AWS are memory and CPU (that's including elasticache, RDS, etc). Bandwidth can be negotiated away. I'm sure you could do some kind of super deal on CPU/RAM, but we've never bothered.
Would it be worth it to rearchitect your app so it doesn't use 2TB of RAM? Probably not. You have a lot of sunk costs, and redoing everything will probably break everything. You guys are big enough that CapEx doesn't matter that much. You just need big boxes with lots of bandwidth.
If you're happy with what you've got, stay with it.
> So let’s say we run our 850 servers
I'm not familiar with this company, but it says on their website We’ve been crawling the web for over 10 years, collecting and processing petabytes of data every day., they have their own search engine (named 'Yep), provide all kinds of reports, etc.
I'm guessing they are running huge Hadoop clusters or something? One of those workloads that just isn't suited for the cloud, and they're just taking the advice of Joel Spolsky
“If it's a core business function — do it yourself, no matter what.”
I think calling them "just a bunch of hadoop cluster" might be disservice to the amount of effort that goes into building a product like this esp. When millions of digital marketing professionals use their service.
The only person ther said “just” is you. And “large Hadoop clusters” can be the most resource-intensive component of their infrastructure, but not necessarily the company’s value-add, or where most of their effort goes.
Stop being adversarial.
Obviously when you started your stuff didn't need 2TB of RAM. If I read your history correctly I don't think you could even buy a box with 2TB of RAM back then. That's enterprise grade hardware, which today costs a fortune. Back in the day it would be a bigger fortune.
Instead, you probably started the way everyone else did, with maybe a 4GB or 8GB linux box at home then built things up from there and a bunch of curl scripts and a local instance.
So why the resource requirement? An in-ram database?
Mainly curious.
If it's a big Hadoop cluster or similar, there's potentially huge performance gains by keeping data in-RAM during processing. They have many petabytes of data, so I doubt the type of processing they are doing today would have been possible when they first started.
Also, Ahrefs bot doesn't handle some things very well. I made an "infinite web site" in PHP a while back, and used Apache mod_rewrite to send every Ahrefs request to the infinite web site PHP program: https://github.com/bediger4000/infinite-fake-website
Ahrefs bot really freaked out, unlike some professional bots like Google's, and even Yandex' bot.
> Also, Ahrefs bot doesn't handle some things very well.
But that was... 7 years ago?
But with the mass layoffs in Big Tech in recent months, this may be an opportunity to re-evaluate the approach to the cloud, consider a reverse migration from the cloud, and hire seasoned professionals of the data center world.
I do think with all these layoffs we can't help but see some interesting things happen from people who have been let go. Either through their work at other places, or new startups.Source: managed systems locally, am now managing systems on AWS
And very less work is required if I get hit by a bus and someone else is now the owner of the huge system that I previously owned.
Also, all of your points apply to non-cloud systems too.
AWS & Co. use cases always look nice on paper, in reality it’s often not as easy as advertised.
I actually saw a very impressive solution to that problem on an indie hacker forum recently. And wow could that person save effort with even just a $5/mo VPS.
That being said, if RDS can do what you need USE IT. You can always migrate later and it will save you the need to hire someone who really knows postgres.
hydrogen033-ext2.a.ahrefs.com
hydrogen042-ext2.a.ahrefs.com
hydrogen106-ext2.a.ahrefs.com
hydrogen326-ext2.a.ahrefs.comhttps://tech.ahrefs.com/how-ahrefs-saved-us-400m-in-3-years-...
Hosting yourself makes sense if you're providing hardware level services like storage or compute. In that case, going to the cloud is literally financing your (potential ) competitor.
If you’re a business leader, you‘re probably just blindly following what your peer business leaders are doing, that’s why you’re on AWS (in most cases at least)
What I'm struggling with is their estimate. I work for an enormous enterprise that runs tons of stuff on the cloud and our budget is less than a quarter of their AWS estimate. We avoid products like EBS unless they are necessary, and use RDS whenever possible.
I worked for a small (~300 headcount) software company that did CFD software. We were told to build out our HPC capacity so that customers could use our hardware to run their jobs instead of having to manage a cluster themselves to run our software. The software was billed per core-hour, and that was the only charge. It made no difference if you ran it on your infrastructure or ours-there was no additional charge for our compute or storage or network bandwidth.
We bought compute in units of between one and four racks fully populated, usually lease with a trivial buyout at the end, or just outright.
In the last 24 months we were an independent company, our SaaS infrastructure drove an additional $24 million to EBITDA. In that increment, we spent $9 million total on hardware, colo, network connectivity and our salaries. The total cost of replicating our compute capacity (for those 24 months) on AWS was ~$31 million. This all came out on the due-diligence that we had to do as we were being bought by a larger firm, so the accountants were satisfied that the numbers were accurate.
IOW, the article seems to be completely plausible.
Also it looks like these are retail prices. You can get cost savings by negotiating but I’m sure there’s an NDA required.
For sure Cloud is more expensive, as it charges a premium for not having to invest upfront in a lot of hardware, however,if it was 1000% no one would be using the Cloud.
I have a fixed day or two each year working out what to buy and getting quotes (which would be the same time whether it's 20 or 200 servers), half a day per 20 servers for installation in racks, and call it half a day per 20 for initial configuration (most of the time to script the process, so it would also be similar for 200). After that, it's at absolute most an hour every couple of months to deal with a failed disc or similar.
For sure Cloud is more expensive, but if you quote 500%, you are overlooking hidden costs.