A lot of the things you said apply just as much to them as to AWS.
In my opinion using AWS is fine, as long as you use cloud agnostic stuff and avoid AWS specific services. That way if you ever have to switch, you just have to rewrite the provisioning part of your infra-as-code, everything else can be easily migrated over with some config changes.
On the other hand, if you use Kinesis, Lambdas, RDS, VPC peering, and IAM accross all of them including EC2 and EKS, then no, none of those other options can even come close. Purely the lack of system-wide IAM alone would be a dealbreaker.
This is also why being 'on the cloud' and 'cloud native' is very different. While the latter is a bit too buzzwordy for my taste, it does show the difference between "playing datacenter" and "getting stuff done in new ways, both cheaper and faster".
Do not use AWS if you just run a few virtual machines and some basic networking. That also applies to GCP and Azure. It's too expensive for what you need, and too easy to misconfigure and shoot yourself in the foot.
On the other hand, if you need API-driven infrastructure with a large gradient between IaaS and SaaS to pick from, then by all means, use AWS.
Hence “you probably don’t need AWS”.
> On the other hand, if you use Kinesis, Lambdas, RDS, VPC peering, and IAM accross all of them including EC2 and EKS, then no, none of those other options can even come close. Purely the lack of system-wide IAM alone would be a dealbreaker.
Hence you are locked in and have zero negotiating position. Your business has an existential risk based on AWS - assuming your code is critical to the business and could not practically be rewritten because of all this special sauce you’re using.
That might be okay but I sure hope your CTO made that decision consciously.
Sure, you can play FUD games about AWS suddenly raising prices, but I doubt you'd have any historical evidence to back that up. So you'd better be able to say why being cloud-agnostic is critical to your business.
You generally pay for what you need and for what you can afford. If you don't need AWS or you can't afford AWS, it wouldn't make sense to use AWS. That goes for in-house software development, infrastructure management and pretty much anything else too...
In my startup we use mostly linode and handle a lot of volume. I have a few VPSs and some dedicated hardware for specific tasks. Monthly bill roughly 200-250USD all included.
The other startup used the AWS free credits. Got extra credits and lived it up. Adopted Kubernetes and all the big AWS stuff. Brought in a few devops professionals. All super smart people.
Monthly spend 7kUSD on traffic/usage that's actually much smaller than the one I have with my startup. Availability is roughly the same (if anything, my startup has better uptime).
So yes, this is an Apples to Oranges comparison. I get that. They can better scale in theory etc. But if they pay 7k right now, I can't imagine how expensive it will be when they scale.
They started just using EC2 but slowly got roped into AWS services so now they have no vendor neutrality and no viable/easy way to reduce these costs.
If you have less traffic then you should spend less.
This sounds like they don't have anyone who actually knows what they are doing, or did any form of cost/benefit analysis which will always mean you spend more than you should.
If you need 'some random compute and a bit of networking' then no, do not use AWS.
Vendor neutrality isn't worth match to a business unless you are the EFF or intend to move vendors all day long.
You don't need AWS or any other cloud for volume (traffic volume, that is), heck I'd say stay away from the clouds for volume since that is where the bulk of your money disappears anyway.
Simple things are simple: need IPv4 networking? Get an IPv4 allocation and route it to a (virtual) port of your choice. That is no different between on-prem, managed-dc, unmanaged-dc, managed server provider and cloud provider. There might be some semantics that are different, where a physical port cannot be unplugged via an API (at best you can down/up a switch port but that's not the same) and routing may or may not be implicit, may need a firewall or routing rule and may need translation. But all of that applies always anyway and gets named differently on who you are paying for it.
If you create something very complex, but it didn't have to be, perhaps it's simply a bad implementation and not a property of the technology that was applied. There are plenty of engineers that think they can roll their own crypto libraries, write their own filesystems and their own RDBMs... but what business case for the 99% out there really requires that? Nearly everything in a business is a generic problem any business has; it's always about products and/or services, and it's always bound to some legal and financial system the business operates in. Everything beyond that is just marketing and trust building.
The only unique thing in most cases is the data you have and the degree of integration of processes within the business. And neither of those are an actual technology problem or complexity driver.
As soon as you have decent amount of data which you care about AWS became dirt cheap compared to on-perm when user right. Look at the list above and imagine how much effort you would have to put on an ongoing basis to support this.
I'm not sure if you're being serious here. It's not about rising prices but making your business completely dependent on another entity and losing any options you might otherwise have.
As for past data, history is rife with stories of companies who made the mistake of trusting a big player their technology will be around forever. So many have invested in technologies like Flash because they were 100% sure they will be there forever.
If I was in charge of AWS, I would avoid sudden price increases, but instead focused on many new features and small charges. These look insignificant, "This is just 20 buck per TB, less than our spending on toilet paper." And then these slowly accumulate and the TCO is getting higher and higher, and again some new* smart features lure you - again, at a small cost. In the end, it turns out you can't leave even if you want.
*Or just your old open source library/framework repainted in AWS colors.
For example, let's take a simple producer/consumer system with some sort of queue. That's an easy enough problem. But add in some additional requirements:
* Data encryption - both in transit/rest * Access control * Audit logs
And it becomes much more complicated. It is a trade off. But the productivity gains available by having a common consistent solutions to issues like this... shouldn't be underestimated.
Doing it "right the first time" in one of the big three, then building a skunk works to replicate efforts as a blackout scenario, seems much more cost effective.
The key term is Total Cost of Ownership. Lockin is a cost. Legacy stacks are a cost. Learning three times the tools are a cost.
There is no real lock-in other than the requirements of your own implementation. A lambda on AWS can just as easily be run on Kubernetes, but that means that on top of running the code that you wanted to run, you now have to run Kubernetes too. You could use a hosted Kubernetes control plane, or fully hosted setup, but according to your same logic you would now be locked in to Kubernetes.
It seems to me that you might have a very special use case, or simply not have had the scaled experience or requirements that made a lot of other vendors non-viable.
Business risk-wise, AWS is a very low risk when compared with pretty much anything else. What do you think will fail first, a custom in-house maintained setup (expensive, non-portable between companies and people) on Digital Ocean, Linode, OVH, Hetzner, etc or all AWS Availability Zones in all AWS Regions world-wide? (and we ignore the lack of IAM, lack of things like RDS, SQS, SNS etc)
You don't have a negotiation position based on your ability to switch even if you are huge. AWS pricing negotiations is almost exclusively based on volumes.
> Do not use AWS if you just run a few virtual machines and some basic networking. That also applies to GCP and Azure. It's too expensive for what you need.
Currently on a contract where I've saved them 6 months of my costs over 3 weeks of part-time development on a single year long project. They saved half of my entire contract on their first foray into AWS compared to Digital Ocean charges.
When you know what you're doing AWS can become very cost-effective. Build it with $PLATFORM in mind and all sorts of savings are possible.
But the tricky part is whether the team is ready for it, or whether they'll spend more time learning the ropes than the company saves over $NUMBER years. Which I think is what you were getting at?
Another term which is more descriptive for cloud native is vendor cloud lock-in, which could be reasonably but descriptively shortened to cloud-locked or cloud-married maybe.
Native sounds cool but completely loses sight of the cost of lock-in, while locked or married talks to both the benefits and the costs.
If you use those, AWS starts rubbing its hands together and cackling in an evil laugh. I bet they even have an mp3 that plays whenever a major player uses a lock-in api product on AWS.
"aws lambda install blah blah" < I obviously don't know the api
Somewhere in Seattle: Mwahahahahahaha! Bwahahahahahah! Kwahahahahah!
Although RDS is just hosted SQL databases. It isn't a pure lock-in. Identity management, key management, firewalls, and networking all exist in some form at most places.
And don't underestimate how bad internal "clouds" are in large enterprises. AWS pure EC2 can still be a huge win, especially if that is the only one that has been "expertly" vetted or negotiated by some lofty level of management on a golf course somewhere.
I use AWS, but I suspect I'm a fair bit older than you.
Those things you say "you can't come close" - they exist elsewhere. There are just software libraries. It's the reverse of "can't come close" - there is a plethora of implementations to choose from, and most are free. Lambdas is just CGI on steroids, RDS is just a database, IAM is just one of any number of authentication / authorisation platforms the most notable one being kerberos which is literally decades old.
It's true that AWS's versions of these libraries all play nicely together, whereas the others are going to require more glue. But on the other hand AWS nickel and dimes you on every use. The price of writing the glue can be high and has to be incurred up front, whereas the price of the nickel and dimes - that's a future self problem. Just like smoking and crack. Not unsurprisingly the cloud vendors use the same tricks as the smoking and crack vendors too - they give away their delights for free to early users by handing out free credits like candy, knowing once you're locked in as a heavy user there is no way to avoid paying the piper.
The price difference in my experience is typically of the order of an an extra 0. That's not a big deal with your primary costs aren't CPU cycles but rather developer time. That's absolutely the case at the start of an enterprises life, and it seems to remain true for many enterprises for their entire existence. So perhaps AWS works out for many.
But personally, I would not take it for my business. There are lots of implementations of all the things you mention, and much more besides out there. Worse, many of AWS's products seem to exist to solve problems created by using other AWS products. Need to monitor you SES reputation? Don't record it yourself - send it to CloudWatch! It's only a small cost. Oh you need to get it to CloudWatch - use SNS - it's only a small cost. Oh you want to trigger stuff from CloudWatch - just use Lambda - it's only a small cost. Oh, you need to control access to all those things - use IAM. Maybe it's true it's all only a small cost - but it a very complex solution compared to parsing you SMTP server logs and writing them to a open source database, which is all that is happening under the hood. At times it seems the only reason AWS structures things in the way they do is to induce you into buying more AWS services.
Which is why things like TerraForm exist. Some are happier to where the cost incur the expense of gluing together the many free alternatives out there, in return for the ability to roll it out to the lowest cost provider out there.
That is a big tradeoff, and it's not something obvious. If being banned from AWS is an important and relatively big risk for your business, sure, go for it.
But if not, then those AWS-specific services are exactly where you get all those productivity gains from, and where AWS (and GCP for that matter) shines in comparison to smaller ones, like DO or Linode, or on-premises.
I'm sure there are tons of developers out there who know a lot about infrastructure and system administration. For those people, it might be worth spending time managing infra, but my time is absolutely better spent using a turn key solution and focusing on product development.
For smaller teams, getting feature-rich products out the door can be more important than building the most cost-optimal product. Bigger teams can afford dedicated staff to optimize infra.
So unless you're going multi-cloud from day one, it's much better to write your services to be specific to the cloud you're using and benefit from them rather than boycotting them because of a theoretical decision of multi-cloud that you might never actually make. Most companies that talk about going multi-cloud never actually do because the benefits of being multi-cloud are actually pretty small for most domains. Being multi-cloud is one of those things that sounds more important than it actually is.
IMO this is exactly when you should not be using AWS. AWS is a premium offering and you can find cheaper alternatives for the commoditized offerings like EC2, RDS (the older, non-AWS-specific features) and (maybe) S3.
It's the high-level services where the real value is in AWS - the cutting edge technologies that only exist in a specific cloud (or multiple clouds but where it's difficult to port across clouds due to subtle differences). This is where the extremely high development cost is able to be amortized over many customers, allowing you to get scalable, easy-to-use, and easy-to-operate services at a much lower cost than would be possible outside of that cloud. A great example of that for AWS is Serverless Aurora.
This is the worst of both worlds. You get the expense of AWS but without any of the things that would make the cost worth it. IMO you should either avoid using their services and move to something way cheaper (OVH, as the article mentions Hetzner), or fully buy into the benefits AWS has to offer.
Back in the ‘old days’ of the cloud you’d hear a lot of talk about “commodity computing”. Compute was fungible. No salespeople, no politics. It was a utopian environment for developers.
But in the decade and a half since, the landscape has changed. Now cloud is infested with salespeople, consultants, and worst of all, third-party solutions. It’s the same old hodgepodge of half-finished incompatible products tossed over the wall that on-prem used to be, only more expensive.
Infrastructure as code is something you can do on smaller deployment environments as well, even in air-gapped environments.
The benefit of AWS, for me, is the reliability and all the capabilities that work out of the box that aren't differentiators for my solution (so it's great that I can offload them to someone else, if security, etc. allows).
I agree that it's a good idea to have a layer of abstraction so you can use other service providers, even though you may take a bit of a performance hit for doing so.
I remember using their email service (SES) and wanting to read logs for emails that have been sent. The only option was to point log events at an SQS queue. So, you had to create one of those, assign proper roles and access controls, etc. After that, you need to dump or consume the queue somehow. I can't remember if SQS could dump to S3, but if it could, then OK: create a bucket, assign roles, etc. Now, to parse the S3 bucket contents somehow... Wait, what was I doing? Trying to read some damn log files or set up a Rube Goldberg machine for the log files?
I get it. It's a modular system, and that can be very powerful. Sometimes I just want to do something simple without going down a multi-hour (or day) rabbit hole. The inability to perform simple actions on AWS without looping in tons of their other products means indefinite job security for a lot of people. IaC and templates for actions like the above alleviate this a bit, but it's still a dreary landscape.
They can offer awesome leverage and benefit. They can also be horrible traps that can break a company. I've seen both. As Doug Comer once told me, Unix is a Power Tool. Power Tools can kill. So can cloud platforms. There is nothing new under the sun...
Total cost 30$ per day ~900$ per month.
You can accomplish the same for a fraction of that cost.
The convenience you talk about comes with a huge bill.
If I add reserved instances/savings plan, probably gonna drop to 600-700$, but still… I can setup that myself on my local lab for around 5$ per month.
So the knowledge required to setup that and hardware is worth around 695$ per month.
Thats why Cloud Providers are biggest employers of old sysadmins ;)
>Total cost 30$ per day ~900$ per month.
Ummm... I think you're doing something incorrectly. The biggest line items on that list are ELB (assuming you mean Application Load Balancer/ALB) and NAT Gateway. But with that setup, I could maybe imagine $100/month with on-demand plans. Not $900.
Managed NAT Gateways can be a very substantial cost depending on your architecture. I believe the recommended VPC layout has 3 AZs so you're paying close to $100/month for a single VPC just for NAT.
Do you really need multi-multi-region failover redundancy? Do you really need multiple NATs in every region (it's easy to share one NAT Gateway across AZs, especially so if a region is only servicing DR)? Your numbers still sound high.
x $1 + y $0.60 to make up $30 is a _lot_ of regions/AZs.
How much compute can you rent anywhere for $5/mo? How much compute can you get even buying the hardware and amortizing over its useful lifespan for $5/mo (~$300 total)?
And it sounded more like a generalisation anyway. ;)
either way my biggest problem with services like AWS isn't really cost but that you end up using all of their SDKs and services out of convenicance and it becomes hard to track down or move away from.. it spiders out of control so fast that sometimes its better to spend an extra few weeks doing things the harder way that's not always an issue though espcially for startups
If you want a simple sight with just core aws, you can set one up on aws on elastic beanstalk for a few bucks a month. Add an rds for 15 a month. Why will you needlessly overcomplicate the system and then cry foul?
Even if it was 50$ per month, its still alot cheaper.
And the cluster we are talking here about is a raspberry box with 10 mashines each 2cpu 8gb, running k8s, replicated storage, my own domain and certificate, metalb load balancers, and few other things.
Its not as redundant as AWS, but it never broke in few years of constant work. Chances of all of those machines to break at the same time are astronomically low in normal circumstances.
If I spread the cost of that setup over two years of reservation, the monthly cost will be a lot less than 50$.
Lets take a small production cluster of 4x r6g.large (default proposal for prod) (0.167$/h) deployed in VPC in private subnet (x2 for HA)
That’s 16$ per day for instances only and 2.5$ per day for nat-gateways per day. So 18.5$ per day and ~550$ per month.
Thats ofcourse if you disable all metrics all logs all alerting etc and it does not include Volumes cost, because its highly dependent on the domain.
AWS is expensive because you get alot of administrative work out of the box (snapshots, restores, etc). Tho sometimes have to still fix in the source. Ive personally pushed quite a few fixes to OpenSearch source, coz version 1.0.0 had a lot of issues.
That also means that a side-project that doesn't need that redundancy would be a mismatch in terms of requirements anyway. On the other hand, if you need a multi zone automatically available system, you're running out of choices really fast.
I think what many fail to realize is that - surprise surprise! It isn't that simple. Many think "if my zone goes out, I'll just migrate to a different zone!" Nope! AWS doesn't have the capacity for it - as we've seen many times, AWS can't seem to handle the migratory load when a zone goes out. Sure, the other zones don't really die, but good luck migrating your workload to them.
But when touting multi zones, no one ever mentions this little nugget of information.
All AWS had to say was “we are working hard on providing additional capacity”.
The same often happens on black friday, where companies scale up their platforms just in case because there might be no capacity on AWS.
It does cost money, but then again, so does not running certain processes. The trick becomes calculating the intersection at which point the costs outweigh the benefits, and that calculation applies everywhere.
Some sort of manual active-standby configuration really doesn't require AWS or a Cloud, that stuff is the same 90's implementation it has always been and practically boils down to attaching your RAID1 USB HDDs from one PC to another PC and booting that bad boy up as 'failover'. (yes, that's an example, and yes it's an extreme one)
If you have capacity planning, and you plan accordingly, you take service provider limits into account, just like you would with anything else. Having two power feeds into a distribution warehouse doesn't help much if neither can't handle 100% of the load in an industrial park. So while having two feeds might seem 'redundant' to a single tenant or customer, it's only really redundant if either can supply all the demand of all connected customers.
The same applies to fiber connections, plenty of fake-redundant connections that are suggested by customers to be 'redundant' turn out to end up at the same PoP and if the PoP goes down your redundant fibers are worthless. In the same logistics distribution scenario, your trucks can't deliver goods if the destination warehouse itself is offline, and now you need redundant warehouses.
That's obviously a weird thing to do at smaller scales, but the fact remains that AWS having an AZ go down is only a small piece of the puzzle, and only really a problem if you didn't plan for it appropriately.
AZ redundancy is like storage backups or backup power, if you don't regularly test it you don't have it.
* AWS MSK (Kafka server) with redundant instances across 3 AZs
* Aurora Posgres instance with redundancy across 3 AZs
* Multiple Lambdas for incoming messages from MQTT to Kafka
* IoT/MQTT and about a dozen iot devices sending telemetry continuously (every second)
* 20+ services with nearly 40 tasks, running on ECS
* All the supporting route53, certificates, VPCs etc. for the above
* A bunch of miscellaneous stuff I've forgotten that's used by other developers.
* All deployed via TF
_THAT_ all costs well under a grand. You're doing it wrong.
My point was just that it doesn't have to cost over $1k on Amazon to run a combination of services which apparently CAN run in a private lab for $5/mo, however that is calculated. I can't run my full stack even without load (due to memory commit) on my 32GB, 3 year old Xeon workstation, and that would cost a lot more than $5/mo at any provider I am aware of.
If you wrote exactly the numbers of instances and types -> we could have a real comparison then.
And yes, bandwidth costs can be brutal, especially if you try to build a hybrid solution that uses heavy services that live outside of AWS.
That said, the above comment didn't get into enough detail to get an accurate read on their bandwidth or CPU/Memory/Storage utilization, but if they can run it in a lab for $5/mo (which I doubt, since that barely covers power), then it can't be much, so I think it's fair to assume it's at least back-of-the-napkin similar scale.
That's an extremely cheap price to pay if you had to hire someone to do it.
I get the pain if you’re at “garage with shoestring budget” stage but $1k/mo isn’t even worth the time to set up a billing alarm.
If you have high compute needs I've had good experience with that and it's cheap. You can put a huge box behind it. I did that (I have ATT fiber to the house) and for the $7/month I could still use all my deploy templates and AWS CLI commands and code to do stuff.
I tried this with an AMD home machine (AMD 5900x / 64 GB ram / nvme) and it worked surprisingly well as did a dell server (beefier). I put ubuntu on them, registered them, and they came right up in the plane.
So yeah, what you're saying is sort of meaningless without understanding the business type/stage/revenue/costs.
What did the hardware cost, and based on depreciation, how long will it last before requiring replacement? Is that also included in the 5$ you are quoting here?
AWS CDK is unparalleled experience. Truly a multiplier in productivity that I haven’t experienced in 20 years working on the web.
>Is it actually good as default choice to host a software system of any scale? Let us go through some arguments against going all-in with AWS and why many companies and individual side-project developers might be better off using something else.
To be clear, the author is suggesting that AWS doesn't need to be your first choice in smaller setups and side projects. I don't think they're intending to say people shouldn't use AWS under any circumstance.
There are plenty of smaller cloud service providers that can offer similar functionality to AWS that probably make more sense for smaller-scale uses. I personally love Linode, and you can definitely pull together their various services to build something without doing it from the ground up on VMs.
I think you've lost touch with where the non-AWS world is up to and are assuming it's still where it was in 2012. It's not.
They had considered switching to AWS or Azure in the past, but ultimately deccided to stay on this setup for cost reasons, and frankly I don't think anyone understood the benefits of infrastructure-as-code let alone taking advantage of all the other thing in the cloud providers.
I came in to help modernize it. Just as we were getting going with some Terraform-based proof-of-concept, a couple things happened: some sales opportunities came up that would require running in another region, and there were some harder questions on the disaster recovery plan (eg: "what happens if the datacenter burns down?"). The company doesn't want to invest in building out another traditional datacenter, let alone redundant ones, so we took this opportunity to discuss shifting this project over to AWS or Azure.
Once built with infrastructure-as-code on the cloud providers a bunch of these problems just disappear, even without making the app itself "cloud-native". Having a dev environment that mirrors production is easy. Having multiple multi-tenant environments in different regions -- and adding new regions -- is also easy. Recovering from a major disaster would also be relatively painless, provided you have a good backup strategy (deploy to new region, restore backups).
Thinking about the resource cost to prop-up, maintain, and cross-integrate a platform anywhere close to AWS is a business in itself.
Using a declarative language is both a pro and con.
Pros of that solution is that you are restricted in what you can do, which often results in simplier and cleaner code.
Cons are that again you are restricted in what you can do.
Now about CDK is just your preferred language library to write infrastructure in imperative way. For example you write java code that deploys infrastructure.
When using a real programming language you have ofcourse a lot more freedom in what you do, but ofcourse that comes at the cost of complexity. If your developers are “creative” it can be very hard to understand wth is happening in code.
—- So yes many ppl consider Terraform as a golden standard (me included) but after many years of working with it, like anything in IT it has own flaws and you may want to consider other solutions for specific projects.
The file is still declarative regardless, and the functions you write to create that file don't per se have side effects outside that file
So without knowing anything about cdk, even if it is declarative, I don’t think your conclusion is applicable.
Terraform, the language, works differently. While it has variables, you can not store state in variables and depend on side effects based on the order you read/write such variables. Every assignment and execution order of resources is purely based on dependency order and not based on the order of the lines in your code.
One can configure the self same resource model using JSON, which can be trivially generated - as you correctly point out - using whatever language and paradigm you like.
However, since the language doing the generation does not influence the order of operations when applying the effects, the model is declarative. This is identical to the CDK configuring the CloudFormation engine via YAML. Indeed, there is a CDK for Terraform too...
- CDK compiles Typescript and other languages down to CloudFormation Templates
- Terraform uses the AWS API directly via an AWS-specific provider, and does not typically use CloudFormation stacks (though this can be useful, in some circumstances!)
- Pulumi uses TypeScript and other languages to drive an engine similar to Terraform (and the Terraform AWS provider is one of the options for provisioning AWS resources).
Hopefully this clears it up?
But this has been shared here before and is still pretty good: https://adayinthelifeof.nl/2020/05/20/aws.html
It covers a small subset of AWS services but it is very comprehensive in what it covers.