Why We Moved from Amazon Web Services to Google Cloud Platform
lugassy.net
lugassy.net
I haven't been dealing with AWS very long but I keep hitting limitations. Stupid limitations.
Sometimes they're a big deal (IPv6 is not available natively. WTF, doesn't AWS run half the web?). Sometimes they're not a big deal, but still annoying (DNSSEC isn't supported by Route53). Sometimes, they're unbelievable (AWS Lambda is Python 2.7 only, no Python 3, in 2016).
Sometimes it's probably something that would be fixed with a single line of code (Ed25519 is not supported as a keypair). Sometimes, the product itself just sucks (Cloudformation), requires a Ph.D (IAM) or has ridiculous limitations (certificate manager). And sometimes, it's just the aws cli that, in between its million subcommands available, cannot do basic things like creating S3 redirect rules or update ACM certs.
But all the time, it's the web UI that sucks the most - I have to use a separate browser for cloudwatch logs and S3 because the native UI is just that bad. And god forbid you enable 2FA, you'll have to get your phone out twice a day.
These are all the limitations I hit in a few months using AWS. I wonder how many more I'll hit within the next year. Is Google actually any better? Seems to me it's just different kinds of limitations.
I mean, what it's doing is hard and I use it over, say, Chef... but it's not a "good" product. Heck, you cannot even truly validate CF templates without running them. The CLI and builder just do a crude structural check pass. Got a case error in a nested structure? Too bad, the entire job is rolling back.
I'm not going to straight up recommend using it, but I will recommend checking it out. It's certainly a lot better than CF, but it still has really glaring issues and very much feels like alpha software (see discussion here: https://news.ycombinator.com/item?id=12213935 and some of my complaints here: https://news.ycombinator.com/item?id=12214358).
I strongly dislike HCL, but that's essentially what it is - templated JSON.
I call it "pray and apply".
https://aws.amazon.com/blogs/aws/new-change-sets-for-aws-clo...
I recommend against making sweeping changes to your infrastructure without thoroughness. Without saying this is what you're doing, I see a lot of developers who get involved with cloud treating changes whimsically. I'm not saying you should write perfect code. But thoroughness and a solid procedure (and backup plan) is a good place to start.
(also, it helps to use nested stacks)
But yeah, CloudFormation just sucks. It does. It's painful because it's a really good idea but the configuration format is inadequate for the task.
CF should have been implemented using a declarative DSL similar to Puppet from the start. I've been tinkering with the idea of creating a DSL with a JSON translation layer but I always assumed AWS would eventually move away from JSON anyway, making the effort futile. Four years later and still no improvements. Definitely my biggest disappointment with their platform.
The format is a weak argument against cloud formation. It has some negatives. This one is easily overcome.
Terraform is great so far. I am running into some duct-tape issues with resources shared between two terraformations, though.
Important tip: If you use one of jetbrains IDE (IntelliJ/PyCharm/...) there is a free plugin to handle terraform files.
It gives the usual syntax highlighting + auto completion of variables and devices + detection of errors. It's very very nice!
If you need managed services Azure and Google Cloud are largely better. If you just need compute and transfer Digital Ocean and Vultr absolutely destroy AWS.
So far it's been working well and support was very responsive on the one occasion when I needed their help.
5$ SSD VPS, 768M of Ram instead of 512 at DO, the reason I moved in at first.
The popularity clearly helps DO keeping up with more 1 click apps and other features than Vultr, but they recently made available reserved IPs and Object Storage, so they continue to be a good alternative to DO.
Google's security team is probably as big as DO's engineering team.
tl;dr: you get what you pay for, and you don't always need our crazy network!
I know it's probably on me to find that out myself.
DO and Vultr are colocated at very large sites. For many locations they sit right next to carrier hotels. I am a bit skeptical of whether the real world difference is that huge, especially if I sprawl my endpoints across both these providers and many locations.
The traffic prices for each of these are very different.
We use VPS providers for compute and S3 for backup and large blob storage. The only other Amazon product we use is Route 53, which is a decent DNS-as-a-service and is very reliable.
What I'd really like to see is Compose.io's full product lineup on Digital Ocean already. As it is, I'm considering seeing how well Azure Tables, and Azure SQL work from DO SFO2 to Azure West US...
On the other hand there are a lot of inconsistencies and odd limitations on the services.
There's really only one thing I can think of that I love is resource groups. It's nice to be able to group together everything in a deployment, for quick access and administration.
Another classic example is Directory Service's security group can only be found if you go to VPC / EC2 console page to search. The SG isn't even mentioned in the Directory Service console. I bought reserved instances for RDS and EC2, one of them accepts custom tagging, so you could name the purchase, but the other doesn't.
If you use CloudFormation you want to make sure the resources are managed by the CFn. Your manual change could break the next CFn update. In some cases, rollback is not possible and will require AWS Support's intervention. Another example, say you have a cluster of 3 Cassandra nodes, you would build 3 EC2 from a single CFn stack. Well, time for server rotation (by that I meant taking the instance out and replace with a new one), one node at a time. Too bad you really can't out of the box. You have to build another stack with 3 instances, and then start configuring one node and swapping one node at a time plus changing the underlying dynamic inventory we got out of Ansible.
How onerous, a security mechanism you actively turned on actually requiring you to do something. My car is just as bad - I drove it at a wall, and the front bit crumpled up!
And if you're getting your phone out twice a day, then perhaps your work hours are a bit crazy (12 hours+) and the problem is less to do with AWS than your workplace?
And as someone who works at home, I work whenever the hell I want, get off my back and go troll elsewhere.
Call me crazy, but that seems like a totally different feature than what AWS two-factor auth is doing. AWS uses 2FA to validate users, while Google uses it to validate devices, and you're making it seem like Amazon is too dumb to use local storage for something.
(Also, I don't think the parent commenter was comparing Google to a car crash. The analogy was that it seems weird to get mad about features performing their stated actions.)
If you're that sensitive about lighthearted hyperbole, then perhaps you shouldn't use it in the first place.
My main nemesis now is the ElasticSearch dashboard. Just opening it gives me a 50-50 chance of my browser tab crashing.
I pull out my phone at least 6 times a day for that just to switch between accounts.
Welcome to the wonderful world that agile/lean methodology has given us. They've delievered shit just to say they delievered something. They won't actually give you what you need until they get enough backing for it.
A though experiment: do you blame democracy, because Trump is now a nominee? If you have lousy people, choice of process does not really matter.
You're right.
This never happened before the agile manifesto was published.
The major downsides that I've noticed are: 1) Documentation is lacking (but improving!) 2) Issues that aren't affecting a lot of customers can sometimes take a long time to resolve. 3) Many services (including App Engine Flexible Environments) are still in beta, meaning no SLA, and they recommend against using them in prod. Unless you have a big paid support contract you'll have no clue how soon (if ever) things will reach GA.
https://cloud.google.com/compute/docs/regions-zones/regions-...
For example at a previous job we were "early" adopters of Amazon Redshift and it gave us no end of troubles. That should definitely be labeled "beta" until they sort those issues out.
1. Networking -> I completely agree with the OP here, The networking on AWS needs to be better. I don't want the strongest machine just to have a better transfer rate. It makes complete sense to have a micro machine for some services, but if those services are accessed or access other HTTP/s services, it will be unnecessarily slow
2. Pricing -> I have the privilege of working at a company that can afford to pay 3 year in advance. Even if you do that though Amazon will keep you on the machine types you purchased and not on the newest parallel machine types. In which case if you paid 3 year in advance you are often "stuck" on previous generation instance types.
3. Someone mentioned CloudFormation here. I completely ditched it in favor of terraform, CloudFormation looks like a tool from the 90s after using Terraform (which itself isn't clear of flaws as well)
4. In terms of VPC and networking I completely disagree with the OP, the networking and security settings on Amazon are great (if you understand them). You can define instances that have absolutely no access to/from the outside world. If you build a secure service, some of your services living in a "sandbox" makes total sense
5. App Engine -> The amount of complaints I heard about this service over the last couple of years are just insane. I heard from multiple users of the platform that it sucks. While I have absolutely zero experience with it myself, I tend to listen to people that suffer from it daily.
GAE is a massive lock-in with outdated software in a world where there's so many better, cheaper, non-lock-in alternatives, even within Google itself.
Within Google, either of the GCEs (https://cloud.google.com/compute/ https://cloud.google.com/container-engine/) depending on your architecture.
I'm not saying "Don't use GAE because it's expensive/google/whatever". I'm saying "Don't use GAE or anything like it".
Cloud Foundry gives you the "just push" experience for both of these (as well as vSphere, OpenStack, GCP and more to come). It's open source, with the IP owned by an independent foundation.
Disclosure: I work for Pivotal, we donate the majority of engineering on Cloud Foundry.
As I see it, at the time App Engine was introduced, there were a lot more comparable PaaS offerings (maybe not with the same languages support, but the style of PaaS that GAE is was more common.) Competition from GAE and Heroku (which originally was focused on a similar PaaS offering) seems to have driven most of the other alternatives to pivot to something else or fail entirely.
Right now, there's very few close substitutes; there's lots of "alternatives" that require more infrastructure work (e.g., IaaS and similar offerings like GCE, GKE, EC2, etc.) or are off in the other direction ("severless" function hosting like Lambda or GCF.) And in some cases these might be better options than an GAE-style PaaS, but they aren't really direct substitutes, and they leave a space better served by GAE.
Since App Engine comes with a proprietary data store, ORM and task queue built in, skipping the Appscale compatibility thing and just porting a large application to run on another platform would be a herculean effort. The data store models and everywhere in the application code that they're queried would need to be re-written for a different ORM, all of the data would have to be migrated to a different store, and all the background tasks and methods that trigger them would need to be rebuilt. This would be thousands and thousands of lines of hand-edited code changes, before even thinking about how much time would need to be spent verifying all of the changes.
At least with Heroku the magic is mostly around deploying and scaling the various pieces, rather than providing a lot of application level libraries that lock you into the platform. Porting to another platform likely would require a bunch of configuration management and some deployment scripts, but fairly limited changes to the application code itself.
Generally we find that any buildpack written for Heroku will work with Cloud Foundry without modification. With a little extra engineering you can create a buildpack that will run in a fully disconnected environment.
Disclosure: I work on the Cloud Foundry Buildpacks team on behalf of Pivotal.
AppScale will allow you to move your application (unchanged) and still reap the benefits of the App Engine model, autoscaling and all. I still believe App Engine benefits trumps any other PaaS out there.
There is an opensource reimplementation of App Engine's interface. That said, I have no clue about its completeness or buglessness.
The GAE team has been (admittedly slowly) pushing the services that were GAE only into being fully-consumable "Cloud Platform" services (e.g., Cloud Datastore is the same Datastore that's been "part of" App Engine forever, but now accessible from anywhere).
It's your choice to use Datastore, and I respect that you would chose to avoid it, but the basic PaaS idea of "here's some code, run it for me as a web endpoint" is still compelling even today.
Disclosure: I work on Google Cloud.
But to answer your question, I don't know anything about that Amazon service but I would feel the same way about similarly-designed competing product. IMHO, no informed person would ever pick proprietary, lock-in-heavy PAAS solutions.
Is your concern the lock-in from the code standpoint, or operational "lock-in" that once you've got all this set up for you, you'd have to go replicate it to get out? If the latter, aren't you saying you're going to do that on Day 1 in the "avoid PaaS" case?
This sounds like your personal bias and is clearly not true. Business doesn't work like this and you would be surprised if you looked at just how much proprietary software runs the world.
Lock-in isn't a thing to fear, it's a natural spectrum of using any service. The more specialized that service, the more work will be involved if you need to move away.
The real question is if the risk is worth it. Do you really need to switch? Are you worried that Google or AWS will somehow disappear before your business does? If not, what is the big deal?
You're conflating two things though: Lock-in as a whole, and unnecessary lock-in like GAE's. Yes, lock-in is a choice to make and not always an incorrect one. But I can safely say no informed business should/would pick Adobe Flash over HTML today for, say, internal web apps. Adobe Flash is fairly old and legacy technology, deprecated in all but name for that scenario and includes a fairly significant amount of lock in due to its proprietary nature requiring you to rework your entire application were you to change to an alternatire. Sounds familiar?
You weigh the risks and see if the potential work (that might not ever need to be done) is worth the benefits offered today by that service. That's what an informed business does.
WRT to VPC and security groups, I wonder if the OP might be talking about some of the surprising and frustrating limitations with them. E.g. VPCs are limited to fairly small subnet blocks, within which you have virtually no control. You can specify simple static routes - but only for subnets outside your VPC block. If you want to run a routing protocol in your VPC (e.g. you want your docker containers to be reachable via BGP), you get stuck pretty quickly - if the BGP subnet is within the VPC block, you can't route to them within the VPC, as the router won't permit more specific routes within the BGP block. If it's not within the VPC block, you won't be able to get to it via VPNs, peering connections, etc. So you're stuck either way.
Security groups are also limited - the number of rules you can have is pretty small, and they are a weird mix of stateless and stateful - they are stateful if the src/dest is an individual IP or group, but stateless for 0.0.0.0/0. And As you can't match TCP states like related or established, properly securing services when you have rules including 0.0.0.0/0 can be surprisingly difficult to achieve.
These are but a few examples of the limitations of the ways in which networking and security are implemented on AWS - there are plenty plenty more.
You can have multiple security groups now, though. Previously you were limited to just the one.
And if you don't want any external communication at all (barring your VPN), just remove the 'IGW' from your VPC (or make a new VPC without one). Or modify the VPC's routing tables. Or don't assign any public IPs to anything in the VPC. Or probably a few other methods :)
The good thing is that you can easily add and remove devices / users to your network. We require an agent to be deployed, which is SoftEther's VPN client (free, open source, known VPN client).
I see this criticism a lot for various things... and the things in question never look like the tools and/or websites I saw in the 90s. Instead, they look like things we'd dream of having in the 90s.
- On GCE the security groups are created automatically and managed by the "role" of instances.
On AWS there is no link between instances and security groups. I currently have to emulate the working of GCE over AWS with extremely complex ansible scripts just so that an instance called "repository" can actually be assigned to a group called "repository".
- ELB (load balancer on AWS) cannot have a fixed IP (we've been waiting for that for YEARS). An ELB can only be accessed over a DNS name, and the underlying IP can change every 60 second. (Have fun with applications which are caching DNS).
By comparison, GCE has had load balancers with fixed IPs forever.
- AWS is a region nightmare. Many resources only exist in one region and cannot be accessed or even acknowledged to exist from another region.
e.g. an AMI only exists in a single region. It's not possible to create a host with it in another region. It's not even possible to take that existing AMI and push it to another region. (As far as a region is concerned, other regions don't exist).
All services on AWS are acutely region centric. I am currently expanding my infra to multiple regions and there are many obstacles to overcome. The networking and interconnections is definitely one of them.
By comparison Google is not region centric like that. It's a lot easier to manage at planet scale.
---
Just three major pain points on the top of my head.
https://status.cloud.google.com/incident/compute/16015 https://status.cloud.google.com/incident/compute/16007 https://cloudharmony.com/status-for-google
Not enough customers to complain / notice.
Gartner report in the last week or two pointed out that Amazon is way in the lead with more cloud capacity than all the other providers combined.
The top two clouds are AWS and Azure. When either of those have major incidents you hear about it all over twitter etc. because that's a large user base that gets impacted. Google is in third place (according to them), but dramatically far behind. They do praise Google's big data tools, but point out even then people use AWS for most stuff and Google just for their big data bits.
Disclosure: I work on GCP
Keeping hobby projects under $10-20/month is sometimes hard...
http://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/CHAP_L...
Along similar lines, I really find Azure SQL tempting to use as well, may test the latency from DO's SFO location to Azure's West US.
Disclosure: I work on Google Cloud, so I'll probably come back to you and hold you to this statement...
Then, you either assign a public IP to each machine, or build yourself a solution...luke a NAT gateway. Google does not offer a service like this.
The machine that are part of the same subnet cannot talk to each other directly.. they have to go via the google gateway. This reduces the posibility of implementing High Avalability solutions that are based on broadcast or Virtual IP (MS SQL, Windows failover, vrrp...etc).
Labels that you apply to VM's or firewall rules are free text without the possibility to select from a labels list/cloud or type a new one...accidents can happen.
The MySQL solution is only reachable via INTERNET!!!. They seems to have a beta for using a local socket if you install a sql proxy on each server where you need sql db access...but not in production.
No support for MS SQL.
I'll be testing some latency between DO and Azure just so I can use some of Azure's services while keeping my applications on DO. Also considering Database Labs (databaselabs.io) for PostgreSQL on DO. Though I do like some of Azure's other offerings... Azure Tables are pretty nice if your needs fit the paradigm, even wrote some node.js wrappers to make access a lot easier a couple years ago, if I go with Azure will likely update them.
The lack of a hosted VPN solution is also weird.
We ended up setting up OpenVPN, which shouldn't be necessary. SSH is great and all, but when you add Kubernetes and other sensitive things like dashboards, you really want to virtualize everything and hide every internal service behind a proper, solid VPN gateway.
That's feasible but only things that do certificate exchange (https etc is easy enough). For what it's worth, part of the reason for Cloud Shell (the little shell within the web console) is so that you can interact with your infrastructure from inside the GCP network. That doesn't solve all your needs (like dashboards or deploying code from your laptop) but is often quite handy.
Generally individual users connecting to a VPN are called road warriors, even if they're not on the road and always connect from the same place.
The feature has been missing for years, it doesn't have any workaround, it's not being worked on anytime soon.
AWS vs GCE: Pick your limitations ;)
Any company who started before had to implement their own NAT gateway (we did). And with the shitty networking (150 Mbps top on large instances), that was the bottleneck for the whole subnet/VPC.
GCE has a lot better networking and 1Gbps on their instances which may them way less sucky than AWS for anything routing related.
Anyway. I migrated part of our cloud infra (one hundred instances, few subnets) to the new NAT gateway recently. It's working well so far but a few notes come to mind.
You gotta use the latest terraform (or cloudformation, tip: TERRAFORM IS BETTER). Eventually, play with the Release Candidates if you started early.
Ansible doesn't support NAT Gateway at all. (And probably that puppet/chef/salt do not either.) I doubt they will before next year.
Just my 2 cents about NAT Gateway. Let's not be too quick to judge google when AWS just released the feature a few months ago and it's only being available in tools about now.
Potentially a deal-breaker for us. We had a LOT of issues accessing various Google services after switching from Google Apps to Office 365, which is cheaper and includes the office suite which the rest of the world use.
That's not exactly right. You need a Google account, but it does not need to have gmail attached to it. Yes, lots of people don't know such a thing exists...but it's easy to set up.
On the "standard" Google account creation page is a link that says something like "I prefer to use my current email address."
As long as you have an email address that Google can validate, it's fine to use that instead of a gmail address.
But I have $15,000 in AWS credits from Stripe's Atlas program, and that was enough to make me reluctantly migrate from GCP to AWS. It means I can do things the "right way" instead of the "cheap way".
It's unfortunate Google couldn't offer something similar.
Could be the best of both worlds there... gitlab's work for CI/CD is interesting as well.
AWS Console is not perfect, but IMO, it's much more clear where I need to go to find things, even without reading a single piece of documentation. And I've rarely had issues with AWS documentation not being clear and correct.
Maybe GCE docs are better, though.
In the end I went with Digital Ocean and had a server up and running in barely a few minutes, much quicker and easier.
I expect more effort to go into this from a technical perspective. Give me some workload / price differentials. Examples of architectures that are simpler on GCP then AWS.
My app relies on a hybrid of Firebase and AWS for the heavy assets. Still trying to justify the move to Google Firebase for the full stack. I'm the only dev so it's time and energy I can't spend on frontend iteration helping real users.
Disclosure: I work on Google Cloud.
This article is addressing that most of the GCE services are simpler and a lot easier to use and manage than the AWS counterpartS. (That final S being important, having many small and unpolished pieces just contribute to the complexity.)
It goes hand in hand with what I had already written about the platform. Expect GCE to be 20-50% cheaper and 1-3x faster:
https://thehftguy.wordpress.com/2016/06/15/gce-vs-aws-in-201...
https://thehftguy.wordpress.com/2016/06/22/a-simple-cost-com...
When you see that GCE has less market revenues, you should keep in mind that GCE could be half the price of AWS in average (yep, no kidding). That means the actual gap in customers is way smaller than it appears at first sight ;)
With that being said. Both platforms are relatively new and spinning services like mad, they both have important features missing (albeit not the same ones.)
Isn't that what AWS was also built for initially? For internal use by developers? I see this statement from time to time in articles like this.
GOOG/Alphabet as a company is going to be around, but will this line of business be with this set of offerings?
Google is very committed to GCP and you see it in their work. Much like Amazon, they have a huge sum of computing capacity they aren't always using. It's an obvious business play.
But what's more, it's become the basis of differentiating their mobile platform. Google is trying to develop and sell-for-cheap the tools to make really good mobile apps (e.g., Firebase) to try and shore up the app ecosystem of their platform.
Given that a lot of the things they offer are 3rd gen variants of their core tech for developing things internally, it seems fairly reasonable to assume they're in this for the long haul.
Digital Ocean and AWS are dedicated to continuing this line of business and keeping those customers addicted. Google is conducting another large and expensive moon shot which fine since it's part of their DNA.
I'm not sure you can say it is an expensive moonshot at this point.
I'm not using them, myself, but I wouldn't consider that aspect of the risk substantial given the fact that I could erect identical infrastructure on AWS in the worst case scenario. And I have had to plan exit scenarios from AWS because ultimately they're a single vendor with arbitrary control over their platform, as well.
As to ease your worry here a little, any launched cloud products have at a minimum a 1-year deprecation policy[0]. Anecdotal: AppEngine deprecated their Master/Slave datastore[1] in April 2012, and it was actually shutdown on July 6, 2015. So that's 3+ years to move to newer tech.
[0] https://cloud.google.com/terms/ (section 7.2)
[1] http://googleappengine.blogspot.com/2012/04/masterslave-data...
You point out examples of longer deprecation, but that doesn't matter one iota to people evaluating the cloud. Track record isn't legally binding, unlike the contractual 12 months.
So you're advocating we should trust the book vendor?
That's why GCE should be as important to Google as AWS is to Amazon while arguably both have a different "core business".
3. They gave you more than 3 years :)
[1]: http://www.networkworld.com/article/3029164/cloud-computing/...
[2]: http://www.recode.net/2016/4/28/11586526/aws-cloud-revenue-g...
AWS is there through first-to-market leadership.
Microsoft is there because they've got a strong sense of what every SME and Enterprise customer in the world wants.
Google are there because they have deep expertise in building these kinds of systems and it would be utterly bonkers not to try and diversify into this field.
There's so much Enterprise Buzzword Bingo that Amazon and Google try to confuse you on as to why there service is magically better, but IMO the biggest selling points for startups are price/bandwidth/performance. And the prices and machine configurations change all the time.
AWS is built by developers for developers as well.
Most of the complaints here about AWS aren't actually about AWS not being "for developers", but about AWS requiring a certain learning curve.
It's a perfect trade-off between power and flexibility vs agility.
That's not to say that these things don't solve real problems, but they are the problems of big organizations. Smaller teams building pure-cloud products don't have the same problems, and may not even have people with Big Corp experience.
AWS is the Java of cloud platforms.
It sounds like a lot of his reason for switching is simplicity. From this article, it sounds like AWS has a lot more options and control, and GAE just does what it wants or thinks is best. I suppose to some, the latter may be appealing, but at least the author is honest with a "your mileage may vary". Willing to bet this simplicity costs a lot of options other companies may wish to use.
I'm convinced Google's infrastructure play will ultimately win the hearts and minds of developers because they understand this.
For most, it depends on the risk of an issue (mostly measured ad-hoc by how many times you hear of others having an issue) vs the potential pain one would feel if an issue arises.
For a business's core infrastructure, I've heard enough so that the risk is high enough, when taken together with the potential loss of everything, that I make it a point to spend time and effort to understand as much as I can about what I rely on.
For walking down the sidewalk, no one is suggesting that you understand the fundamentals of concrete curing. The risk of something happening to you must be practically zero, and the potential harm if something does also seems quite low.
Of course, there's always going to be a success story in the midst of the crowd where someone flew by the seat of their pants and everything worked out.
I happen to think that's a minority, and I also don't feel comfortable gambling with my business like that. But to each their own.
If GCP can bring 80% of the benefits of AWS at a reasonable price for shops without the requirement of deep AWS & devops expertise, isn't that a really excellent place to occupy?
People here often seem to have an implicit assumption that AWS's complexity is either inevitable or inherently of value. I'd point to Azure as another alternative rejection of that viewpoint that serves a wider audience of technical skill levels than GCP. And honestly, I really like working with Azure.
I've built a few businesses on AWS now, and I know and trust it. But it has a ton of infuriating features and many many things they've promised have been delayed years. I'm not opposed to a new CSP at all. Why would anyone be?
AWS IAM policies are a good example - yes they are complicated and hard to navigate, but when you need to do complicated grants of access to your resources, then you'll appreciate the complication.
If you are part of a larger business with an established system administration function, then the flexibility of AWS is probably appealing, because you can transpose your existing set-ups and practices on to it.
For smaller organizations, particularly new "cloud-native" ones that may have never owned a server and might not have big administration teams, AWS looks like a bad fit.
Learning AWS is now more complicated and a bigger time investment than any technology that I've encountered for years. We operations people are used to doing a lot of learning up-front for gnarly and sometimes old tech, but it's not how developers engage with new platforms today.
If GCP can deliver the 80% without the huge time investment that mastering AWS now requires then over time, it's going to win the green-field projects.
A couple of things that baffle me about AWS though -
* You can purchase reserved instances from anywhere in the world in a few seconds, but WHY when I want to on sell an unused reserved instance in the marketplace, do I HAVE to have a US bank account for Amazon to pay my revenues into?? Why can they just not credit my existing account which I pay thousands per month to them?
* We still battle with occasional high latency spikes for EC2 instances talking to RDS instances within the same VPC! Continuous "MySQL server has gone away" log messages. Ironically, we have absolutely no trouble with Non-VPC EC2 instances that use classic bridging to talk to the same RDS server within the VPC!
We thought to simplify our architecture by putting everything in one VPC for better security and ease of maintenance, but had to go back to a cobbled half and half solution to maintain performance. Hence why we now have a few reserved 'VPC only' instances which we cannot recoup costs by onselling in the marketplace... :(
Some thoughts:
## More focus on core products
- Still no IPv6, it's 2016.
- EC2/VPC (compute) didn't have a simple NAT-instance until a few months ago.
- S3 (storage) still has nothing like Nearline (GCE) and Glacier is so confusing and complicated that few bother using it.
- ELB (Load Balancer) still has scaling problems (with bursts).
- Still only a single public key per instance, why?
- Can only use ACM certificates for CloudFront if they are in us-east-1 (despite ACM being available in other geographies and CloudFront is global). Nothing major but weird.
- S3 has like a handful of different ways of setting permissions (per file, per bucket, iam, and some weird old xml-format). Why not simplify this?
## Interfaces are terrible
- Basically the entire console UI is low quality.
- Same with APIs, they could use a lot more polish.
- This is the API call to create a CloudFront (CDN) distribution: http://pastie.org/10931494 (no offence, but it looks like a group of schizophrenic monkeys designed that API).
## Think less about press-releases
- Apparently AWS force everyone to write the press-release for every product before designing[1].
- While it's good to keep the end-goal in mind I think many important details are lost.
- The press-release-thinking focus too much on features and too little on quality and (even more important) refinements to existing things.
## Release fewer features
- Every year on ReInvent [AWS conference] they tout how many hundreds of new features they've released[2]. I wish they'd release far fewer features and make them great instead (+ refine existing stuff).
[1] http://www.allthingsdistributed.com/2006/11/working_backward...
[2] https://s3-eu-west-1.amazonaws.com/vpblogimg/2015/04/Aws-Sum...
>- Basically the entire console UI is low quality.
I wouldn't bet on them fixing this. Amazon in general are not big on design and in particular they aren't about making things look good. I think in general they see it as a waste of time, something that can always be done later after they have beaten their competitors. Obviously there's a threshold for this sort of neglect but they seem pretty expert at riding it.
Frankly what you said about it is polite. The AWS web interface is horrendously ugly and just barely functions well enough to be used. It's a testament to how little they care about good design but then again their consumer-facing website is no peach either.
Managed NAT is new, but the nat instances that previously existing could be spun up without any configuration outside of disabling src/dst check and configuring a route to point to them, and they've existed for many years. http://docs.aws.amazon.com/AmazonVPC/latest/UserGuide/VPC_NA...
>- S3 (storage) still has nothing like Nearline (GCE) and Glacier is so confusing and complicated that few bother using it.
S3 Infrequently Accessed? https://aws.amazon.com/s3/storage-classes/
Still, I think Nearline easily beats Glacier on simplicity and clarity. With Nearline it's super-obvious how long time retrievals take and what the costs are.
--------------
## Nearline
Q: How much does it cost?
1 cent per GB & month in storage. Plus 1 cent / gb of transfer (retrieval).
Q: How fast can I retrieve data?
Access times sub 1 second.
SOURCE: https://cloud.google.com/storage-nearline/
--------------
## AWS Glacier
Q: How will I be charged when retrieving large amounts of data from Amazon Glacier?
You can retrieve up to 5% of your average monthly storage, pro-rated daily, for free each month. For example, if on a given day you have 75 TB of data stored in Amazon Glacier, you can retrieve up to 128 GB of data for free that day (75 terabytes x 5% / 30 days = 128 GB, assuming it is a 30 day month). In this example, 128 GB is your daily free retrieval allowance. Each month, you are only charged a Retrieval Fee if you exceed your daily retrieval allowance. Let's now look at how this Retrieval Fee - which is based on your monthly peak billable retrieval rate - is calculated.
Let’s assume you are storing 75 TB of data and you would like to retrieve 140 GB. The amount you pay is determined by how fast you retrieve the data. For example, you can request all the data at once and pay $21.60, or retrieve it evenly over eight hours, and pay $10.80. If you further spread your retrievals evenly over 28 hours, your retrievals would be free because you would be retrieving less than 128 GB per day. You can lower your billable retrieval rate and therefore reduce or eliminate your retrieval fees by spreading out your retrievals over longer periods of time.
Below we review how to calculate Retrieval Fees if you stored 75 TB and retrieved 140 GB in 4 hours, 8 hours, and 28 hours respectively.
First we calculate your peak retrieval rate. Your peak hourly retrieval rate each month is equal to the greatest amount of data you retrieve in any hour over the course of the month. If you initiate several retrieval jobs in the same hour, these are added together to determine your hourly retrieval rate. We always assume that a retrieval job completes in 4 hours for the purpose of calculating your peak retrieval rate. In this case your peak rate is 140 GB/4 hours, which equals 35 GB per hour.
Then we calculate your peak billable retrieval rate by subtracting the amount of data you get for free from your peak rate. To calculate your free data we look at your daily allowance and divide it by the number of hours in the day that you retrieved data. So in this case your free data is 128 GB /4 hours or 32 GB free per hour. This makes your billable retrieval rate 35 GB/hour – 32 GB per hour which equals 3 GB per hour.
To calculate how much you pay for the month we multiply your peak billable retrieval rate (3 GB per hour) by the retrieval fee ($0.01/GB) by the number of hours in a month (720). So in this instance you pay 3 GB/Hour * $0.01 * 720 hours, which equals $21.60 to retrieve 140 GB in 3-5 hours.
First we calculate your peak retrieval rate. Again, for the purpose of calculating your retrieval fee, we always assume retrievals complete in 4 hours. If you request 70GB of data at a time with an interval of at least 4 hours, your peak retrieval rate would then be 70GB / 4 hours = 17.50 GB per hour. (This assumes that your retrievals start and end in the same day).
Then we calculate your peak billable retrieval rate by subtracting the amount of data you get for free from your peak rate. To calculate your free data we look at your daily allowance and divide it by the number of hours in the day that you retrieved data. So in this case your free data is 128 GB /8 hours or 16 GB free per hour. This makes your billable retrieval rate 17.5 GB/hour – 16 GB per hour which equals 1.5 GB/hour. To calculate how much you pay for the month we multiply your peak hourly billable retrieval rate (1.5 GB/hour) by the retrieval fee ($0.01/GB) by the number of hours in a month (720). So in this instance you pay 1.5 GB/hour x $0.01 x 720 hours, which equals $10.80 to retrieve 40 GB.
If you spread your retrievals over 28 hours, you would no longer exceed your daily free retrieval allowance and would therefore not be charged a Retrieval Fee.
Q: How is my storage charge calculated?
The volume of storage billed in a month is based on the average storage used throughout the month, measured in gigabyte-months (GB-Months). The size of each of your archives is calculated as the amount of data you upload plus an additional 32 kilobytes of data for indexing and metadata (e.g. your archive description). This extra data is necessary to identify and retrieve your archive. Here is an example of how to calculate your storage costs using US East (Northern Virginia) Region pricing:
Your storage is measured in “TimedStorage-ByteHrs,” which are added up at the end of the month to generate your monthly charges.
Q: How long does it take for jobs to complete?
Most jobs will take between 3 to 5 hours to complete.
I'd like to add one piece I always forget: with S3-IA, you pay for a minimum of 128 KB, which for apps with tons of small objects really adds up!
Disclosure: I work on Google Cloud.
* http://siliconangle.com/blog/2016/08/05/aws-microsoft-azure-...
day to day i work with AWS, Azure, and Google quite a bit -- from what, in terms of investment, i would say gartner is accurate.
i'm not sure why digital ocean wasn't taken as serious.
1) Didn't mean to hand-wave. For the last few months GCP was the new shiny toy for me. I got over excited and plan to do deeper post, battling 2 or more products (i.e Kinesis vs. Pubsub)
2) Not anti-AWS at all. I still find it wonderful: https://lugassy.net/search?q=aws. what really flipped me was this: https://twitter.com/mluggy/status/727764607176159232 and other small anonyances
3) Downsides to GCP. Certainly! added to the article
4) For developers, by developers. I realize this was tacky. Let me try again with 4 examples: a) Trace/breakpoints in production b) Connecting through virtual socket files instead of tcp c) writing and collecting logs centrally with console.log and specifically d) DATAFLOW which is architecturally brilliant (AWS add "Streams" to every new product where GCP simply treat every product as source and/or sink)
5) Yes I favor GCP for simplicity and dev-friendliness. I don't want to train people for AWS devops/gotchas. I don't need sophisticated IAM and VPC to feel smart. GCP ui/quickbar is nicer. For example networking is all on one page. Most products are properly named (CDN instead of CloudFront, DNS instead of Route53, Pub/Sub instead of Kinesis)
6) GAE. I was exposed to some horror stories about its and datastore early days. History aside, I use it today through Flexible VMs which is really docker under the hood + ability to SSH + easy deployment/versioning + cool tracer/debugger. My code is Node.js and haven't coded any GAE specific necessity
Again loved the comments. keep it going!
Uhh...no. Why would anyone assume some random Gmail account is secure by default?
Disclosure: I work on Google Cloud, so if you sign up I indirectly ... take your money?
Google Compute Engine instances will always be more expensive than DO/Linode because GCP offers so much more. GCE instances will only be cheaper than DO droplets if you will shut down the instances when you don't need them as you don't get charged for them when they are off unlike DO.
Disclosure: I work on Compute Engine and Cloud and want your business.
I would move to them in heartbeat if they had anything like postgres RDS ... But Google seems to love it's own version of MySQL.
Anybody know if this is on the roadmap ?
Between everything else, database is the single most critical aspect - and there's a huge value for a high availability system.
Rough patches are:
- Live migration is sometimes not seamless.
- Pub/sub is missing some core features.
This point is oddly weak in comparison to the others.
> Amazon has one of the most confusing IAM. While it is nice to set up a role to only allows usage for a particular resource from a specific device and times of day, you end up spending most of your time debugging policies.
Haven't used GCP, but I didn't mind IAM. Missed it when I waas trying to figure out Azure stuff.
> We moved because we wanted to work on infrastructure that runs YouTube, Gmail and Google Analytics. We moved because Google is fair, much more tech-savvy and launch products that just works.
ugh. Perfectly fine article now just reeks of google "fanboy-ism".
> Google lets you set up simple Firewall rules. Amazon gives you VPC, security groups, network access control lists and a big, fat headache.
AWS also gives you firewall rules, in addition to the other options. If you don't get why you would need them, you don't really have to use them.
Agreed - it was senseless. SNS/SQS aren't particularly complex and they scale like crazy. Kinesis is something else entirely (N-hour record retention w/arbitrary consumers and offsets) and doesn't even belong in the comparison.
> ugh. Perfectly fine article now just reeks of google "fanboy-ism".
Do Google's core products actually run on GCP these days? I was under the impression they do not.
Not really. They are influenced by, but the same, code as runs Google. (for instance, Kubernetes is based on Borg)
Also, another random oddity: somewhere in one of its built in libraries there was a function for validating email addresses. It was returning true on a bunch of very obviously bad address formats, so I looked up the source and found that all the function did was verify its argument wasn't the empty string.
I imagine GAE is better now, but my first experience with it left a bad taste in my mouth.
My favourite comment of the whole thread.