The Cloud Is Raining Cash on Amazon, Google, and Microsoft
bloomberg.com
bloomberg.com
Cloud computing means the entire DC is programmable, and I use it as such.
Moreover the unprecedented level of automation means I can spend a lot more time on creating customer value rather than faffing around with admin. The shift I've seen in the last three decades* has been phenomenal. Teams aren't smaller but they are vastly more productive.
* yes I have been in tech that long :~
In 100% of my experience to date, "cloud lock-in" is a myth trotted out by server huggers and hardware salesmen. Some SaaS providers may be data prisons, sure, but that's a different conversation.
If the economics of establishing and operating off-cloud resources ever made sense for us, we'd go for it, but it looks increasingly unlikely.
Edit: I should note that we do, where appropriate use 'cloud' services including Rackspace, AWS and Azure where appropriate. Azure has had significant performance issues and has a lot of provable downtime especially due to internal network routing and DNS problems that they fail to acknowledge and we've found their support to be a joke if you know what you need / are doing, even their own O365 service has weekly outages that can take several minutes to resolve. Rackspace's support has been good but they do have a lot of small outages again often network related. AWS' has been alright be very costly unless you're doing either very small deployments or at the other end of the scale massive, horizontally scalable deployments, however their storage performance is woeful. For our mission critical or high performance deployments using our internally hosted platform is significantly fast and almost always cheaper. Our uptime across the platform is fantastic and it generally 'just works' while we watch our cloud hosted services suffer from inconsistent performance and service 'blips'.
(I'm genuinely curious but I feel like there's a missing upfront cost that's not being included here)
The advantage of a lot of these is that you don't need any of that. You can have just a dev team, nothing else, and be up and running at any scale. At that point cost comparisons may come out in favor of a cloud solution (or at any rate are so close that the tradeoffs become acceptable, hence their popularity)
Has anyone done this in practice at significant scale? Every shop I know of with a big cloud deployment has armies of devops people managing the deployment, and further reserve armies on pager-duty standby in case something blows up. They're certainly doing different things than if you had an in-house datacenter (less hardware maintenance, more cloud orchestration), but I'm not convinced the sysadmin/devops headcount has actually gone down.
TwitPic relied on it and operated as basically a one or few person shop the entire time. At the peak, before Twitter came up with their own solution, it was a relatively large service.
Instagram, Imgur, and Reddit all had/have (Instagram of course moved to FB) small teams operating at vast scale with the help of AWS.
Slack has probably benefited a lot from leaning on AWS for scaling purposes, given their rapid growth. I'd place a bet that they have managed to achieve their scale with a relatively small team managing their infrastructure.
Also Twitpic didn't reduce their sysops, they simply kept it low, AWS didn't do that, a server which serves much more complex stuff and many more people than twitpic can be run by one person, see stackoverflow.
There are a lot of people out there who don't realize what can be done with one or two dedicated servers and one or a handful of good developers without ever mucking around with cloud. Cloud can just add yet another point of failure to a small business if it wasn't worth it in the first place.
Sure, its not 5 9's, but its a far cry from "notoriously bad at staying up"
There's over 400 million Snapchats sent per day. I'd say that's pretty big scale, all done in the cloud.
Now you're talking about DevOps, which is short for Development Operations, which is a developer role, not an IT role. DevOps people automate your buildchain and that kind of stuff. No one is saying that using cloud services means you don't need DevOps, though there are some cloud services that will handle at least part of what is traditionally in the realm of DevOps for you.
If you're using an automatic container scaling solution, such as AWS' Elastic Beanstalk, you can still benefit from devops, but I'd argue you don't need it; all the difficulty resides in structuring your application to be able to handle that environment, a problem for your software devs. The devops burden is low, and the time to communicate what a developer needs to the devops is likely going to trump the time taken for the developer to just do it.
If you're looking to use a containerless solution using cloud resources (like AWS' Lambda, API Gateway, and Dynamo to create a CRUD app), you don't need devops. All your difficulty resides in reducing and/or handling state between functions (well, and any other shortcomings in the actual implementation of the service); again, a software problem.
Basically, Amazon, at least (what I have the most experience with) seems to be looking to remove the need for devops by creating standard workflows and mechanisms to bind services together with arbitrary code, and to be able to inherently scale out. The remaining devops burden is sufficiently small, and so tightly integrated with the nature of the software involved, that it's oftentimes more effective to just have the devs handle it. While sufficient amounts of code might turn that into a devops role, what I meant by scaling out is a particular app handling a given amount of load. In a classical datacenter environment, moving from an app on one box to an app that spans many is something both the software devs and devops have to concern themselves with, but in the cloud it's mostly just a dev consideration; if the app is written to handle multiple copies of itself, spinning up those extra copies should be close to if not actually trivial. That was all that I was saying; the move from one to many doesn't require devops any longer, as "how do I make sure all of these boxes are set up properly, get deployed onto, are kept in sync, load is shared between them, etc" are problems that cloud providers have provided tools to solve, and what they leave out doesn't require dedicated devops to address.
Anyone that says moving to the
cloud saves you money is likely
wrong.
Evidently you've never worked at a place where database disk space costs $31,000 for a terabyte. :)What kind of storage are we talking here? Even 1TB server SSDs don't cost nearly that much.
I don't see the allure of the costlier cloud other than the "nobody ever got fired for" factor so common in enterprise purchasing. Amazon is the only one that might have a stronger case for it based on its huge managed service stack, but much of that is not too terribly hard to duplicate with other tools and more a la carte services. There also really isn't a reason you can't use some of Amazon's stuff while also using more commodity options.
On a more principled level I'm starting to see huge proprietary cloud as a potential threat to the open Internet. It's not quite there yet but at some point I could see it, especially with the walled garden plays you see around IoT.
A lot of the difficulty also comes from over engineering and premature scalability obsession. You often just don't need all that. I swear over engineering is the bane of software and devops these days. We've gone from java factoryfactorysingletons to "how many distributed systems fads can I use in one stack?"
This implies that a $40k per month bill for AWS[2] could pay for three DevOps engineers and save approximately $10k per month in the vast majority of the US.
1 - http://www.indeed.com/q-Devops-Engineer-l-Boston,-MA-jobs.ht...
2 - derived from your statement that a staff of $20k per month would save "~50% off one's AWS bill"
IMHO, a reasonable estimation for fully loaded cost per employee (excluding facilities expendetures) is approximately 1.4 * ES, where "ES" is the employee salary.
The "three DevOps engineers and save $10k" estimation was based on working backward from the 92% of available jobs in Boston being less than half of the stated $20k per month cost. Assuming a Gaussian distribution where 0.5 * $20k per month represents the high end of two standard deviations (since Boston ranks quite highly in S/W salary nationally), most DevOps engineers will be paid roughly half of that as well.
This yielded an estimation of $6.5k per month per DevOps employee or $19.5k per month for three.
Since all of this was off-the-cuff, I figured it best to throw in a bit of "fudge factor" and present a $10k per month savings.
As always, YMMV and I could be completely wrong about all of this :-).
Edit: BigQuery, for example, allows you to rent 10,000 cores for 5 seconds at a time. This type of stuff is impossible to do with VMs at all.
PubSub is also a GLOBAL service. Not only are you protected from zone downtime, you are protected from regional downtime. Is there an equivalent to this level of service anywhere in the world?
I'm not too familiar with Kafka's fully managed service, but Kafka-on-VM is a whole other ball game. YOU manage the service. YOU guarantee delivery, not Kafka.
From an availability standpoint, I don't disagree with anything you mention, but the difference between the consistency models means that PubSub is solving a different set of problems than Kafka, thus my opinion that comparing them is problematic.
There are several ways to look at it, but I'd opine that a "mostly ordered" fully-managed truly-global service that's easy to unscramble on the receiving end is more "guaranteed" than something that is single-zone and relies on the health of underlying VMs that YOU have to manage.
edit: Kafka and PubSub have a lot of overlap, but they each have qualities the other one doesn't. I suppose you gotta choose which qualities are more important for you.
Also, just so we are on the same page. Kafka is a software product that can be run on hardware or VMs, not a managed service. Possibly, you are thinking of the Amazon Kinesis product which does offer a managed service with strict ordering.
No confusion on second point. My argument was that Kafka adds significant complexity and delivery risk because it's software that you must run on hardware/VMs, rather than a fully-managed service. You have to pay a whole lot of eng time to make Kafka truly "guaranteed delivery" because there's always risk of underlying hardware/VM/LB dying.
Pubsub guarantees delivery regardless of what happens with underlying infrastructure. In a sense, the bar has been raised dramatically.
Could you point to some of the documentation that describes more about its reliability model and SLA? I glanced through the documentation and couldn't find out any information about this.
It seems like a service that has this kind of global availability would have to make a trade off in latency for writes and potentially reads. If it's a multi-region service, then all writes need to block until they're acknowledge by at least a second region, right? It seems like that will add latency to every request and may not necessarily be a good thing. Similarly, at read time, latency could fluctuate depending on which region you query, and whether your usual region has the data yet. I'm just speculating though, not having read any more about the service. It does sound nice to have the choice to fall back to another region and take the latency hit, instead of an outage. On the other hand, regions are already highly available at existing cloud providers (with zones being a more common failure point).
Is PubSub mature? The FAQ suggests that you should authenticate that Google made the requests to your HTTPS endpoint by adding a secret parameter, rather than relying on any form of HTTP-level authentication.
> If you additionally would like to verify that the messages originated from Google Cloud Pub/Sub, you could configure your endpoint to only accept messages that are accompanied by a secret token argument, for example, https://myapp.mydomain.com/myhandler?token=application-secre....
This feels rather haphazard. If I'm exposing an HTTPS endpoint in my application that will trigger actual behavior upon the receipt of an HTTP request, then of course I "would like to verify that the messages originated from Google Cloud Pub/Sub", so that they're not coming from some random bot or deliberate attacker who happened to learn my URL.
- PubSub is used by Google internally to power everything from Android notifications to Hangouts messages. So it's certainly proven.
- A lot of your questions are answered in docs:
https://cloud.google.com/pubsub/
https://cloud.google.com/pubsub/docs
You can always reach out to me, and I can get you in touch with a PubSub SME.
In the "Delivery contract" section:
"For the most part Pub/Sub delivers each message once, and in the order in which it was published. However, once-only and in-order delivery are not guaranteed: it may happen that a message is delivered more than once, and out of order."
So it is at-least-once delivery as far as I see.
Point is, higher-level cloud-native services unlock very interesting use cases that are applicable for both small-scale startups and large companies, use cases that are impossible with just VMs.
What Amazon and kin have done is offer developers a new sexy way of over engineering. The AWS stack is the new Java OOP design patterns book. Yes, there is occasionally a time when an AbstractSingletonFactory is a good thing but I guarantee you most of those you see in the wild are not those times.
The real genius was to build a jungle gym for sophomore programmers to indulge their need to develop carpal tunnel syndrome where everything bills by the instance, hour, and transaction. If Sun had found a way to bill for every interface implemented and every use of the singleton pattern they would have been the ones buying Oracle.
Your argument can be summarized thus as this - do not give people incredible computing capacity at never-before-seen economic efficiency, because they will use it inefficiently. I'm afraid this argument gets made every time the world gets disrupted technologically (horse vs car anyone?).
Edit: I may argue that if "carbon footprint" is your prerogative, then economies of scale + power efficiency should tilt the scale towards cloud, no? AWS is certainly on the dirtier side, but there are other, greener clouds.
Since I do data analysis and machine learning (sometimes), a common one I see is people using "big-data analytics" stacks when they don't have anything remotely in the range of a big-data problem. Everyone really seems to want to have a big-data problem, but then it turns out they have like, single-digit gigabytes of data (or less). And they want Hadoop on top of some infrastructure to scale a fleet of AWS VMs, so they can plot some basic analytics charts on a few gigs of data? They would be better served by the revolutionary new big-data solution "R on a laptop". But somehow many people have convinced themselves they really need Hadoop on AWS.
Though I haven't used it yet, BigQuery does seem interesting in comparison, because it at least seems like it doesn't hurt you much. The Hadoop-on-VMs thing is objectionable rather than merely unnecessary, because you get this complex, over-architected system for what is not a complex problem. BigQuery at least seems like, at worst you end up with basically a cloud-hosted RDBMS with scaling features you don't need, which isn't the end of the world as long as the pricing works for you.
edit: Just to clarify, I'm not the person you were replying to, just someone who also has opinions on this. :)
I am not arguing that there are no great use cases for these systems. But I would be willing to bet that those are less than half the total load.
It's like big trucks. How many people who drive big trucks actually need big trucks? Personally I like my company's Prius of an infrastructure. :) And of course we've architected it so it can be a fleet or an armada of Priuses if need be, with maybe just a bit of work but if we get there I will be happy to have that problem.
But I think you might underestimate the amount of use-cases that do legitimately benefit from and desire a greater degree of reliability and automation. When one of my machines dies, I don't want to be notified, and I don't want to have to do anything about it. I want a new virtual machine to come online with the same software and pick up the slack. Similarly, as my system's traffic grows over time, I want to be able to gradually add machines to a fleet, to handle my scaling problem, or even instruct the system to do that for me.
Plenty of use-cases may not require this, but I'm not convinced that the majority of systems in the cloud do not. Every system benefits from reliability, and it's great to get it cheaply and in a hands-off way. In the cloud, I can build a system where my virtual machine runs on a virtual disk, and if there's a hardware failure, my VM gets restarted on another physical machine and keeps on trucking without my involvement. As an engineer and scientist, I can accomplish a lot more with a foundation like this. I can build systems that require nearly zero maintenance and management to keep running, even over long time scales.
I don't think I disagree with you that some people overengineer systems, but I think I disagree with you about how much effort it requires to achieve solid availability and a high level of automation. It's not a lot of effort or cost, and it's a huge advantage. Once I build a system I never want to touch it again.
A certain segment of users are adopting these technologies because they want to be prepared to scale. One of the advantages of "big data" products even for small use-cases is: all successful use-cases grow over time. If you plan for success and growth, then you may exceed the capabilities of a traditional technology. If you use a "big" technology from the beginning, then you can be confident that you'll be able to solve increases in demand by scaling up, rather than by rearchitecting. As these platforms mature and become easier to use, the scales begin to tip, and they no longer require more engineering time than the alternatives; a strong hosted platform actually requires less time in total, especially when you consider setup and maintenance. Many of these technologies do an excellent job "scaling down" for simple use-cases too. While they have been difficult to use, they're getting easier. For example, MapReduce-paradigm technologies are becoming fairly easy with Apache Hive, and fast with Spark. They're becoming easier to set up due to hosted variants like AWS's ElasticMapReduce or Google Cloud Dataproc, etc.
[0] http://www.businessinsider.com/snapchat-is-built-on-googles-...
Feel free to ping me if you want more info :)
Cloud Storage considers retrieving an object through its HTTP url is considered a class B XML request type which is priced at $0.01/10,000 ops. This is 1.3x-2.5x more expensive than cloudfront and s3 respectively. I think this is the only google cloud service which is more expensive than the AWS equivalent
We built our recommendation engine for Recent News (https://recent.io/) on Google App Engine in Python. There was some tricky engineering involved in making sure that it would work inside that particular environment, but it's paid off in terms of scalability. We're not worried about Google shutting down App Engine; in fact it's being continuously improved.
http://www.gearthblog.com/blog/archives/2015/01/google-maps-... The move should be seen as Google transitioning customers to already existing alternative products, especially Google My Maps (formerly Maps Engine Lite) which has come of age and now has most of the important features of Google Maps Engine
A better example would be if Google were to discontinue its paid Google Maps API?
The only companies that don't terminate projects when it it no longer serves a strategic business purpose to support them are companies that instead terminate the projects because they go out of business.
Google is, if anything, unusually good at providing warning and a migration path off a product when they decide to terminate it.