I think I need this translated back into tech-speak.
I think I need this translated back into tech-speak.
I’m not sure what else needs to be translated? Nothing, I think?
And then people also pick it up for non-persuasion, because it also sounds like a catchy name for an engineering approach we already had.
Of course it can still be used for persuasion for awhile, but will grow baggage over time, as efforts linked to the term don't play out that way.
The term didn't sound familiar to me (though the concept was), and the term might not have been familiar to some others.
People might not want to contradict an assertion because of language like "The term’s in wide use, talk to anyone involved with cloud-anything and they’ll be familiar with it. [...] I’m not sure what else needs to be translated? Nothing, I think?"
> > LinkedIn was having a hard time taking advantage of the cloud provider's software. Sources told CNBC that issues arose when LinkedIn attempted to lift and shift its existing software tools to Azure rather than refactor them to run on the cloud provider's ready made tools.
The only other terms I can see that are jargon are "cloud provider" and "refactor", and those are already technical (more or less) so don't need to be translated into technical language.
As for the other bit, I just meant that it's a widely-used term so one may continue to encounter it in these contexts. It truly is ubiquitous in discussion of and around "enterprise transformations" to the cloud, and among cloud practitioners more generally, so anyone connected to that space will know what it means. It's also kinda already a technical term, in that developer/devops and SRE sorts throw it around and do mean a specific thing by it, which doesn't need to be translated for other technical folks in that area.
The original person might've instead asked for an explanation in a way that didn't come across as criticizing the article.
But probably best not to insist that everyone should already know the term; just explain it.
It's usually done to avoid the engineering cost of making the services more cloud native. What tends to happen a lot is that after a considerable portion of the migration is completed, the cost of the lift-and-shift effort start to overtake the savings, and the projected costs, dwarf the future savings.
I suspect this is what happened with Linkedin.
I'm used to organizations moving out of the cloud when they realize that it's more expensive if you don't have very peaky load demands.
And it's difficult to make that as expensive as a cloud deployment.
Because you are paying someone else for them.
This is considered rational because those operators are presumably more productive in a pool of people using similar skills to support many customers rather than just one. It is similar to hiring a cleaning service rather than employing individual cleaners in a department of cleaning because cleaning things is not a core competency of business.
It might be less irrational if some amount of compute is part of the core competency of the business. Since "software is eating the world," compute is a core competency of all businesses except for the ones that don't realize it yet.
I've not really seen this work out well. I think it might be true for simple set-ups, letting a tiny developer team also handle infra and support without going nuts doing it, if they set it up that way from the beginning, but more-complex setups always seem to have so damn many sharp edges and moving pieces that support ends up looking similar to what a far more DIY approach (short of building one's own datacenter outright) would, in terms of time lost to it.
... and so does downtime, for that matter.
In the real world, for baseline load, the big advantage for many large companies isn't price, but the massive lack of alacrity of many inhouse ops teams. If it takes me 3+ months to provision compute for the simplest, lowest demand services (as is the custom in many large companies full of red tape and arguments about who bears costs), letting teams just spin up anything they want and get billed directly is often a winner, even if it's more expensive. Having entire teams waste months before they can put something in prod is a very different kind of expense in itself.
The cloud native approach would be to modify your app so that it can be scaled up and down so you keep a few machines always running, and scale up and down with your traffic so you only run at capacity when you need it.
Anecdotally, when my previous company was looking at costs, cloud unequivocally came out significantly more expensive, and that wasn’t even a large company (only 2,000 or so employees).
I will grant that we did not have globalization problems to solve (but I’d also wager that lots of businesses prematurely “what if” this scenario anyway).
If you neeed 4 CPUs for your peak load for 4 hours per day, and only 1 of them for the other 20 hours a day, you can save by scaling down to 1 cpu for 85% of the day.
It's also extremely uncommon to have loads that spiky.
And when you do, hybrid is often a solution (use a provider that can provide colo or managed servers for your base load and cloud instances for your peaks, or scale across providers).
Even at that, you said yourself that you can use "cloud" to scale into your spikes.
It takes very unusual load patterns for cloud to win on cost. It does occasionally happen, but far less often than developers tend to think.
There many reasons to choose cloud services, but cost is almost never one of them.
Guy asked what is a cloud workload, I responded. Nitpicking every tiny detail doesn't help.
> There many reasons to choose cloud services, but cost is almost never one of them.
It's cheaper to pay me to manage IAM roles for lambdas and ECS instances for 5% of my time than it is to pay someone full-time to manage some sort of VMware or other system. It's easier and cheaper to find someone with experience with AWS who can provide value to the team and product than it is to find someone who can manage and maintain a cobbled together set of scripts to update apps. There are click and go options for deploying major self hosted services like grafana, k8s with secure details that I can use without spending any time (and time == $$$) learning about the developers preferred deployment scheme.
> It's cheaper to pay me to manage IAM roles for lambdas and ECS instances for 5% of my time than it is to pay someone full-time to manage some sort of VMware or other system.
True, but it's a false equivalence, and one I often see used from people unaware of the ease of contracting this out on a fractional basis.
I used to make a living cleaning up after people who thought cloud was easier, who ended up often spending a fortune untangling years of accumulated cruft that just never happened for my non-cloud customers.
Sure, sure, you’re about to say something about discounts? Granted, that’s available, but only for commitments starting at one year or longer!
Okay, fine, I actually agree that there are savings available by reducing head count. The entire network and storage teams can be made redundant, for starters. Even considering that DevOps and cloud infra engineers need to be hired at great expense, this can be a net win…
…but isn’t in my experience. Managers are unwilling or unable to make many people redundant at once, so they stick around and find things to do…
… things like reproducing the mess that kept them employed, but in the cloud.
I’m watching this unfold at about a dozen of my large enterprise customers right now.
Got to get back to work and send the fifty seventh email about spinning up a single VM. Got to run that past every team! It’s no rush, it’s only been about fourteen months now…
Which… isn’t to say anything about which way we should expect that to swing things. But it seems quite unusual, as most companies have not been bought by a cloud provider. Yet…
Nope, LinkedIn executes completely independently.
The pressure to move to a cloud and to Azure specifically both comes entirely from MS. Linkedin was perfectly happy then and now running its own setup, this 4 years of trying to move is because of MS ownership.
If you treat us-west-1 as a single data center, you may find you are spending a lot on traffic between AZs.
A lift and shift might treat us-west-1 like a single data center. A more sophisticated strategy might treat it as three.
When we were acquired by MSFT we had the same project. We had to move from AWS to Azure. I made them all stop saying "lift and shift" because in reality it is "throw away all of your provisioning code and rewrite it using Azure primitives which don't work the same way as AWS ones".
It is more akin to writing an iOS app to work on Android.
I mean, if you're already with AWS using their services (besides EC2 for hosting) such as RDS or S3; moving to Azure SQL (or DB for MySQL or whatever) and Blob Storage is not just lift-and-shift anymore, since you are actually changing from a cloud provider to a different one.
AFAIK an actual migration to the cloud would involve rewriting some parts of the application to be cloud-native, such as using Service Bus for queues instead of a local Redis/RabbitMQ instance, using GCS instead of local disks, and using RDS instead of hosting your own single MySQL server.
I've always read it as being roughly analogous to "like for like," and dependent on the specific circumstances and status quo.
Similar terminology is "forklift"... been hearing that one for well over a decade.
Migrations are oftentimes an opportunity to revisit scaling, configuration, build and deployment pipelines, platform primitives, etc. Every migration I've been involved in has a (probably necessary) tension between getting the job done efficiently, while not repeating all the mistakes of the past.
And it’s not just the service fees. I blanche to think of the opportunity costs we accrued by focusing for that long on infrastructure to the exclusion of new product and features. It’s truly breathtaking.
And then there’s the burnout, and the ruffled feathers.
Even if done skillfully with valid rationale, they don't show any value until you come out the other side successfully.
They were worried the old vendor might go under. My own track record with predicting company failures is pretty bad, so I suspect they'll still be around ten years from now.
I've been on enough Cloud Archeology expeditions into the land of VMs where nobody knows what they do, it might as well be my job title now.
If refactoring is too hard for a Microsoft owned company what am I to think about my tech stack?
Reality has a surprising amount of detail, and any non-trivial, customer-facing system will have accumulated weird code paths to account for obscure but nonetheless expensive edge cases. A codebase built across >20 years, scaled to support millions of concurrent users is going to be absolutely filled to the brim with weird things.
When you add the need for live migrations with zero downtime, done every few years to account for next order of magnitude loads, you end up with a proper Frankenstein's monster. It's not called "rebuilding an airplane while flying" for a lark.
Every round includes a long, complex engineering effort of incremental live migration. With parallel read/write patterns between old and new systems, and all their annoying semantic differences. And then, to add insult to injury, while your core team was going through the months-long process of migrating one essential service, half a dozen upstream teams have independently realised they can depend on some weird side effects of the intermediate state and embedded its assumptions to a Critical Business Process[tm] responsible for a decent fraction of your company's monthly revenue. Breaking their implicit workflow will make your entire company go from black to red, so your core team is now saddled with supporting the known-broken assumptions.
Then you get to add wildly differing latency profiles to the mix. While you were running on your own hardware, the worst-case latency was rack-to-rack. Implicit assumptions on massive but essential workloads may depend, unknowingly, on call-to-call latencies that only rarely exceed 100 microseconds. In a cross-AZ cloud setting you may suddenly have a p90 floor of 0.2ms. A lot of software can break in unexpected ways when things are consistently just a little bit too slow.
Welcome to the wonderful world of distributed systems and cloud migrations. At some point the scars will heal. Allegedly.
https://arstechnica.com/civis/threads/i-thought-hotmail-was-...