We don’t use Kubernetes
ably.com
ably.com
> Ably is a Pub/Sub messaging platform that companies can use to develop realtime features in their products.
At this time, https://status.ably.com/ is reporting all green.
Althought their entire website is returning 500 errors, including the blog.
It is very hard not to point the irony of the situation.
In general I would not be so critic, but this is a company claiming to run highly available, mission critical distributed computing systems. Yet, they publish a popular blog article and it brings down their enire web presence?
This reaction is a storm in a tea cup left inside of a sunny cottage.
Downvoters: I don't see anyone screaming at this stock trading app that also stopped using Kubernetes [0].
I run on k8s and it works really well
I suspect the sort of individual visiting this website and reading the blog will know that their services and blog do not run on the same machines.
95% of the comments here parade about a hiccup of this service having a blog going down when that is not even their main service but continues to use an unreliable service like GitHub even when that goes down every month and they also don't use K8s either.
GitHub (and GitHub Actions) was proven to be very unreliable and the whole service went down multiple times and somehow that gets a pass? Even GitHub's status page returned 500s due to panicking developers DOSing the site. Same with GitHub pages.
But a single blog getting 500s but the whole service didn't go down? Nope.
Seriously, a Wordpress blog goes down (like that’s never happened under HN load) and all of HN is saying “lol what dogshit if only they used k8s for basically static content, the morons” which tells you a lot about the psychology going on here. A LOT. Otherwise smart engineers just can’t help themselves and redefine “irony” even though it contributes absolutely nothing to the discussion and doesn’t even address the fundamental criticism.
We are in a thread that essentially started with “the blog is down, how can I trust the product?” If you can figure out that logic please let me know. That’s basically saying “I don’t understand operations at all,” but here we are, listening to it and its Strong Opinions on a resource allocator and orchestrator.
This thread basically confirmed for me that Kubernetes is losing air. I already knew that, but the signal is getting a little bit stronger every month.
Kubernetes isn’t perfect but it’s the best we have. Losing air to what?
"the following question comes up a lot:
“So… do you use Kubernetes?”
This has been asked by current and potential customers, by developers interested in our platform, and by candidates interviewing for roles at Ably. We have even had interesting candidates walk away from job offers citing the fact that we don’t use Kubernetes as the reason!"
Why do you think this is an "edgy blog post"? This is a piece of Marketing, both recruitment and regular. The implied subtext of the post is: "we don't use Kubernetes but its not because we don't know what we are doing".
Separately, what does having studied distributed systems or not have to do with Kubernetes?
The distributed state algorithms that underlying kubernetes are what makes it what it is. I see very little added to the discussion. I see a hacked together AWS system that will be subject to the laws of entropy.
Plus this doesn't look like a heavy microservice environment. With a monolith or small number of services that aren't changing often with dedicated LBs and autoscale groups I'm guessing their environment is fairly easy to monitor and manage.
Best system for what?
In addition to insight in their infrastructure practices, this gives you a unique opportunity to look into how they deal with outages, whether they update their status page honestly (and automatically), and how fast they can solve technical issues.
It also highlights a growing pet peeve of mine: the uselessness of status dashboards, if not done very well.
They're much harder than they look for complex systems. Most companies want to put a human in the loop for verification/sanity check purposes, and then you get this result. Automate updates in too-simplistic a way, and you end up with customers with robots reacting to phantoms.
My experience with monitoring 100% of outgoing API calls is that many services have well below a 100% success rate. Sometimes there are just unexplained periods of less than 100% success, with no update to the status page, and sometimes there's even a total outage. (I had a sales call with one of these providers, and they tried to sell me a $40,000 a year plan. I just wanted their engineers to have my dashboards so they can see how broken their shit is.)
The one shining light in the software as a service community is Let's Encrypt. At the first sign of problems, their status page has useful information about the outage.
I'm not surprised when a company's status page only reports on (highly available) services, not the website or blog—which are likely run by marketing, not engineering.
Still, it's simple and free to set up status pages so there isn't much excuse.
I've been here in a similar situation, and my guess is they've:
1. Reverse-proxied /blog to a crappy Wordpress instance run without caching.
2. HN traffic killed the blog.
3. Cloudflare, in their infinite wisdom, if they see enough 5xx errors from /blog, will start returning 5xx errors for every uri on the site, with some opaque retry interval.
4. Voila, the entire site appears to be dead.
1) The realtime service and website are different things. The blog post is talking about the service, which has been continuously available.
2) Oops, the website fell over. We'll fix that. Thanks for all the advice :)
I can't say for certain that this is what happened here, and the irony is definitely there, but overall it's not valid to judge a company's reliability on a piece that is not what they are selling, and are likely to not even be managing themselves.
> Our blog and website are experiencing load-related issues, leading to slow loading and 5xx errors.
All backend services (realtime and REST apis) are on entirely separate infrastructure and are unaffected.
You may expect that company blog have a bit lower uptime than services it offers. As someone on purchasing side, I don't give a f** if a company website or blog is down for weeks, if their services are operational.
As a disclaimer - never heard of Ably, but wholeheartedly support non-kubernetes environment. Being CTO in large e-commerce company we do not use or even plan to use kubernetes.
Perhaps my team runs a simpler cluster, but we have been running a Kubernetes cluster for 2+ years as a team of 2 and it has been nothing less than worth it
The way the author describes the costs of moving to Kubernetes makes me think that they don't have the experience with Kubernetes to actually realize the major benefits over the initial costs
It doesn't bear much weight though, and I've had experience with other toolings largely to the same effect.
I agree with the sibling comment about upgrades - this is IMO where GKE really shines as a cluster upgrade is mostly a button-pressing exercise.
Just use GKE or whatever the equivalent managed offering is for your cloud provider. A couple of clicks in the ui and you have a cluster. Its really easy to run, I rarely have to look at the cluster at all.
Was brought on as a consultant years ago and found their bespoke setup of random ec2 instances and lambda scripts to be far more difficult to understand than just spinning up a managed Kubernetes cluster and having a generic interface to deploy the application, as well as monitoring, logging, metrics, etc.
This, to me, is the biggest advantage of kubernetes. Yes, you can do all the things yourself, with your own custom solution that does everything you need to do just as well as kubernetes.
What kubernetes gives you is a shared way of doing things. By using a common method, you can easily integrate different services together, as well as onboard new hires easily.
Using something ubiquitous like kubernetes helps both your code onboarding and your people onboarding.
Helm charts are used by 99% of the open source projects I've seen that run on top of Kubernetes. They are all written in a similar style so transferring settings between them is fairly easy. Helm will even create a barebones chart for you automatically in the common style.
You used to work with Ably?
I've been running my own kubernetes cluster on a Raspberry pi. Does my cat count as an engineering team?
I'm sure the big cloud providers make it easy for end users to use, but that doesn't help me.
Or is this just the Jenga Model of Modern Software Development in action?
K8s was a lot more simple earlier on. It's actually dying from adoption: There's a ton of dumb ideas bolted on top, that have become "standard" and "supported" because of demands from bad customers. The core is very clean, though, and you rarely need to interact with that.
The networking part is always the most challenging. Everything between your router and your kubernetes cluster should still be routed and firewalled manually. However, if you can live with your home router firewall and a simple port mapping to your machines/cluster, then routing the traffic and setting up the cluster should be relatively painless.
Knowing nothing about K, I’m constantly wondering how it could be simpler than dumping a binary into a PaaS (web) hosting context and letting it do its thing. I’m interested to learn.
At the end of the day it's not necessarily Kubernetes that's difficult/time consuming, it's all the context around it.
How do you monitor/observability? Security? What's your CI/CD to deploy everything? How do you fully automate everything in reproducible ways within the k8s clusters, and outside?
I've been doing infrastructure management / devops for well over 23 years. Built & managed datacenters running tens of thousands of servers.
Kubernetes and AWS have made my life easier. What's really important is having the right senior people fully aware of the latest practices to get started. Greenfield now is easier than ever.
By myself.
But, I have to say that kubernetes is not the devil. Lock-in, is the devil.
I recently underwent the task of getting us off of AWS, which was not as painful as it could have been (I talk about it here[0])
But the thing is: I like auto healing, auto scaling and staggered rollouts.
I had previously implemented/deployed this all myself using custom C++ code, salt and a lot of python glue. It worked super well but it was also many years of testing and trial and error.
Doing all of that again is an insane effort.
Kubernetes is 80% of the same stuff if your workload fits in it, but you have to learn the edge cases, which of course increases tremendously from the standard: python, Linux, terraform stuff most operators know.
Anyway.
I’m not saying go for it. But don’t replace it with lock-in.
[0]: https://www.gcppodcast.com/post/episode-265-sharkmob-games-w...
Mostly it’s an issue of perception too, a cloud saves me time. If it doesn’t save me time it is not worth the premiums and for our case- it would not save time. (Due to the complexity mentioned before).
But here’s part of list (with project specific items redacted):
3am topics:
* Project name (impossible to see which project you're in, usually it's based on "account" but that gets messed up with SSO)
* instance/object names (`i-987348ff`, `eip-7338971`, `sub-87326`) are hard to understand meaning of.
* Terminated instances fill UI.
* Resources in other regions may as well not exist, they're invisible- sometimes only found after checking the bill for that month.
Time cost topics (stumbling things that make things slower):
* Placements only supported on certain instances
* EBS optimised only supported on certain instances
Other:
* Launch configurations (user_data) only 16KiB, life-cycling is hard also, user-data is a terrible name.
* 58% more objects and relationships (239 -> 378 LoC after terraform graph)
* networking model does not make best practice easy (Zonal based network, not regional)
* Committed use (vs sustained use) discounts means you have to run cost projections _anyway_ (W.R.T. cost planning on-prem vs cloud)
* no such thing as an unmanaged instance group (you need an ASG which can be provisioned exclusively with a user-data (launch script in real terms)
* managed to create a VPC where nothing could talk to anything. Even cloud experts couldn't figure it out, not very transparent or easy to debug.
Sticky topics (things that make you buy more AWS services or lock-in):
* Use Managed ES! -> AWS ES Kibana requires usage of separately billed cognito service if you want SAML SSO.
* Number of services brought in for a simple pipeline: https://aws.amazon.com/solutions/implementations/game-analyt...
* Simple things like auto-scaler instances being incrementally named requires lambda: https://stackoverflow.com/questions/53615049/method-to-provi...
* CDK/Cloudformation is the only "simple" way to automatically provision infra.
In fact; cost did not factor at all.
For almost everyone, I'd say: just pick a cloud provider, stick to it, and (unless your whole business model is about computing resource management) your time is almost certainly better spent on other things.
I'm working at a company that had moved into AWS before I joined and I don't see it ever moving out of AWS. Of course we have some issues with our infrastructure, but "we're stuck at AWS" is the least of my concern. Any project to move stuff out of AWS is not going to be worth the engineering cost.
Usually I’m not working in startups. Usually I’m responsible for 10s-100s of millions of Euro projects.
Of course being pragmatic is a large part of what has led me to a successful career; and in that spirit of course whatever works for you is the best and I’m not going to refute it.
I would also argue for a single VPS over kubernetes for a startup, it’s incredibly unnecessary for an MVP or a relatively limited number of users.
But I wouldn’t argue for the kind of lock-in you describe.
I have seen many times how hitching your wagon to another company can hurt your long term goals. Using a couple of VPSs leaves your options open.
As soon as you’re buying the kook-aid of something that can’t easily be replicated outside the you’re hoping that none of the scenarios I’ve seen happen again.
Things I’ve seen:
Locayta: a search system, was so deeply integrated that our SaaS search was permanently crippled. That company went under but we could not possibly move away. It was a multi-year effort.
One of our version control systems changed pricing model so that we paid 90x more overall. We could do nothing because we’d completely hitched our wagon on this system. Replacing all build scripts, documentation and training users was a year long effort at the least. So we paid the piper.
This happens all the time: Adobe being another example that hasn’t impacted me directly.
It’s important in negotiations to be able to walk away.
I was wrong. What I didn't understand is how easy kubernetes is to use for the application developers. It's the most natural, seamless way to do Docker-as-deployment-packaging.
If you're going to have infrastructure guys at all, might as well have them maintain k8s.
Orchestration, in general, isn't needed. The major cloud providers are already doing it with virtual machines, and have been for a long time. It's better isolation than Kubernetes can provide.
Sure it's not needed! But cloud providers aren't needed in much the same way.
Whatever convenience you think Kubernetes provides over EC2/autoscaling (which Kubernetes uses, by the way) is several orders of magnitude less than the convenience of using a cloud provider.
That you would draw on-prem versus cloud as an equivalency to Kube vs other deployment methods reeks of inexperience, to me.
edit: Oh no, I seem to triggered a drive by Kube fanboy or two. Yes, stealth downvote because you disagree without defending your position. You will do so much to show how right you are.
I'm a definitely fan of K8s, but I'm not defending it here, however, to saying orchestration isn't needed is silly. In a way what AWS provides is orchestration, it's just for VMs instead of containers.
As a devops engineer I've worked with a lot of individuals and with a lot of tooling, and so far I can only say my opinion for container orchestation has only grown stronger. I recall having to explain to certain developers how they have to first figure out more than 5 different services for AWS, then use packer to build an AMI, which is provisioned using Chef, then they have to terraform the relevant services, then use cloud-init to pull configuration values. All in all I had to do most of the work, and the code was scattered in several places. Compare that with a Dockerfile, a pipeline, and some manifest in the same repo. I've seen teams deploy a completely new application from first commit to production in less than a week, with next to zero knowledge of K8s when they started. The same teams, who weren't pros but had a bit of experience with AWS, took several weeks to deploy services of similar complexity on EC2. Saving two weeks of Developer's time and headaches is a lot, considering they're one of the most expensive expenses a company has.
That sounds horrible. You can actually do all that from a single repository, using just ansible alone, with one cli invocation:
- use packer to build an AMI => ansible
- provisioned using Chef => ansible
- terraform the relevant services => ansible
- use cloud-init to pull configuration values => baked into the image
People think that Kubernetes is just for service orchestration and it's true, it's very good for that. But what really sets it apart for me is the ability to extend kubernetes through operators and hooks that enables platform teams to really start to abstract away most of the underlying platform.
A good example is the custom resources we've created that heavily abstract away what's happening underneath. The apps gives us a docker image and describe how they want to deploy and it's all done.
1) Churn + upgrade horror stories,
2) Apparent heavy reliance on 3rd party projects to use it well and sanely which can (in my experience, in other ecosystems) seriously hamper one's ability to resist the constant-upgrade train to mitigate problem #1.
Basically, I'm afraid it's the Nodejs (or React Native—folks who've been there will know what I mean) of cluster management, not the... I dunno, Debian, reliable and slowly changing and sits-there-and-does-its-job. Does it seem that way to you and is just so good it's worth it anyway, or am I misunderstanding the situation?
But service discovery and zero downtime upgrades are not that hard to implement and maintain. Our company's implementations of those things from 12 years ago have required ~zero maintenance and work completely fine. Sure, it uses Zookeeper which is out of vogue, but that also has been "just working" with very little attention.
A thought experiment where we had instead been running a Kubernetes cluster for 12 years comes out pretty lopsided in favor of the path we took for which one minimizes the effort spent and complexity of the overall system.
Using deployment scripts + docker alone would be insane, even at our small scale.
> Packing servers has the minor advantage of using spare resources on existing machines instead of additional machines for small-footprint services. It also has the major disadvantage of running heterogeneous services on the same machine, competing for resources. ...
Have a look at your CPU/MEM resource distributions, specifically the tails. That 'spare' resource is often 25-50% of resource used for the last 5% of usage. Cost optimization on the cloud is a matter of raising utilization. Have a look at your pods' use covariance and you can find populations to stochastically 'take turns' on that extra CPU/RAM.
> One possible approach is to attempt to preserve the “one VM, one service” model while using Kubernetes. The Kubernetes minions don’t have to be identical, they can be virtual machines of different sizes, and Kubernetes scheduling constraints can be used to run exactly one logical service on each minion. This raises the question, though: if you are running fixed sets of containers on specific groups of EC2 instances, why do you have a Kubernetes layer in there instead of just doing that?
The real reason is your AWS bill. Remember that splitting up a large .metal into smaller VMs means that you're paying the CPU/RAM bill for a kernel + basic services multiple times for the same motherboard. Static allocation is inefficient when exposed to load variance. Allocating small VMs to reduce the sizes of your static allocations costs a lot more overhead than tuning your pod requests and scheduling prefs.
Think of it like trucks for transporting packages. Yes you can pay AWS to rent you just the right truck, in the right number for each package you want to carry. Or you can just rent big-rigs and carry many, many packages. You'll have to figure out how to pack them in the trailer, and to make sure they survive the vibration of the trip, but you will almost certainly save money.
EDIT: Formatting
Or even those with caching.
A proof of concept demo: https://github.com/MatthewSteel/carpool
PHP in an ordinary deployment behind Apache2 or nginx is very fast. If you're writing a Web service and your top two criteria for the language are "fast" and "scripting language" and PHP doesn't at least make your short-list, you've likely screwed up. It's the one to beat—you pretty much have to go with something compiled to have a good chance of beating it at serving dynamic responses to ordinary Web requests, and even then it's easy to screw up in minor ways and end up slower. Persistent connections, sure, maybe you want something else, but basic request-response stuff? Performance is not a reason to avoid it, unless your other candidate languages are C++ or very well-tuned and carefully written Java or Go, something along those lines. Speed-wise, it is certainly in a whole other category from Ruby or Python (in fact, even "speedy" Nodejs struggles to match it).
Now, WordPress is slow because a top design goal is that it's infinitely at-runtime extensible. And lots of PHP sites are slow for the same reason that lots of Ruby, or Node, or Python, or Java, or .Net, or whatever, sites are slow: the developer was very bad at querying databases.
EDIT: I'm not saying these softwares can't do more, it's just that usually on most default configurations, people don't bother to use a CDN, optimize the database, use a caching plugin, etc. They just install it, then install another 39 plugins and then ask why everything is slow. It's very common to see WordPress websites failing under load (and I've helped many to optimize their installs so I know it's possible, it's just not what the average WordPress install looks like).
I used to have a 5mil pageviews per month site running on 5$ instance from DigitalOcean
I have known teams move to back to very simple default wordpress hosting from more advanced stacks.
My point is it harder to get your marketing dept to upgrade tech even if we could do it.
I have learnt over the years that solutions that have to be for the people who will use everyday, even if it's poorly managed WordPress.
In your example You or your SRE/devops will do all the basic configuration tunning pretty much out of the box. Your IT dept may not be able to at all or unless someone tells them to
It meant that we -never- had naked requests hitting the underlying CMS, so our security footprint there was miniscule (to access it you'd need to already be in the network), there was no chance of failure (i.e., even with a caching layer, a bunch of updates and enough traffic could lead to a spike hitting the CMS, and in fact, if there was a bad edit in the CMS that lead to part of the site being malformed, the crawler would die, alerting us to the bad edits before going live), and we could host using any static site/CDN we wanted to, with the only downside being deploys took a while. Even that downside we could have worked around, if we'd wanted to get into the CMS internals enough to track modifications; we just really didn't want to.
I was tempted to building something like like you did, but came to the same conclusion many do : it is not our core business and marketing will still have issues as the offering won't be as good as a professional one.
Even a solution to bridge the two setups needs dev time to maintain. End of the day managed solutions like Hubspot comes in cheaper in total cost of ownership even though all of this kind of architecture is technically superior.
The move for us was before netlify became popular so never could consider it.
That's why I was tempted to build, then I realized bulk of the work will be to build the graphical UI editor and templating tools - the parts I hate to begin with.
Awfully presumptive there aren't you. :P
I've had software solutions delivered by consultants in past jobs that were billed as being 'complete', that were missing that. 10 rps = Django fell over, site is down. That was the first thing we fixed.
I’m no WP expert but maybe they have a good reason to leave caching out of the core.
Unpatched WP instances full of vulns are also rampant for this reason.
[1] See, for example, https://ably.com/support , which is currently still up.
[2] The cf-cache-status header says "DYNAMIC", which means "This resource is not cached by default and there are no explicit settings configured to cache it."
I'm pretty sure cf-cache-status is added by CF—that's not the site saying not to cache it, it's CF reporting that it's not cached (which, again, I think is the default, not something the site owner deliberately turned off).
Need to use a "Page Rule" telling CF to cache everything, and provide the required Cache-Control header values from app side.
It's not an option in the caching settings (or need a enterprise plan?)
Enterprise is largely about very high levels of bandwidth use, serving non-Webpage traffic, and all kinds of advanced networking features. It's also a hell of a jump in price from the $200/m top-tier "self-serve" plan (which is also the only self-serve plan with any SLA whatsoever, so, the only one any business beyond the "just testing the waters" phase should be on)
But the Cache Settings page, it doesn't have the "Cache Everything" option, or it's reserved to Enterprise plan.
It's covered in what I'd consider to be one of a handful of docs pages that're must-read for most any level of usage of the service:
https://support.cloudflare.com/hc/en-us/articles/200172516-U...
(it also links to a page that describes exactly how to create a "cache everything" page rule)
Then again, most orgs have a ton of stuff they ought to do and haven't gotten around to. You just hope it doesn't bite you quite so publicly and ironically as this.
If you're going to make a big deal about being contrarian with your technological choices, you either have to be really good at what you do, or be really honest about what tradeoffs you're making. I'd be extremely embarrassed if I pushed this blog post out.
I was serving static HTML off a $5 VPN and later off Cloudflare Pages. The response times were stable in both cases for all users.
Now, application servers and database tiers, that's a different story. Epic novel, really.
You gotta admit, it’s a pretty bad look to come here with a boldly titled systems architecture blog post and have your website crash…
A junior web developer nowadays can easily build a system to serve static content that avoids all but the worst hugs of death.
If this wasn’t a hug of death then you have service availability issues, regardless…
And cloudflare caching of some previous successful page fetch seems to be turned off as well. A cf-cache-status header of "DYNAMIC" means "This resource is not cached by default and there are no explicit settings configured to cache it.".
Having this happen to a post bragging about infrastructure is just adding insult to injury.
Edit: As another comment mentioned, this site is actually behind Cloudflare, with cf-cache-status: DYNAMIC (i.e. caching is disabled). I don't know what to say...
I don't even know what those scripts are supposed to do. This is a static blog post.
You don't need Kubernetes for a blog. If you do then something is wrong with your whole backend stack.
Perhaps they are the smart ones who want to save money and not burn millions of VC money.
The one that REALLY needs to use Kubernetes is actually GitHub. [0]
All I had to learn is Docker and Kubernetes. If Kubernetes didn't exist I would have had to learn myriad tools and services, cloud-specific tools and services, and my application would be permanently wedded to one cloud.
Thanks to Kubernetes my application can be moved to another cloud in whole or in part. Kubernetes is so well designed, it is the kind of thing you learn just because it is well designed. I am glad I invested the time. The knowledge I acquired is both durable and portable.
Jokes aside, it sounds like they should just use ECS instead.
> A small custom boot service on each instance that is part of our boot image looks at the instance configuration, pulls the right container images, and starts the containers.
> There are lightweight monitoring services on each instance that will respawn a required container if it dies, and self-terminate the instance if it is running a version of any software that is no longer the preferred version for that cluster.
They've built a poor-mans Kubernetes that they won't be able to hire talent for, scales slower and costs more.
> Functionally, we still do the same, as Docker images are just a bunch of tarballs bundled with a metadata JSON blob, but curl and tar have been replaced by docker pull.
but, yes, I agree; it's a hand-made Kubernetes
This is one of the first cases where I think that maybe Kubernetes would be the right solution, and it's an article about not using it. While there's a lot of information in the article, there might be some underlaying reason why this isn't a good fit for Kubernetes.
One thing that is highlighted very well is that fact that Kubernetes is pretty much just viewed as orchestration now. It's no longer amount utilizing your hardware better (in fact it uses more hardware in a many cases).
I would have agreed with this statement 2 years ago but now I think K8s has been commoditized to the point where it makes sense for many organisations. The problem I have seen in the past is that every team rolls their own orchestration using something like Ansible, at least K8s brings some consistency across teams
Absolutely, but we still need to lose the control plane for it to make sense for small scale systems. Most of my customers run on something like two VMs, and that's only for redundancy. Asking them to pay for an additional three VMs for a control plane isn't feasible.
I think we need something like kubectl, but which works directly on a VM.
I had something like a desktop PC with 1GB of RAM, maybe 20 years ago, don't remember exactly how long ago. It was average, not too much but not too low either. Once I installed Oracle DB, it was idling at 512MB. Absolutely no DBs and no data set up, just the DB engine, idling at half the RAM of an average desktop PC.
If your scale isn't high enough though I think you're right, it's often not worth the complexity and Kubernetes' overhead might actually be detrimental.
Instead it is easier to to critique the low hanging fruit point rather than discussing their actual reason for not using this 'kubernetes' software.
So is their blog the main product? If not, then the 'their blog gone down lol' quips are irrelevant.
I found their post rather interesting and didn't suffer any issues the rest of the 'commenters' are facing.
* Well four, but one of the mentions of "custom x" is talking about k8s way of doing things.
ably cto: Go talk to the CMO, I can't help.
ably ceo: What!!
ably cto: Remember when you said the website was "Totally under the control of the CMO" and "I should mind my own business"? Well I don't even have a Wordpress login. I literally can't help.
Google Cache doesn't work either https://webcache.googleusercontent.com/search?q=cache:YECd_I...
Luckily there's an Internet Archive https://web.archive.org/web/20210720134229/https://ably.com/...
In fact there is probably a growing problem that the ways this could have been solved in Kubernetes is slowly approaching the number of ways that it could have been solved on good ol' baremetal/ec2's. Kubernetes != non-bespoke system.
I like the fact, however, that there do at least exist standard primitives in Kubernetes that could approach horizontal scaling (if that was the issue as is insinuated by the aforementioned laughing) since the best thing Kubernetes does for me is helping me keep my cloud vendor (aws,gcp,azure,etc) knowledge shallow for typical web hosting solutions.
They have essentially semi mocked Kubernetes without any of the benefits.
If your orchestration needs aren't "use every last % of server capacity" you might not need k8s.
And if you eventually need to use Kubernetes, you can always spin up EKS. Just don't rush to ship on the Ever Given when an 18 wheeler works fine.
It's OK to not use k8s. We should normalize that.
I also now have a k3s cluster at home. The learning curve was insane, and I hated it all for about 8 weeks but then it all just clicked and it's working great. The arrogance to roll your own without assessing the standard fully speaks volumes. Candidates figured that out and saw the red flag.. Writing your own image bootstrapper... What about all the other features, plus the community and things h things like helm charts.
ECS Fargate has been awesome. No reason to add the complexity of K8s/EKS. We're all in on AWS and everything works together.
But this... you guys re-invented the wheel. You're probably going to find it's not round in certain spots too.
And we got to maintain the code all by ourselves too! It might take a bit too long to implement a new feature, but hey! its ours!"
Really, Kubernetes is complex, but the problem it solves it even more complex.
If you are ok solving a part of the problem, nice. You just built a competitor to google. Good luck hiring people who come in already knowing how to operate it.
Good luck trying to keep it modern and useful too.
But I totally understand the appeal.
Where I disagree with this article is on Kubernetes stability and manageability. The caveat is that GKE is easy to manage and EKS is straightforward but not quite easy. Terraform with a few flags for the google-gke module can manage dozens of clusters with helm_release resources making the clusters production-ready with very little human management overhead. EKS is still manageable but does require a bit more setup per cluster, but it all lives in the automation and can be standardized across clusters.
Daily autoscaling is one of those things that some people can get away with, but most won't save money. For example, prices for reservations/commitments are ~65% of on-demand. Can a service really scale so low during off-hours that average utilization from peak machine count is under 35%? If so, then autoscale aggressively and it's totally worth it. Most services I've seen can't actually achieve that and instead would be ~60% utilized over a whole day (mostly global customer bases). The exception is if you can scale (or run entirely with loose enough SLOs) into spot or preemptible instances which should be about as cheap as committed instances at the risk of someday not being available.
It _forces_ you to become a "yaml engineer" and to forget the other part of the systems. I was interviewed by a company and when I replied the next step I could do was to write some operators for the ops things, they simply rejected because I'm too experienced lolz
I celebrate a diversity in opinion on infrastructure but… if I was a CTO/VP of engineering and I read that line, that would be enough to convince me to use kubernetes.
This type of language isn't something that should come out of a company, and may be a signal that there are other reasons developers refused to offer their services other than they just don't use K8s.
I'm saying this because while their architecture seems reasonable, albeit crazy expensive (though I'd say it's small-scale if they use network CIDRs and tags for service discovery), it also seems like they wrote this without even trying to use Kubernetes. If they did, it isn't expressed clearly by this post.
For instance, this:
> Writing YAML files for Kubernetes is not the only way to manage Infrastructure as Code, and in many cases, not even the most appropriate way.
and this:
> There is a controller that will automatically create AWS load balancers and point them directly at the right set of pods when an Ingress or Service section is added to the Kubernetes specification for the service. Overall, this would not be more complicated than the way we expose our traffic routing instances now.
> The hidden downside here, of course, is that this excellent level of integration is completely AWS-specific. For anyone trying to use Kubernetes as a way to go multi-cloud, it is therefore not very helpful.
Sound like theoretical statements rather than ones driven by experience.
Few would ever use raw YAMLs to deploy Kubernetes resources. Most would use tools like Helm or Kustomize for this purpose. These tools came online relatively soon after Kubernetes saw growth and are battle-tested.
One would also know that while ingress controllers _can_ create cloud-provider-specific networking appliances, swapping them out for other ingress controllers is not only easy to do, but, in many cases, it can be done without affecting other Ingresses (unless they are using controller-specific functionality).
I'd also ask them to reconsider is how they are using Docker images as a deployment package. They're using Docker images as a replacement for tarballs. This is evidenced by them using EC2 instances to run their services. I can see how they arrived at this (Docker images are just filesystem layers compressed as a gzipped tarball), but because images were meant to be used by containers, dealing with where Docker puts those images and moving things around must be a challenge.
I would encourage them to try running their services on Docker containers. The lift is pretty small, but the amount of portability they can gain is massive. If containers legitimately won't work for them, then they should try something like Ansible for provisioning their machines.
This is not too surprising. Candidates want to join companies that are perceived to be hip and with it technology-wise in order to further their own resume.
Personally, you don't have to be hip, but I want to be able to have another job in hand based on the tech I was working with and the work I was working on in a matter of days when I decide to bounce (or the decision is made for me). This is just good risk and career management.
(disclosure: infra roles previously)
You left out "Their bespoke infra work tends to be the genesis of the commodity infrastructure other people are using."
It's not hard for a Googler familiar with Borg to pick up Kubernetes. It's not hard for a Facebooker with some UI work under their belt to figure out React.
This is another way of saying that they don't want to read too much documentation.
If a company says they don't use any modern JS framework and stick to JQuery, I'd be out before they even explain why. Not because I want to be hip and further my own resume but because I'd hate my job if I worked there.
p.s. Not claiming it was the right tool for the job in this case, it would all depend on context.
So maybe if people hear, hey our infrastructure is running on some 10 year old Frankenstein cobbled together from Docker and Virtual Machines and AWS and on premise servers they just pass.
I’m tired of having to learn an new infrastructure management/deployment tool every time I shift jobs. Tired of having to work with company custom platform. If you’re deployed on kubernetes it would take me a day to figure out how/where everything is deployed just by poking around with kubectl. Probably around that much time to deploy new software too.
K8s is just another giant footgun. Good and bad orgs will still be good or bad, with or without it, but it definitely amplifies the bad.
If we’re looking at the median, most orgs will more or less work with k8s manifests, maybe with some abstraction layer on top to make it easier for developers.
I would walk away if they were using a technology not fit for the job, or were actively building an in-house version based on out-dated knowledge (the use of the word "minion" in Kubernetes was deprecated in 2014 IIRC).
Think about this in another realm, like plumbing. If you were a plumber, would you accept a plumbing job if the company told you that they didn't buy materials and tools, and were instead designing and building their own in-house version?
If they decided in 2014 not to use k8s, how much effort do you expect them to spend on keeping up-to-date with the latest and greatest of a technology that they're not using?
Ably using ghost blogging platform (https://ghost.ably.com) seeing requests in network console.
Also Employers: Some of our applicants turned down our job offer, when we revealed that we don't use specific technologies!
https://www.sebastianbuza.com/2021/07/20/no-we-dont-use-kube...
Another pub/sub startup publishing almost the identical blog post also today, July 20, 2021.
seems like a me too, i know more than you type of project. looks geared towards cryptos though.
This reads as simply a choice to have less complexity in their stack, which I think is admirable.