Google admits Kubernetes container tech is too complex
theregister.com
theregister.com
It's taken us a little over a year. Partly because K8s has a steep learning curve, but also because safely transitioning services without disrupting product teams adds a lot of overhead.
The investment is already yielding great returns. Developers are happy. Actual quote: "Kubernetes is the biggest quality-of-life improvement I've experienced in my career."
The same is true for me as a DevOps engineer. I never have to write a script to orchestrate a rolling deployment ever again. I'll never have to touch Puppet or wrangle OS upgrades. I no longer need to use Terraform to scale out a service.
Now that we've completed this transition, one of my goals for the next year is to get the team to a place where we spend 75% of our time on higher-level projects -- working on developer tools and Kubernetes extensions and setting site reliability standards, instead of the firefighting and operational work that has traditionally taken up so much of our time. Couldn't be happier about this change.
... That is until management decides we're overhead and fires us to free up budget for feature developers. ;)
I'm on the DevOps side as well going through the same transition, k8s also allows insane customization, and I have some colleagues that are delaying our rollout unintentionally so they can play around with developing more tooling for deployments which is really frustrating. The k8s scene seems to be filled with constant scope creep and refactoring to get it just perfect before use. Either way, I agree the benefits far outweigh this annoyance that I've experienced. I'm so excited to work on developing tooling instead with my time.
However, I don't think we're free entirely from managing servers the old way with Chef / Puppet / Ansible, unless you're purely hosted there's still the rule of thumb you shouldn't run services that hold state in k8s. But with persistent vol's I do see that changing, though I'm not sure if everyone agree's that's a good idea.
My impression is that the primary purpose of Kubernetes is to give SRE teams political air cover to rewrite a lot of their existing processes. Whether Kubernetes is actually required for that, or even net superior seems questionable. This unsexy work becomes justifiable because it's coupled to a mainstream accepted tech modernization.
You see this same phenomenon with database migrations. Where what the team really needed to do is just rewrite an app to use the existing database properly. But no one is going to approve that work. So what happens is people convince themselves that the existing tech sucks and use that to rationalize doing the rewrite. The result ends up not always being net superior, because sure you did the rewrite but you are also eating the operational cost of integrating a new technology into the org.
I’ve also seen a switch from rdbms to Hadoop because a company had “millions” of rows. Luckily on this one I only had to rewrite a handful of queries.
Wat. That's gross. It probably costs them more per query now than the rdbms did.
and in short order you will reap the savings of being able to hire people who already know your devops/infra tech stack, and can hit the ground running. not to mention being able to benefit from the constant improvements that come from outside your org.
That certainly is a thing that happens, but you could use that to dismiss any technology at all. In the case of Kubernetes, it makes operations a lot easier to the (important) effect that the development teams can do a lot of their own operations work. This is important since they're the ones who are empowered to solve operations problems and it also eliminates the blame game between ops and dev. Further, it eliminates a lot of coordination with a separate ops team--the dev teams aren't competing to get time from an ops team; they can solve their own problems, especially the most common ones. This also has the nice property of freeing the SREs to work on high-level automation, including integrating tools from the ecosystem (e.g., cert-manager, external-dns, etc).
Kubernetes certainly isn't the final stage in the evolution, but it's a welcome improvement.
No, you can't; you need three (-ish) factors:
1. The technology is sufficiently incompatible with what you're currently using that you need a rewrite to use it (eg, this generally doesn't happen with gcc -> llvm, for example).
2. The technology is sufficiently (faux-)popular that it's possible to convince a pointy-haired boss that you need to switch to it (eg, this won't work with COBOL anymore, though unfortunately it successor Java is still going).
3. The technology sucks.
And really, if you want to dismiss a technology, point 3 ought to be enough all on its own (particularly since that's presumably the reason you want to dismiss that technology).
Anyway, we're trying to assess Kubernetes' value proposition (i.e., to answer "does it suck?"). If your system for answering that question depends on already knowing the answer, it's not a very useful system.
Well, I'm not, since I already know that, but if you don't know that yet, then your position makes more sense. (That is, using "dismiss" in the sense of finding out that it sucks, rather than (as I read it) in the sense of justifying a refusal to use technology that you already know sucks.)
Unfortunately, due to market-for-lemons dynamics, it's usually not possible to convey knowledge that a particular technology sucks until things have already gone horribly wrong. See eg COBOL or (the Java-style corruption of) Object Oriented Programming.
We are running a handful of stateful services in K8s (things like MongoDB for which GCP doesn't have a compelling and affordable managed offering). It's definitely more complex than transitioning a stateless service, but so far our experiences with StatefulSets and PersistentVolumes have been good. And this allows us to sunset Puppet/OS management completely. I should note that we _are_ being extremely careful about backups. We also run each stateful service in a dedicated node pool for isolation. Who knows, maybe a year from now we'll be shaking our heads and saying "that was a TERRIBLE idea" but for now, so far so good.
We're running on GKE, so lots of things that would be hard in on-prem environments (ingress, networking, storage) are easy.
Agreed. The on-prem story is still really messy, but I think there's a lot of third-party work to build on-prem distributions that are cut and dry. Unfortunately, there are lots of them right now and it's not clear what the advantages and pitfalls are of each. Things will settle and this problem will be solved with time, but for now it's quite a pain point.
They hire other people with the supposedly same skillstack and then have them rebuild it from scratch.
Snake eating it's own tail.
But not only do I not understand K8s, I'm also an idiot.
Can you elaborate? What exactly about Kubernetes improved the developers life?
It reminds me of docker-compose file, except I can publish it too.
We didn't have these before. Yes, you can implement them without K8s (I have, at other companies) but to get the full set of features K8s provides, such as deploying N services in parallel, taking no more than X% of your capacity offline at a time, short-circuiting in the event the app is dead on arrival, connection draining with timeouts, you end up with a VERY complicated multi-threaded codebase.
2. Seamless horizontal scale-out.
Want to scale up your app in a test environment from 3 replicas to 6 to do some performance testing? This used to be a DevOps ticket that would take a few days -- DevOps engineer tries running Terraform, but oh no, a CentOS package update seems to have broken our Puppet manifests and we have to fix that. Now the developer makes a PR to the GitOps repo where they adjust a single YAML setting.
3. GitOps/ArgoCD.
ArgoCD is possibly my favorite piece of software of all time. It provides incredible visualizations for what's happening with your Kubernetes infrastructure. It really increases your confidence and trust in the system to be able watch a deployment rollout or scaling operation happen in real time. ArgoCD makes it spectacularly obvious when something has gone wrong -- you still sometimes need to go spelunking through the GCP console or use kubectl to inspect resources, but to a much lesser degree. I cannot emphasize enough how magical it is.
These are all things that our initial implementation has delivered, we are also planning to try to leverage K8s for things like on-demand pull request preview environments, hosted developer environments, and canary deploys that are MUCH harder to implement in a world without K8s (trust me, I've done it).
Although, to be fair, the article is about GKE.
But literally every cloud provider out there has a managed solution so at this point you really only need to do it for DC work or if you like to do it.
Amazon - EKS Google - GKE Azure - AKS
anything beyond those 3 is a rounding error but...
Linode - LKE DigitalOcean - DigitalOcean Kubernetes IBM Cloud Kubernetes Service Rackspace - KAAS
Creating a big-ole cluster of app servers was never really the hard ops problem which is what k8s does really well.
The benefit of k8s is the orchestration of those clusters. Spinning up 6 new http servers and getting them added to the load balancer automatically. Generating 4 new memcached nodes and getting them registered in DNS so clients pick them up and add them to the hashring.
The benefit of k8s is the scaling and elastic capabilities. It can trigger vertical scaling by spinning up larger pods or horizontal scaling by adding/removing pods.
Anyone thinking that people are using k8s because it can create app servers doesn't understand why anyone is using k8s. If all we needed to do is create a cluster of app servers, we wouldn't be using k8s.
That being said, a cluster of app servers still needs orchestration and config management and we had a ton of crazy solutions for that prior to k8s.
You can roll your own anything with enough time and manpower. Whether it makes sense to do so depends on your circumstances.
For the last several years, due to static load we've been fine with 2 instances per service. Now that we want autoscale (even though all the real load is really just the DB), it seems we could get ourselves onto AWS autoscale in ~1 month, though it would require some coding.
Spending 3 devops on 1 year doing something is a red-flag to me. If your changes save devs 1 hour per deploy, it'll pay off 6,000 deploys from now.
I think it's great you guys have 2.5 dedicated people operating 6 services across 12 prod instances, with that kind of ratio our team would be 15 people!
When you make a change manually, there's a chance that you forget something, have a typo, etc. Those problems disappear when those same things are automated.
I've used Puppet, Chef, and Ansible in production over the last decade. For me, Ansible replaced Puppet and Chef five or six years ago. For the last two years the only thing I have used Ansible for is to manage desktops and Android boxes AT HOME. Configuration Management systems like Ansible are fun and cool, but have been mostly obsoleted by the combo of Terraform and K8s. If you're on AWS, Kops can replace a lot of what Terraform does.
Kubernetes makes running services easier and more reliable. I've spent more time learning Istio that K8s itself. The learning curve of K8s is minimal compared to Drupal.
Ansible or Puppet still excel at that kind of work.
Then you can focus on containers which can be run, tested and built wherever without the fear of broken updates or one thing stepping on another. We found back in the days of ansible and chef that we had very low confidence in upgrading hosts live. So we would then do immutable hosts and blue green deploy them to production. But why think in the scope of hosts and VMs when really you have some application that needs to run somewhere.
K8s IMO isn't the end all, I think eventually we will get to something that doesn't need containers at all and you run just processes. But it is a good step for now. Also once you have your stuff containerized it makes other non k8s stuff easy like AWS Lambda
Edit: Also yes you can use those to set up generic k8s nodes but when we ran bare metal we used kubeadm to make coreOS immutable nodes. I don't think that is used anymore haven't checked but really the best way to set up k8s is to deploy really thin hosts that have nothing but Docker and k8s. VMware and others have solutions for this too where you don't have to mess with building hosts.
So if you had a fleet of ten containers running the same application in a load balanced config, I'm guessing you'd need to upgrade all of them at once (with downtime) rather than upgrading them one by one (because then the database would be inconsistent)? I'm assuming that since the containers are immutable the data is stored elsewhere.
That depends entirely on your application and the upgrade itself. Assuming we are discussing 10 different containers (e.g. 10 micro-services), k8s will normally update them in parallel but not atomically, it would be up to apication or deployment time logic to ensure they are updated as a single 'transaction'. If they are 10 copies of the same container, then k8s itself has tools for rolling upgrades where you can control the rollout.
Also, depending on application logic, the upgrade could be done in such a way that there is no need to synchronize the services, they could work with the DB as is.
Who is preparing Dockerfiles? Developers and system administrators / security people do not generally prioritize same things. We do not use k8s for now (therefore I know very little about it), so this might not be relevant but how do you prevent shipping insecure containers?
But since you are not running multiple things or users in one space in a container, something such as an out of date vulnerable library can't be leveraged to gain root access to an entire host running other sensitive things too.
In Kubernetes and docker in general one container should not be able to compromise another, or k8s. But there are other issues if an attacker can access a running container such as now having network access to other services like databases. But again these are all things that can be locked down and should be even if provisioning hosts running things.
Are you still using Terraform to specify your deployments? From experience, you need something there to manage all your yamls: deployments, configmaps, secrets. Especially if you have multiple environments.
It doesn’t end there... after complaining for a really long time two things happen. First off so much time has passed that you get these domain experts (Devops people) whose entire job is to mess with kubernetes. Second the complainers have been complaining so long that people get tired of it.
You now get people who are so tired of listening to people complain whose entire job entire job revolves around kubernetes that these people now start complaining about the complainers.
It happened with JavaScript. Javascript was around for so long people started complaining about the complainers. In fact it’s been around so long that a whole generation of people who’ve never used any other language was literally born. These people started off as the complainers against the complainers but now they outnumber the complainers so you rarely see people talk shit about JavaScript anymore.
Actually JavaScript has been around so long that the entire language has changed and part of the terribleness was fixed by making another language (typescript) compile into JavaScript.
Which brings me back to kubernetes. Kubernetes is a bad tool with no alternative precisely because it requires a 3 man dedicated team a year to get things up and running.
A good tool would be something like allows me to to get it up and running in a week just by reading some docs. Even better an hour. Could such a tool exist and replace Kubernetes? Yes. Does such a tool exist? No.
I am a complainer and you are a complainer of complainers. What will likely happen some time from now is two possible things. Kube will be so integrated into the infrastructure ecosystem that wrappers will be written on top of kube just like how react and typescript have replaced JavaScript. If that doesn’t happen then a whole new tool will replace it.
I’m sorry to say but the ideal we are shooting for here is a tool that will ultimately make Devops a general thing that all developers can deal with rather then an entire specialist team. Again no such tool exists yet but it certainly can exist, especially when the inventor of the tool has become a complainer.
If the inventor of the tool becomes a complainer, that validates the complainers. And now the complainers of the complainers have nothing left to say.
Your 5-person startup does not need a DevOps engineer. Your 50-person org should start thinking about it. Your 500-person company probably needs multiple DevOps engineers -- having every dev team independently figure out how to handle things like deployments, reliability, and security is chaotic and wasteful. There are a lot of details that only start to matter at scale, and both Kubernetes and dedicated infrastructure teams are for this use case.
Instead of complaining, why don't you build this tool? That's the problem I have with complainers.
Instead of complaining why don't you build me the tool to stop me from complaining? It's the same reason why I'm not building the tool.
That's the problem I have with complainers complaining about other complainers. Why don't you guys do something about my complaining rather then complain about it?
For me the scariest part would be that it's not tech I would use (i.e. update the descriptors) regularly, so what if something goes wrong, how quick would I be able to identify the problem. I have no answer to that because I'm off the project.
Just wait until the developers discover Google App Engine, Heroku, or DigitalOcean App Platform.
I think Heroku and DigitalOcean App Platform are still going to be popular for small setups (as will things like Amplify) but when you outgrow those (or realize you are paying too much for them) then Kubernetes is a reasonable option.
Need to run a bespoke database, yeah it can do that. Need to migrate an old service running in a VM that needs a disk, it can do that too.
I do agree that kubernetes is a pleasant experience from an application developer perspective with an existing cluster, but in my experience it was not without excessive pain and long hours by those administering the cluster. A year doesn't surprise me in your case, which brings to mind this question.
It's a real risk, at least perceptually. Mitigate it by documenting the number of developers' hours you save with automation and better infrastructure. It's pretty hard to argue against a trend of increased developer productivity. Include your own hours; you're targeting saving 30 hours a week of your own time and that frees you up to improve other things even faster.
If you're determined to be skeptical of the value-add from Kubernetes, my random post on HN ain't gonna convince you :) The developer experience improvement is obvious to people actually interacting with K8s.
I do want to clarify that this isn't the only thing we accomplished in the past year. Just the thing we shipped that was our biggest priority and had (by far) the biggest impact.
- I can store anything in a secret? Let's have thousands of cat images. Etcd then stops working because we have over 2GB of funny cats in the key store.
- I can run a root Pod? Lets mount the docker socket and start building images with it. Oh and by the way, I never clean those up and my Node simply fills up. Also I add some additional docker networks that break Pod to Pod networks.
- Istio is nice - why we don't add automatic injection for Pods in all namespaces? Including kube-system? And then they brick kube-proxy and the cluster stops working.
- I can use validating webhooks for better security? Lets watch on all resources. To keep it more secure lets set the failure policy of the webhook to Fail, so we never admit any modification without the apiserver to make a call to out webhook. Whats that? My single replica webhook has was evicted from the Pod (we didn't add any resource requests and limits) and now it cannot even be created or scheduled because kube-controller-manager and kube-scheduler cannot update their lease and they lost leadership and now are idling, effectively bricking the entire cluster.
Google would reduce the pain points with this change, however they would still face countless other issues with Kubernetes.
Sure just put your secret data on a file then we'll use your file name as the key of the secret.
Cronjobs sometimes have weird bugs as well.
A lot of its complexity is due to the fact it's an evolving system, that's fine. But I see that some things end up way more complex or unreliable than it needs due to overengineering or use cases no one needs
However playing the devil's advocate here: If you actually took the steps of learning the basic abstractions, then for me it's really hard to see what you could still get rid of.
If you actually go all-in and fit your application to the principles of Kubernetes-native applications (instead of the other way around), then it works nothing short to amazing.
We're running 120 microservices in GKE and the difference to our custom-built setup before is night and day. I let my Infra team go surfing together for two weeks because without changes it flies mostly on autopilot.
Let's not kid ourselves, distributed computing is _hard_ and Kubernetes is a testament to that. I'm not saying it can't be made more accessible by further standardization, but there are fundamental limits to how easy it can be made.
Which by the way is leading to my only pet peeve with it: I feel most of the complexity of K8S comes from the fact that it got hyped as an enterprise product and then lots of features were built that support shoving your non cloud-native workload into Kubernetes even if it was never designed for it.
If you don't do or need all of that, the amount of interface, complexity and footguns shrinks significantly. Maybe it's time to better pull them apart in the documentation.
This argument basically sums up to "Developers just need discipline, and stop blaming the tools". While this is a sound argument on paper, the intrinsic complexity of software systems make it hard to pin the blame on developers. BTW This is the same argument Uncle Bob makes which is not so popular with many mainstream developers.
You're right about feature creep in k8s though.
If your goal is to build highly reliable and available services to end users that are secure and scalable with a team of more than 10 engineers, eventually you will run into more than 50% of the concepts in Kubernetes anyway and end up re-inventing them.
Scaling up and down, node draining, finding out whether services are healthy, RBAC, resource distribution, secrets management, service hardening, introspection capabilities, explicit declaration of dependencies and endpoints and many, many more.
My point is: Sure, if your goal isn't that, it doesn't make sense to start out using Kubernetes.
But if at least eventually that's what you need, imho it's way preferable to just learn and apply well proven abstractions instead of reinventing the wheel along the way and end up with a less maintainable, capable and standardized solution you won't find anyone for maintaining.
If I hear about some of the comments here suggesting to "just spinning up docker-compose with Traefik in front" (disclaimer: I really like Traefik), then that reminds me of how some of the ops mess started that I historically had to care for.
It’s always easier to shift the blame somewhere else.
Thats why some radical but correct concepts are so hard to push.
To add a counter-example, I have lots of Ruby experience and I've just joined a Go team. I won't tell them to use Ruby, I will just do it where it makes sense and saves us time. (And then we'll have two problems... enter "limiting blast radius")
Point of my counter-example is, I'm extremely skeptical that all the world's problems can be solved by adopting a new mono-culture, whatever it is. There are 100% always gonna be some problems that are better solved in a different language. PHP is the best way to run Wordpress, for example (ok, so it's the only way to run Wordpress, but you get the idea... "Wordpress is the best way to..."), but I've been in high-functioning IT organizations that won't touch that with a ten foot pole, because "it's another language to support, and PHP is icky."
We also got rid of a perfectly fine Wiki in favor of centralized Knowledge-base software for similar reason. "Better to just have one KB. We don't need to be hosting another thing." So the chances of moving everything over to BEAM VM are next to nil, unless you are a product-focused company with just one product, or happen to have an absolute champion leading the effort to migrate all the things. For all the other things, you need to have a consistent answer too.
No tool is one-size-fits-all. Where Kubernetes shines most is under any environment that isn't running a single monolith or building a software monoculture and/or can't manage that for whatever reason (because those are all basic use cases that are frankly easy enough to manage without adding on top the additional complexity of Kubernetes; don't need it, don't use it!) IMHO, diversity in infrastructure is a plus though, and Kubernetes is a technology that it turns out enables this.
For example you still need a way to get BEAM onto hosts, still need to manage the OS on the host, still need to setup networking, RBAC etc.
In case anyone's interested, here's a pretty funny and educational talk by Bryan Cantrill about that particular incident:
GOTO 2017 • Debugging Under Fire: Keep your Head when Systems have Lost their Mind
Why would someone want to store non-secret information as a secret?
— Douglas Adams
Top reason given to me by developers: "I don't want to spend time thinking about the distinction."
If you have a managed object store or even relational database outside of k8s, the thought of storing arbitrary data in secrets probably doesn't come to mind. But if your enterprise spools up a cluster and tells you to use nfs PVCs with no other storage solution, suddenly you might start getting creative.
One of the goal of K8S is normalization/standardization of a complex topic to better share knowledge
Just in case you haven't figured out the proper way to do this, you should use docker:dind.
For our current stack, the answer has been to make the entire business application run as a single process. We also use a single (mono) repository because it is a natural fit with the grain of the software.
As far as I am aware, there is no reason a single process cannot exploit the full resources of any computer. Modern x86 servers are ridiculously fast, as long as you can get at them directly. AspNetCore + SQLite (properly tuned) running on a 64 core Epyc serving web clients using Kestrel will probably be sufficient for 99.9% of business applications today. You can handle millions of simultaneous clients without blinking. Who even has that many total customers right now?
Horizonal scalability is simply a band-aid for poor engineering in most (not all) applications. The poor engineering, in my experience, is typically caused by underestimating how fast a single x86 thread is and exploring the concurrent & distributed computing rabbit hole from there. It is a rabbit hole that should go unexplored, if ever possible.
Here's a quick trick if none of the above sticks: If one of your consultants or developers tells you they can make your application faster by adding a bunch of additional computers, you are almost certainly getting taken for a ride.
It's a somewhat odd solution for a too common problem, but any solution is still better than dealing with such an annoying problem. (source: made docker the de facto cross-teams communication standard in my company. "I 'll just give you a docker container, no need to fight trying to get the correct version of nvidia-smi to work on your machine" type of thing)
It probably depends on the space and types of software you 're working on. If it's frontend applications for example then its overkill. But if somebody wants you to let's say install multiple elasticsearch versions + some global binaries for some reason + a bunch of different gpu drivers on your machine (you get the idea), then docker is a big net positive. Both for getting something to compile without drama and for not polluting your host OS (or VM) with conflicting software packages.
The list goes on and on, it’s bizarre to me to think of this as the true, good way of doing software and think of docker as lazy. Docker certainly has its own problems too, but does a decent job at encapsulating the decades of craziness we’ve come to heavily rely on. And it lets you test these things alongside your own software when updating versions and be sure you run the same thing in production.
If docker isn’t your preferred solution to these problems that’s fine, but I don’t get why it’s so popular on HN to pretend that docker is literally useless and nobody in their right mind would ever use it except to pad their resume with buzzwords.
./configure
make
At a prompt, then install missing libs. Unless you have to maintain updates regularly, “It’s just so hard” seems like a damn meme.
Yeah, it happens with .so files, .dlls ("dll hell"), package managers and more. But that's where things like containers come in to help: "I tested Library Foo version V3.4 and that's what you get in the docker". No issues with Foo V3.5 or V3.6 causing issues... just get exactly what the developer tested on their box.
Be it a .dll, a .so, a #include library, some version of Python (2.7 with import six), some crazy version of a Ruby Gem that just won't work on Debian for some reason (but works on Red Hat)... etc. etc.
There really isn't any middle ground (except to not use third-party libraries at all).
That makes sense for a hobbyist community but not so much for production.
In a former job we needed to fork and maintain patches ourselves, keeping an eye on the CVE databases and mailinglists and applying only security patches as needed rather than upgrading versions. We managed to be proactive and avoid 90% of the patches by turning stuff off or ripping it out of the build entirely. For example with openSSH we ripped out PAM, built it without LDAP support, no kerberos support etc. And kept patching it when vulns came out. You'd be amazed at how many vulns don't affect you if you turn off 90% of the functionality and only use what you need.
We needed to do this as we were selling embedded software that had stability requirements and was supported (by us).
It drove people nuts as they would run a Nessus scan and do a version check, then look in a database and conclude our software was vulnerable. To shut up the scanners we changed the banners but still people would do fingerprinting, at which point we started putting messages like X-custom-build into our banners and explained to pentesters that they need to actually pentest to verify vulns rather than fingerprinting and doing vuln db lookups.
Point being, at some point you need to maintain stuff and have stable APIs if you want long lasting code that runs well and addresses known vulns. You don't do that by constantly changing your dependencies, you do it by removing complexity, assigning long terms owners, and spending money to maintain your dependencies.
So either you pay the library vendor to make LTS versions, or you pay in house staff to do that, or you push the risk onto the customer.
You lost me there already. Why should there be missing libs, and why would you not have to maintain updates regularly in production environment?
So let me see if I got this right, it's basically:
1. ./configure 2. make 3. ??? 4. profit!
Doesn't sound like a perfectly good solution to me.
Refactoring your application so that it can be cloned and built and ran within 2-3 keypresses is something that should be strongly considered. For us, these are the steps required to stand up an entirely new stack from source:
0. Create new Windows Server VM, and install git + .NET Core SDK.
1. Clone our repository's main branch.
2. Run dotnet build to produce a Self-Contained Deployment
3. Run the application with --console argument or install as a service.
This is literally all that is required. The application will create & migrate its internal SQLite databases automatically. There is no other software or 3rd party services which must be set up as a prerequisite. Development experience is the same, you just attach debugger via VS rather than start console or service.
We also role play putting certain types of operational intelligence into our software. We ask questions like "Can our application understand its environment regarding XYZ and respond automatically?"
I built a service that is installed in 10 lines that could be ran through a makefile, but I assume specific versions of each library of the system and don’t intend to test against the hundreds of possible system dependencies combinations or assume it will surely be compatible anyway.
The dev running the container won’t building their own debian installs with the specific version required in my doc just to run the install script from there, they just instanciate the container and run with it.
How big should the one server that serves the whole of netflix be..?
Apparently netflix has that many customers. Then again, if you split Netflix into regions and separate all the account logic from the streaming, the recommendations engine and the movie-content, you could perhaps run the account logic for one region in one server.
What is the point of running one server?
Do you also object to them running in the cloud in VMs and not on physical hardware that they own? Sounds like an old man's "kids these days" rant..
The link you shared just says they manage their OS layer, ofcourse they do. Everyone running on AWS VMs is responsible for their own OS layer. Wether they want precise control over their OS doesn't change their preference for who owns and manages the hardware..
You can pretend to be Netflix or Google, and build your tech-stack like they do. Or you can stop wasting your resources setting up a tech stack that you’re never going to get a return of investment on.
Stack Overflow is not a unit of measurement that anyone would be able to take seriously or find useful? How many stack overflows is one asana? Or how many stack overflows is one trello?
Horizontal scaling, docker K8s have their own benefits that are many and obviously to the industry. you don't need to be google to deploy and use them. If you deploy one server for each app and each team vs deploying a common K8s cluster where is the higher investment? You claim more ROI with more physical hardware and more servers?
I shall not express any opinion on this topic other than to say that Trello is not a good example to bring up. The entire customer base of Trello is not using a single shared board and thus they could scale in any direction they wanted to maximize ROI.
Because it’s less dangerous and cheaper?
> Horizontal scaling, docker K8s have their own benefits that are many and obviously to the industry.
Which is why SO runs on more than one IIS...
You don’t need a tech-stack, that is apparently even too complex for google considering the article, to scale horizontally.
> If you deploy one server for each app and each team vs deploying a common K8s cluster where is the higher investment?
The investment comes from the complexity. We’ve seen numerous proofs of concepts in my country, and in my sector of work, where different IT departments spent one or two 2-5 full years worth of man hours trying to adopt a perfect devops tech-stach.
Maybe that’s because they were incompetent, you’re free and possible right to claim so, but that’s still professional teams expending real world resources and failing.
From a management perspective, and this is where I’m coming from much more than a technical perspective mind you, the most expensive resource you have is your employees. If software is so complex that I need one or two full time operators to run it, well, let’s just say I could run more than a million azure web apps, and have our regular Microsoft certified operators handle it.
> You claim more ROI with more physical hardware and more servers?
I haven’t owned my own iron since 2010. All our on-prem servers, and we still do have those, are virtual and running on rented iron.
I think we may be speaking past each other though. My point is financial and yours appear to be mostly technical. If you can set up and run your K8s without expending resources, then good for you, a lot of companies and organisations have proven to be unable to do that though, and in those cases, I think they would’ve been better off not doing it, until they needed to.
Ofcourse transitions can fail. People can think yea let's do this small thing and end up chewing off a much bigger problem than they thought they were getting into. But that problem is in the whole of tech. "Let's just use our present people and switch from all proprietary to all open source in 3 months.." yea, best of luck with that... You need a solid team and going all in on K8s is hard, you need technical talent and leadership to drive this.
Agreed, maybe it may not be for everyone. Benefits are both technical and financial, less compute resources used, more reliable deploys, more resilient services. The problems being solved by this are not trivial. There are tangible benefits. Is it a risk? Ofcourse it is. The risk is not in the technology, the risk is in the competence of the team deploying it. If it can't change and adapt, maybe a lot more fundamental things need to change in that organization than just deploying a new orchestration layer.
My only point is, this shift from dedicated servers to VMs and now to containers is a fundamental shift in how things are done. People can hate on it all they like, but it's a better way of doing things and everyone will catch-up eventually.
Maybe if your app handles < 10k concurrent connections. Otherwise it is the most cost efficient solution and exists because it solves the scaling problem in the best way as of today.
But...
Splitting a horizontally scalable workload across a dozen virtual servers that are barely larger than the smallest laptop you can get from Best Buy, you are just creating self-inflicted pain. Chances are the smallest box you can get from Dell can comfortably host your whole application.
The fact remains the odds of you needing to support more than 10K simultaneous connections are vanishingly small.
Even before. If you want low latency. And banks handle more than 10K concurrent every day.
Cost example: https://pt.slideshare.net/markmyers106/vertical-vs-horizonta...
Chances are very high that the problem domain you are working in does not.
Like the author said “1%” I think maybe 5% to 7%
The point being that masses of software is developer everyday on a cargo cult adoption of solution they do not require.
This is certainly true, but there is a possible benefit: standardization. Having a standard skillset allows employees greater flexibility since they can jump employers and still expect to be rapidly useful. Similarly, if your company uses a standard toolkit, there's going to be less training overhead for new hires. Now, the devil is in the details, and I'm inclined to agree that you'd be better off hiring someone that can think outside the box and keep the tooling simpler. But using the standard toolkit will work reasonably well across several orders of magnitude in scale.
I used to do UNIX integration work in the late 1990's early 2000's and containers weren't really a thing. So you had to make sure libs from one program didn't crap on another program. And developers had to be conscious of what dependencies they included in their code. Nowadays they don't have to care as much because of containers. Every program can have its own dependencies. Thereby solving the integration problem.
A better solution would be to actually integrate programs and their dependencies into working systems, but no one has time for that. Software bloat is fine. Computers are cheap and fast. And actually understanding what we're doing would be too expensive. So just wrap all your have finished crapware up in a giant black box and dump it on a server.
Separating the dependencies between programs allows you to test and release independently to allow incremental upgrades. IMO that is better.
What I'm not interested in, is this kind of walking uphill in the snow both ways:
> So you had to make sure libs from one program didn't crap on another program. And developers had to be conscious of what dependencies they included in their code.
ie. having to understand what everybody else is doing in order for my software to run properly. No thanks. That's not why I'm here.
I'll put the exact dependencies I want, in the versions which work best for my software, into a Docker image or whatever tool offers a similar level of isolation, and I'll be working on my code while everybody else spends their time fighting over the ABI compatibility of C system libraries.
But it's not the case in linux, or at least non-enterprise / ldap linux. Installing mysql/redis/elastic/dotnet core/etc on each machine may be different, and has different installer available.
With docker I just need to instruct them to install docker, setup docker compose and everything is handled via containerization.
Don't you just end up with a less hackjob version of a container when you do that?
Which we later replicated in Red-Hat with RPM.
Bare bones OS install + bunch of OS packages => done.
And in what concerns containers I was working with HP-UX Vault in 1999.
You need to start with the same base operating system, and you need to make sure that you pull in the same versions of packages in case there are backwards-incompatible bugs, or version bumps such that the dynamically-loaded library is no longer detected (hence the common albeit dangerous workaround of "add a symlink").
(If you're using rpm directly, then you need to bundle the actual packages that you're installing, or point to specific packages that you're confident won't change. And at that point, what's the difference between your approach and Docker?)
The challenge that I believe Docker solves (or, at least, attempts to) is environment reproducibility: without it, you have dependency hell.
We use both together, since Nix is the only sane way to package Docker images, in my opinion.
And don't even get me started on having instances labeled "large" that have less memory and CPU capacity than my personal backup laptop (currently on loan to my 8yo for COVID reasons)...
Maybe. But when the time comes 512 MB doesn't seem like much anymore, what do you do? Do you pick the next larger instance or do you split the load across more 512 MB slices of a computer?
But you can use other cloud providers with better value.
At the risk of nitpicking, docker images aren't the equivalent of VM images, as they don't include a kernel.
Docker is not virtualization, it's just an abstraction that makes some Linux process isolation features easier to manage.
It also allows you to bundle whatever dependencies you have in the same bundle, but that is not the same as having a VM.
[1]: https://en.wikipedia.org/wiki/Operating_system-level_virtual...
Threads, processes and containers exist on a continuum.
See also cgroups: while this feature is used by the container run times, it predates Docker, and can be used standalone with normal processes.
If only I had realised that could have been useful for more than testing npm packages...
Both if implemented to spec will be logically equivalent and drop-in replacements for one another.
Mind you, that's still Docker, though, not Kubernetes.
> If one of your consultants or developers tells you they can make your application faster by adding a bunch of additional computers, you are almost certainly getting taken for a ride.
Eh. There's a redundancy play in there somewhere too, if you know how to pull it off. (Big if.)
That said, I haven't really used docker as a runtime container.
People have since bought into the marketing reasons for 'being in the cloud' and having 'infinite scalability' but that largely misses the point (and the pain) that caused many of these technologies and patterns to be developed in the first place.
The best example of how to scale without buying into this pattern I know of is Stack Overflow. At least circa 2016 - https://nickcraver.com/blog/2016/02/17/stack-overflow-the-ar...
There's your first miss.
I'm not interested in being "not lazy." I only care about user value and ability to provide user value (tech debt/cost/velocity).
For the price of an M5A.16xlarge to get those cores, I can get like 27 m4.larges and have way more fault tolerance.
I think your disdain is misplaced.
Imagine of your only unit of compute is a single bulky machine. You don't fully saturate it, but you need a second machine to avoid downtime anyway. Now you spin up a second or third service and suddenly you need 5 or ten machines and your compute utilization is 20%. You can pack things in tighter. But then you have a knapsack problem, and that's easier to solve efficiently with many small blocks even if it costs you a 1% overhead or whatever.
Cassandra is a big PITA, but it hasn't gone down (knock on wood) in the 6 years I've been using it. PARTS of it have...
How many distinct services do you have in your monorepo? A monorepo is fine...
until it isn't.
I've seen this (a long time ago) in the education world market. Very small school with a STEM program. They had specific scientific software they wanted undergrads to use (and some of it was pretty proprietary and used to interface with lab equipment) + a pre-configured IDE.
Instead of going through the compatibility matrix of OS and their versions they just gave all students a VM image that would "just work". Everyone could bring in their own devices and as long as you could run an hypervisor everything would "just work".
I don’t think it suits all teams and use cases, but for us it’s absolutely fantastic and without going down the rabbit-hole of cloud-provider specific tools and recreating half the issues it solves, I’m not super sure what we’d use.
I've already climbed most of the learning curve so YMMV, but as a team of one and dozens of WordPress, MySQL, and bespoke app servers, kuberenetes makes ops manageable so I can spend time on things that really matter.
Deploying new web apps is trivial, declarative manifests are easy to reason about, TLS certs are issued and renewed automatically (cert-manager), backups are cheap and reliable (daily GCP snapshots), making changes to the cluster via declarative terraform is a breeze, etc etc. No way I could manage all the ops without leaning so heavily on the core foundation provided by k8s.
I think that's the first thing with k8s: it all starts with an app that requires several physical nodes.
Well _technically_, sure, we could have run a bunch of those products on a single machine, but there goes your durability and the memory overhead on some of them was quite Hugh, and properly fitting them onto a single machine would have required more optimisation and technical skills than the devs I was working with had or were inclined to do.
Most of the value I get from k8s is the hands-off nature of it - I get slack notifications (prometheus+alertmanager) if anything is happening I need to address (e.g. workload down, node down, API not responding, etc). Otherwise I can safely ignore my cluster and know everything's good. Spinning up a new WP site takes 10m with backups, TLS, monitoring, etc built in.
You are absolutely spot on because this is how not to pass the behavioral interview for Engineering Manager.
After a week of almost full-time work, I threw in the towel. Admittedly, I also had to learn concepts like reverse proxies alongside, too, so I was by no means well-equipped to begin with.
Yet, tossing together some docker-compose.yml files and "managing" them with a Python script has worked very well. Kubernetes really scarred me in that sense, and I am healed! Also, Caddy has helped me in actually enjoying configuring the webserver.
I am talking about a homelab, a single server at home, for home use. It's much safer now, with Docker compose, because I understand it and I wrote the core exposed part's configuration, the Caddyfile, myself, manually. I know exactly what's exposed, and it's exactly right the way it is!
The remaining risk comes from the services themselves having security holes, but k8s has that very same risk.
Does mean that anything that upsets the ingress controller is an outage, but for experimentation, that's probably OK.
There are very strong financial incentives for every individual developer and sysadmin to adopt Kubernetes, regardless of the impact it has on the organisation as a whole. In a sense this is engineering reaching the level of corporate maturity of the sales department who will optimise everything for their commission regardless of the organisations ability to deliver it at a profit, or even at all.
Then that organization is doing a terrible job of aligning incentives. I'm guessing their pay structure isn't terribly merit-based nor high enough that people aren't constantly thinking about other jobs.
If this is about FAANG (your comment wasn't, but others were), perhaps part of this is exposing larger problems in many smaller orgs. (note: I'm ex-FAANG and happily so)
That issue is not endemic to Kubernetes, but rather to any larger system past a certain age, you learn stuff as you go along and would do stuff differently if you did it again today - but you can't easily, because you cannot break compatibility for everybody using your stuff.
As a concrete example from the Kubernetes world, there is a talk by Tim Hockin [1] about how today, they would fundamentally design the api-server differently and base pretty much everything on CRDs.
Also, I don't much like Go's templating syntax.
Not every addon/tool for the k8s ecosystem is worth it. I also don't bother with the ever-growing list of service meshes... not enough value to me for the overhead.
K8s is definitely the simpler alternative for me but there is still a lot of essential complexity in k8s due to the nature of the problems it's trying to solve. Mostly I like building on top of a solid foundation of standardized k8s API objects (pods, services, volumes, etc).
Tldr; Bring in only the add-ons and tools you really need so you don't add more complexity than necessary. Don't get swept up in the hype and marketing from other devs and cloud vendors.
Really when looking at tools in the k8s ecosystem, it's better to approach it as you would importing a new library into your application. Most decent devs wouldn't blindly import a new lib so that they can copy/paste a single line of code they found online for a business critical function, and k8s tools should be no different. We must think about what value does a given tool bring, and is it worth the cost of learning/maintenance? Sometimes the answer is a resounding "yes", but too often the question isn't even asked.
If you're looking to have your small app eventually grow into a large one, read up on K8s and just make sure you're not blocking future-you from making your app work on it. E.g., work well in a container (which is useful for automated testing, deps management, etc), have a simple 'ping' endpoint to make sure the app is up, have a better config story than "recompile to change these variables", use a logging library, and tolerate any other services you're using to sometimes be down.
All useful things for a grown-up app to do anyways, all a bit of a PITA, and all better than trying to operate an app that doesn't do them.
If you have one monolithic backend service (and most web applications really should start out this way), Kubernetes offers almost no benefits over alternatives.
Nomad is way simpler to get a cluster up and running, has a great configuration syntax (I'll take HCL over YAML anyday) and had first class Terraform/Consul/Vault integrations.
Onboarding devs is fairly straightforward, if they can write a docker-compose.yml, it's an easy transition to a nomad job specification.
It took me by myself ~4 months to get our current hashistack(Vault/Consul/Nomad) stood up using Terraform+ansible. Two members of my team have been working to replace the hashistack with a self hosted K8's deployment and they just went over the 1 year mark and we still do not have something capable of hosting the workload currently running on the Hashistack.
This got a little long winded but I feel like this "it's docker compose or K8's, take your pick" mentality had led to a bunch of needless time being spent by smaller teams/companies on solutions that just aren't right for them.
This feature helps a lot with that problem by bringing GCP closer to where AWS has been with Fargate. k8s will still be more work than using AWS ECS but it might also be preferable if you dislike using the provider’s components and want the control of, for example, doing your own load balancing and storage management.
Also you can get red/green deployments and rolling deployments with little to no effort, which can be very nice, nice to have.
Also why Docker in the first place? I’m genuinely wondering - in the stacks I run (Express / Python) it doesn’t seem necessary at low scale. Elastic Beanstalk, Heroku, Digital Ocean etc all offer facilities for single-command deploys that work out of the box.
Docker make it easy to run the same version of code in different places and let’s things run next to each other without version conflicts.
Also, I think you’re in a very small minority not to care about $720/yr increases in your hobbies.
I replied that beyond a single instance, you can probably get away with not hiring a K8s devops person and just spinning another instance. I'm not sure you've read this whole thing right.
And yes, I certainly wouldn't mind paying an additional $720 / yr for a project that had revenue; I almost certainly wouldn't want to spend money hiring a specialist, or spend time hyperoptimizing that myself - I make that in about a dozen hours of work, so counting how far one can go down the rabbit hole of optimizing server costs, and the associated cost of opportunity, the economics are crystal clear.
I don't have any successful personal projects but I have significant experience working with clients, and they are sold on the reasoning pretty much every single time ("I can charge you $3,000 for developing this feature, or we can use a paid service for $720 a year").
I also don't see how Docker is going to save you that much money; if you need a certain amount of compute, you need a certain amount of compute. AWS ElasticBeanstalk for instance charges nothing for spinning up an additional instance compared to EC2; there is no overhead for the PaaS aspect of it, like there would be in Heroku. Digital Ocean app platform is the same as EB, AFAIK.
That makes me think they have multiple services of variable workload packed onto a single host, eg web server, async, and DB all on a single host via Docker.
That’s the antithesis of EB, which can only do horizontal scaling. Docker provides a way to replicate those multiple services in a deployment configuration when you want to set up a host image.
Not using Docker as an easy way to pack it all onto a single box as long as they can is just wasted expense.
I previously avoided Docker Swarm for ages since I assumed it involved the same level of complexity as k8s. I also initially figured that managed k8s would be a safer bet than managing my own Swarm cluster, but if you’ve used anybody’s managed k8s (or read https://k8s.af), you’ll realize that every cloud provider has their own closed source fork of k8s with plenty of nasty bugs that you can’t do anything about.
My only complaint with Swarm is there isn't an easy way to expose containers directly on the network (like host networking). I have a few containers (wireguard and minidlna) which need this, so those are running through docker-compose. I've tried macvlan but wasn't able to get that working in Swarm mode.
networks:
- host
https://docs.docker.com/network/host/Ideally I'd like to give each service its own IP on the network, which was possible with how I had k8s setup.
Asking because I never saw the point in multiple container replicas for simple self-hosted stuff. One container each has served me well so far [0] (Nextcloud, Bitwarden, GitLab), and if they crash, they just get restarted. Multiple containers increase throughput, supporting more users, is that it? It just sounds nightmarish in regards to storage and parallel, conflicting writes.
[0]: One container per component (web, db, cache, ...).
Hashicorp's Nomad is the best choice on the complexity for features scale IMHO, and that's why i'm writing an article how great it is, how easier some things are and what's missing compared to Kubernetes.
Hashicorp's Nomad is good if you have a strong engineering department or need to run mixed workloads (e.g. both containers and native processes) because its abstractions are well suited for this, but HCL, their DSL for describing deployments, doesn't map nicely to neither Docker, nor Docker Compose files, knowledgebases or tutorials. Nomad's integration with Consul is a major boon, but the need to run your own CA for safe communication between nodes, Nomad's read-only Web UI, and the oddity of HCL at times also makes it a non starter for me and some other people.
At the end of the day, these are just two data points, sadly the job market for Kubernetes also dwarfs everything else and sadly many companies will be burned by this and will learn nothing at the end of the day. Ideally, i think that the best route would be evaluating the orchestrators and other technologies that you want to use by doing pilot projects and such, and looking at them in real world circumstances, to determine their fit for your goals and needs (Web UIs will matter for some, but not for others, for example; as will onboarding and the need for long term investment vs plug and play).
Edit: as for Kubernetes, personally i find the K3s distribution to be an almost reasonable alternative to Swarm/Nomad, if the situation calls for it: https://k3s.io/
Edit #2: it would actually be pretty awesome to read more about your experience in this article that you're writing!
> Swarmpit and Podman
I actually meant Swarmpit ( https://swarmpit.io/ ) and Portainer ( https://www.portainer.io/ ).
Podman is another container runtime that acts as an alternative to Docker (even if it is not feature complete), so i misspoke.
Nomad's UI has a good number of functionalities that can be controlled through it. Sure, there are some more lower level operations that are CLI only(though this seems to be something they are actively working to improve on) but most of that probably won't be needed but someone just trying to run a couple containers on a single node.
In places with 10+ developers and 10+ services on a single cluster it works surprisingly well from the user (developer) side.
Definitely a learning curve but honestly not bad if you like ops. Absolutely possible and beneficial for any size team. As always, it depends on what your goals are and how you want to use k8s.
If you were self-hosting k8s, then I'd agree with you.
I like Cloud Run for the same reason because I can use it without needing a lot of devops skills in my team or without sacrificing my own time (because I have those skills but have more valuable things to do). It allows me to focus on keeping my CI/CD pipeline (cloud run sets that up with a button click) busy with new functionality. And our hosting cost are close to 0$ because we stay below the freemium layer until we actually need to scale.
Edit. corrected the typo 700->70
for these case, i still have gke cluster around.
(Still agree that Cloud Run isn't for everyone.)
There's (almost) no limit to what you can run in one cluster too, and Kubernetes namespaces can help to separate different environments to allow for sharing.
Cloud Run sounds like the perfect solution for your workloads though!
This way:
- I can easily move from dev to prod by using the private container registry
- I have apt-get on the server and just use the default Ubuntu
- No distributed/network state
I used to use Core OS, but all these container OSes are here today and gone tomorrow, and normally have their own config standard. At least with Ubuntu they have LTS and a bunch of Google pages for fixing stuff.
For long I was scared of it because so many people say it's crazy complex. But actually it took no more than a dozen of hours for me to learn it and get a working setup on aws.
Maybe being full stack and having a strong knowledge of Linux and Docker helped.
Now I'm not pretending to be an expert with it, and there are certainly traps and mistakes that I didn't experience yet. But I don't understand what people find to be hard about it.
So, not simple.
Google did not say "Kubernetes is too complex" but rather, they are making this new tool - called Autopilot - that is an abstraction layer on top of Kubernetes for certain types of applications / companies.
This Autopilot system still uses Kubernetes AFAICT.
But it is a surprisingly negative headline. Autopilot sounds like a cool tool to simplify container orchestration!
That's the Register's schtick, they're snarky about everything.
GKE Autopilot is still fundamentally Kubernetes and supports the k8s APIs. GKE has simplified a lot of Kubernetes tasks that can be a pain, such as upgrades & scaling, but as many people said on other threads Kubernetes definitely is not for all workloads. For workloads that actually do need to scale then Autopilot removes a lot of the time consuming tasks required to get Kubernetes running and staying running. There are restrictions so check the docs on if it would work for how you use k8s.
Disclosure: I run product for GKE at Google so I'm definitely not a neutral voice on this...
Interested in having someone who has had to set up 1000s of containers in k8s and hated all of it?
K8s is too complex, and I say it as k8s dev(ops) who dreams the next great thing will come soon which will save me from piles of YAMLs and will bring back programming fun ;) /s
it's just that a lot of people think they have to use k8s because it's trendy even it doesn't fit at all their needs and scale. Then they whine it's too complex.
classic PEBCAK issue.
The problem for me is that these project are not under full time development. Most of them see maybe one or two deployments per year. It seems like every time I do a deployment, some part of the YAML config structure has been deprecated, with no clear path on how to migrate. In the end I have to spend a day figuring out how to rewrite the YAML configs to do _exactly the same_ as before. This is really hard to explain to a customer.
But it's not just K8S itself that suffers from this, third party components do this too. Stuff like ingress control is a nightmare, I've seen the nginx ingress controller (not to be confused with the 'other' nginx ingress controller) grow from a 50 LOC base config to multiple KLOC of YAML config you are supposed to pipe to your production environment. The ACME/LE extension (cert-manager) is even worse, the default config is 26KLOC long! [0]. And you are supposed to pipe this config straight into your environment. The amount of new and amazingly complex components that are added with each new release is staggering.
If you are an average Joe developer and you want to use containers in production, just stick with a machine with docker-compose on it. It's a lot easier to maintain and has far fewer surprises down the road. It is also much easier to make a quote on maintaining a VPS box that having to get a crystal ball to predict when Google will deprecate parts of your setup.
[0] https://github.com/jetstack/cert-manager/releases/download/v...
This is exactly what I am doing. It's easy to configure and almost same configuration between local test and production machines. The only downside is it's hard to monitor whether service is up without additional tools, which I lazy to research.
Additionally if you don't have a good enough ci cd yet you can deploy with volume-linked container.
Simple load balancers without auto-scaling work fine!
For smaller non-critical services sometimes even one reliable server is an acceptable trade-off.
Do most companies really need kubernetes given the complexity and difficulty it introduces?
But I grant you, it has been a very complex, painful journey to get this working right, and in fact I'm still making tweaks and adjustments. Plus, often times it's really unclear if I should scale vertically or horizontally.
TBO, it does not sound like it's working great for your use-case =) Wouldn't something like Cloud Run, AppEgine Standard or Heroku be much simpler and cheaper? Or just a single $5-15 per month* VM?
* depending on how much oomph you need for those 400 rps
I guess the ultimate goal of services like GKE autopilot is for you not to have to worry about kubernetes at all, just give them your workload and have it run on whatever resources they think is appropriate?
I do think it's important to recognise though that there are lots of ways to host services, and simpler with less abstractions is often better and more reliable and certainly easier to debug when things go wrong.
If I hear more reports like this I might just have to try out GKE :)
We have chosen k8s and I would again, because its nice to use. Its not necessarily easier, as you point out, the complexity of managing the cluster is considerable. But if you use a managed cluster like EKS or DO's k8s offering, you don't have to worry too much about the nodes and the unit of worry is the k8s config and then for deployment you can use Docker.
I like Docker, because its nice. Its nice to have the same setup locally as you have remotely.
In my experience the tooling around k8s is nice to manage declaratively, I never liked working with machines directly because even tools like Chef or Ansible feel very flimsy.
The other thing you can do is run on ECS or similar, but there the flexibility is a lot lower. So k8s for me offers the sweet spot of being able to do a lot quickly with a nice declarative interface.
I'd be interested to hear your take on how to best run a small cluster though.
For smaller setups (say 1-10 services) I'm quite happy with cloud config and one VM per process behind one load balancer per service. It's simple to set up, scale and reproduce. This setup doesn't autoscale, but I've never really felt the need. We use Go and deploy one static binary per service at work with minimal dependencies so docker has never been very interesting. We could redeploy almost all the services we run within minutes if required with no data loss, so that bit feels similar to K8s I imagine.
For even smaller companies (many services at many companies) a single reliable server per service is often fine - it depends of course on things like uptime requirements for that service but not everything is of critical importance and sometimes uptime can be higher with a single untouched service.
I think what I'd worry about with a k8s config which affects live deployments is that I could make a tweak which seemed reasonable in isolation but broke things in inscrutable ways - many outages at big companies seem to be related to config changes nowadays.
With a simpler setup there is less chance of bringing everything down with a config change, because things are relatively static after deploy.
how do you deploy your static binary to the server? (without much downtime ?)
Services behind a load balancer so one node at a time replaced then restarted behind that, and/or you can do graceful restarts. There are a few ways.
They're run as systemd units and of course could restart for other reasons (OS Update, crash, OOM, hardware swapped out by host) - haven't noticed any problems related to that or deploys and I imagine the story is the same for other methods of running services (e.g. docker). As there is a load balancer individual nodes going down for a short time doesn't matter much.
Ask yourself how would you solve this problem if you deployed by hand and automate that.
1. Create a brain-dead registry that gets information about what runs where (service name, ip address:port number, id, git commit, service state, last healthy_at). If you want to go crazy, do it 3x.
2. Have haproxy or nginx use the registry to build a communication map between services.
You are done.
For extra credit ( which is nearly cost free ) with 1. you now can build a brain-dead simple control plane by sticking an interface to 1 that lets someone/something toggle services automatically. For example, if you add percentage gauge to services, you can do hitless rolling deploys or cannery deploys.
Docker introduced a great level of abstraction and reproducibility over platforms. However, Docker (or OS-based containers) are the most atomic unit of computation on Kubernetes. Which causes centering scaling on the Instance, instead of the Application or even the functions.
This leads to a lot of unintended side-effects. Of which complexity is the most evident since now you need to handle scaling, monitoring and compliance over the VM layer rather than on the functions or the app.
I believe a VM centered on the app (WebAssembly VMs) or functions (serverless approach) is the right computation unit to allow proper scalability and a simpler and more powerful system on the long term.
Agreed, although at some point in a not very far feature most of those missing features will resolved. So in my mind is just a matter of time.
The Wasm Community group is doing an awesome work on that :)
The scaling of their primary business (amazon.com, google.com) is at least partially an infrastructure problem - so they had to solve this anyways. Why not try to sell it too?
But it's very rare that you can scale a system by only scaling infrastructure. Ironically, probably the only way this will work is if you scale vertically - which means you don't need K8 and want to avoid the cloud like the plague.
There are all types of application-specific consistency issues that need to be treated as first class parts of the system for horizontal scaling to work.
That should be why Red Hat took an early investment in Kubernetes. Operating systems did matter for them. Applications did not.
> I believe a VM centered on the app (WebAssembly VMs) or functions (serverless approach) is the right computation unit
I agree with that. The thing I worry about is a k8s variant like Krustlet, running Wasm instead of container, might introduce the same complexity into the Wasm world. We need a more app-centric scaling solution.
Docker's co-founder agrees:
"If WASM+WASI existed in 2008, we wouldn't have needed to created Docker. That's how important it is. Webassembly on the server is the future of computing."
Are we just making rabbit holes out of rabbit holes of abstraction using kubernetes?
I just use and push code to Heroku and I'm done for the day, simple. NoOps I call it.
I wish more tools and platforms were like this.
Can't afford to test my luck until I've finished migrating all my accounts off my gmail.
Honestly, creating a second google account might itself increase the chances of getting banned. Nobody knows with Google, and that's the problem.
I created a new google account with no relation to my main one and added it as an owner to my GCP project, so hopefully in the event of either account being banned I can still access everything.
Luckily, there are lots of other Google services for that that you can plugin for that sort of stuff. Cloudrun is great if you are planning to use those things.
Kubernetes is what you use when you want to mix stateful and stateless stuff so you can avoid depending on those services. That makes sense if you need to support multiple clouds or on premise installations. But otherwise, it's a lot of extra complexity and devops even before you consider the overhead of managing the kubernetes cluster. There are lots of companies that talk themselves into needing this where the need is arguably a bit aspirational. I've been on more than one expensive project where we served absolutely no traffic at all with hundreds of dollars worth of kubernetes clusters idling for months on end that had no realistic hopes of ever getting more than a very modest amount of traffic even if everything worked out as they planned.
Just to illustrate the point, dokku's docs[0] for logging say "Warning: The default docker-local scheduler will "store" these until the next deploy or until the old containers are garbage collected - whichever runs first. If you require the logs beyond this point in time, please ship the logs to a centralized log server."
Alright, well, that's gonna be a whole lot more work than two click in Heroku to send all my logs to any logging provider of my choice.
https://azure.microsoft.com/en-us/services/app-service/conta...
Edit: fixed name
When you start having a lot of pieces to manage eg. cache, database, auth, multiple applications then Kubernetes comes into its own.
Because then you can scale, monitor, trace, debug, log, backup, audit, encrypt and visualise all of those pieces in exactly the same way.
And do it irrespective of which cloud you use or whether it's even in the cloud at all.
I’m not telling people what to do, if you like K8s or Docker or what have you knock yourself out, and I mean it. For example, people keep telling other people on HN not to use React and that’s a hill I’d die on - and could write a dissertation defending it. I’m just wondering what the dev experiences of others are so I may learn from them.
Running a cluster across your IOT fleet (1k+ devices; 5-10 apps each) gives a nice interface for pushing out tasks, choosing applications for a device, configuring supervisor-device relationships, etc. It turns out from scratch Docker containers on ARM are very portable.
I think you’re being dramatic — embedded is notorious for crazy builds of weird config flags to even get “hello world” to compile, and you’re waving your arms pretending that concepts from Erlang abstracted away from the language are too much.
Docker is just cgroups and namespaces, with a zip file of code. More or less literally.
Orchestration is the same mess it’s always been — back to at least the telephone days, when Erlang used the same concepts.
Because everyone else is doing it.
Things have gotten so complex that we need complex solutions, but w're afraid to make and own the solution ourselves because of cost, focus on core business or recruitment/knowledge concerns. So we need to find the next best open source solution to leverage the collaborative effort to reduce time and cost and have a large community to fall back on for support. And because everyone jumps on board of the same train all use cases must/will be accounted for (else the solution will fade into a niche), meaning the solutions becomes a problem in and of itself.
Repeating the cycle once again.
Isn't this just labeling the knowledge that you don't have (and could probably read up on) as potentially unnecessary complexity?
I mean, every time I hear about embedded, I keep hearing about byte boundaries, RTOSes, compiler chains, musl, JTAG, and a million other things that make my mind numb. But I assume embedded engineers need to know all those things, because of the constraints unique to their field. Some of them could be just "a complex way for engineers to keep themselves spinning their wheels and not actually working on an application", but a lot of them provide value in the real world and solve some specific problem.
Cloud infrastructure management frameworks are the same.
I guess I was comparing it to how things were (or more precisely, how I perceived things) 20 years ago. You'd rack a server or two, install Linux, stick the application in /usr/local/bin, make sure Apache was set up, and you were off and running. Simple enough. It probably didn't scale to 2020-sized Internet user counts, though.
That sounds incredibly complex! :) Where would you find a place to rack a server? Are there minimum rates for such a place, or can you rack a single server for one month? Would you be able to rack it yourself? What if there was some problem with the PSU, and the server lost power when the data center was fine? What if the network card failed? What about if a hard drive in the server blew out?
Cloud infra in 2021 infrastructure has pretty turnkey answers to all of those (Kubernetes is not one for any of the ones I mentioned, except maybe for “what happens when any node goes down“). So yes, the complexity might be high, but so are the capabilities.
I mean, maybe you are running a mom-and-pop shop with a server rack in the basement for some reason, and have pretty daytime-specific hours. That's fine, and I still think you can do the server rack thing. Or at least that's the way I see it.
Perhaps one piece of meta-commentary here is that server hardware has not become substantially cheaper, faster, smaller, and easier to set up since 20 years ago. I can't really set up a server in my small bedroom that's capable of serving 10k concurrent users, and doesn't deafen me with its noise. Maybe if I did, all this cloud stuff might be less ubiquitous?
All the tools you mention are not integrated into your app neither change how you write code. In fact your application shouldn't even know what service mesh you are using.
I can write a back end application once and then have it run with istio/linkerd/traefik/whatever with zero code changes.
That doesn't happen in the Javascript world. The choice of framework directly affects your code.
I just enter my Honda Civic and drive to my office. Simple. I call it "Sedan".
(Honda Civic driver looking at an 18-wheeler in the highway and not understanding why somebody would use that)
In all seriousness though, Kubernetes solves a specific set of problems. Just because you personally don't have these problems, doesn't mean that Kubernetes is bad, or that people who use it, don't need it.
Lens has been a huge boon for helping us manage our kube cluster and regular devops operations, and even 15 minutes with it helped me grok a number of complex kubernetes concepts that I've struggled with for awhile now.
Everyone who works with k8s for a living should at least know of this tool imho its fantastic with prometheus
As for cloud tech: https://www.amazon.com/Mastering-Kubernetes-container-deploy... will cover 90% of it for you.
This is sad.. I don't understand why providers do it. It makes it very expensive to run small staging clusters. There is some reference on the current state here:
https://docs.google.com/spreadsheets/u/0/d/1yhkuBJBY2iO2Ax5F...
The pricing is... not cheap: https://cloud.google.com/kubernetes-engine/pricing
Comparable-ish with Fargate. You probably wouldn't be using this with an eye to save money unless you think (or have measured) that you'd spend more on operations, security or compliance otherwise.
Direct link to the docs: https://cloud.google.com/kubernetes-engine/docs/concepts/aut...
Google has a tool by the same name (autopilot) internally that does close to the same thing. There was a paper published on it 10 months ago talked about here: https://news.ycombinator.com/item?id=22980467
But likely they shared little in common.
Things are always fine when you're starting, but if you don't understand the system it's hard to troubleshoot when it breaks. Understanding Kubernetes is hard. Obviously there are ways to outsource cluster management, but we don't have the budget of a VC funded startup.
We ended up choosing Hashicorp's products -- Nomad for orchestration, Consul for service discovery, and Vault for secret management. Each one is just a binary with a <20 line config file, and it takes ~two days to read and understand the docs well enough that we have been able to troubleshoot issues quickly.
IMHO, most of what K8s offers can be obtained cheaper and in a simpler way using more "traditional" DevOps approaches and systems. I'm currently operating an IT infrastructure consisting of more than 20 different components, using only basic Linux technologies and open-source packages (ssh, iptables, ipsec, ferm, (r)syslog,...) plus Ansible to orchestrate it all. Never encountered a problem that I wasn't able to debug and fix within a few hours, and managed to have more than 99.99 % uptime so far. I understand that this approach might not work for large companies, but it seems to me a lot of startups are going down the K8s route just for the sake of it, and their DevOps processes become incredibly brittle and slow as a result.
1. Someone has an idea. It's alright. Really good for their use case. Someone else hears about it, likes it, and adapts it to a similar use case. So and so forth until the idea has a large user base
2. Employees of large companies hear about the idea and implement it
3. Marketing gets a hold of the idea, gives it a flashy name, and uses it in promotions
4. A majority of the loudest voices in the industry get on board so everyone starts forcing this idea on every imaginable use case
5. A lot of people realize this is all too much extra work without much benefit so they start looking for more appropriate solutions
6. return to 1
It's happened with mainframes, object oriented programming, services, micro-services, web-all-the-things, blockchain, steaming data (kafka), sql, nosql, containers, agile, VMs, cloud, and many other things.
It's just technology people are the worst at getting caught up today's fad and I wish we could try to not do that as much.
Most of the software tech world is bullshit. Most best practices are bullshit. And, if I were feeling a bit conspiratorial, I'd say the large tech companies do this on purpose. They promote fads that overburden any smaller company with more modest budgets, thus keeping competitors at an arm's length. Resume-driven-development plays a role in this as well.
A lot of companies following the FAANG cargo cult are digging their own grave and don't even know it. They don't realize they don't have the manpower for microservices. Or the calendar time to make it work and still get product out the door before their competitor that doesn't even unit test eats their entire lunch.
But it's no good to be big company when Whatsapp can be build by one person you missed to hire.
After a half-decade in the industry, I am beginning to realize there is a lot of bad engineering hiding behind marketing and emoji. So many of these web tools have nightmarish interfaces and add complexity that, for most of the industry, is unnecessary. But their docs pages are full of rocket ships and confetti.... I swear, seeing a page with rocket ship emoji has become such a turnoff. Why do we need our engineering handed to us sprinkled with decorations like a cupcake?
And the imposter syndrome that the entire industry seems to share dictates that a large percentage of developers feel like they need to be using these tools as a badge to wear saying, "Yes, I am with it and hireable."
Meanwhile, the technology itself keeps churning because there is now a profession full of people believing that creating and open sourcing the Next Big Thing in ____ Technology is the best way to move their career forward. People keep reinventing wheels because everyone is focused on making a name for themselves with the new; no one notices the person doing mundane maintenance on Rails or whatever.
I think it's been this way since the 70s/80s to some extent, after having done some reading on the history of the profession. I think it's just scaled with the number of programmers and the Internet has applied its intensification effects.
1. Build an opinionated solution that you control fully (e.g. difficult to fork).
2. Convince everybody to adopt it. Nobody gets fired for choosing $bigcorp
3. Grow it complex and expensive to maintain over time.
4. SELL software and services to manage its complexity.
...and the industry is getting more marketing/fad/resume-driven every day.
It makes a lot of sense to build portable infrastructure, but you can scale a long ways with much simpler technologies.
Way too complex for my tastes, but if you're convinced you need a highly automated control plane for those things then pushing .deb or .rpm packages doesn't even come close to solving that problem.
As usual it comes down to how you define the problem. You can't dissuade people from using k8s by comparing and contrasting k8s to alternatives; you've already ceded the debate over how to define the problem at that point. W'ever its relative merits, k8s is a reasonable approach to the problem of automating the control plane for "scaleable" services.
Not saying it is the answer to everything, but it can do a whole lot more than software. Most of our problems are self made we can unmake them by changing the rules we operate under.
> As usual it comes down to how you define the problem.
Totally.
Redefine the problem until the solution is tractable and simple. K8s is the problem to a solution.
It's an insanely complex solution to a very niche problem - scaling stateless web app backend nodes written in scripting languages.
Stray even a little bit off the garden path and you start feeling pain.
K8s sounds like a good idea on paper - you get reproducibility and resilence "for free" - but then I found out I have to effectively roll my own everything with k8s anyways and went with Jenkins instead.
K8s solves a far wider problem space. Need to run a data store? Use a StatefulSet and PersistentVolumes. Need an occasional task? Jobs and CronJobs. Need to know what's happening? Metrics and logs have APIs. Load balancing? Ingress? Firewalls? Security?
If someone knew nothing of operating systems other than a class on MINIX, I suspect running a massive datacenter would be easier with k8s than running a medium system of debian boxes.
- App Engine has a bunch of weird limitations and slow deploy times. Qualitatively, it feels like the spotlight has moved on.
- Running my own compute instances felt like reinventing Kubernetes, especially once you roll your own deploy mechanism and throw load balancing in the mix. I also don't buy that it's easier.
- Cloud run is promising but for database heavy apps it's a non-starter.
- GKE was pretty smooth. It feels like it gets a lot more love than App Engine. The UI was functional with lots of depth. Once I push a docker image, GKE updates the nodes to serve the latest version. Load balancing was a matter of ~3 yaml files at 10 lines a pop.
There are a lot of businesses that will never have to deal with dynamic scaling, engineering autoscaling into those solutions is pointless. If you need that kind of scaling then K8r is fantastic. My point is a lot of people turn to K8s well before they need to or without understanding why they might need it.
Containers add vast complexity, add another layer of complexity on top.
So far that's Docker, or containers to put it more generally.
Now if your web page is so amazing that it receives a lot of traffic, your little container is going to get overwhelmed. And if it falls over, then it's dead and nobody can see your web page until you bring the container back up. Fortunately there are tools that let you manage this aspect of the image, called orchestration. You can tell orchestration tools how to figure out if an image is unhealthy and needs replacement, and if it falls over whether to bring it back, and importantly, how many copies of the image to run to handle the traffic. And if you need to push an updated image with an updated web page in it, how to gently make that new container available to the world without interrupting the traffic.
There's more to orchestration, you also tell containers how to talk to each other if needed, how to manage secrets, encryption, load balancing. There are lots of aspects of hosting that fall into this.
The two main orchestration tools I know of are Docker Swarm and Kubernetes. Docker Swarm is bundled with Docker already. It's pretty easy to shift from normal Docker use to Docker Swarm use, it works well enough for small-medium deployments. Kubernetes is a tool for much larger and highly flexible use cases, and it has a lot of levers and buttons and swiss army knives with its own swiss army knives. Many aspects of Kubernetes like the load balancing and secrets are all pluggable and you can use different tools in there.
Now you're at this article's topic, which is Kubernetes. K8s as it's called has a larger mindshare of the ops world, therefore everyone wants to use it, but it's very complicated, so a tool has been introduced to try to simplify it.
Kubernetes basically lets you define your references not as "c:\jon\reports\fy2020_final_final_2_comments_review_Bob_final.xlsx", but as "fy-report", with "fy-report" being defined elsewhere.
This is required to help you run the same program in different circumstances without breaking everything. You can say "run it on this slow computer with this test data", or you can "run it on many big computers with real data", but the program is exactly the same.
What makes it so complicated is that Kubernetes tries to abstract everything a given program would need, so it's not just references to external files, but practically the whole computer with all its network connections that must be defined elsewhere using Kubernetes' special language.
A lot of this special language is the same for 99% of programs, like in Excel you want your VLOOKUP to work on a $-pinned range with the last parameter set to FALSE 99% of the time. This makes people make the same stupid mistakes and finding them is hard.
And of course, this special language means you have to relearn a lot you know about running programs on computers, like when you move from Excel formulas to VBA.
This was years ago, so maybe they greatly simplified things. But somehow I doubt it =/
[1] " Docker, by default, punches massive holes through your firewall in non-obvious ways. People don't realize that with a default Docker configuration, containers are ignoring any normal firewall rules you may have setup with iptables or ufw." - https://news.ycombinator.com/item?id=25834444
* Built the app (into a self container .jar, it was a JVM shop)
* Put the app into a Ubuntu Docker image. This step was arguably unnecessary, but the same way Maven is used to isolate JVM dependencies ("it works on my machine"), the purpose of the Docker image was to isolate dependencies on the OS environment.
* Put the Docker image onto an AWS .ami that only had Docker on it, and the sole purpose of which was to run the Docker image.
* Combined the AWS .ami with an appropriately sized EC2.
* Spun up the EC2s and flipped the AWS ELBs to point to the new ones, blue green style.
The beauty of this was the stupidly simple process and complete isolation of all the apps. No cluster that ran multiple diverse CPU and memory requirement apps simultaneously. No K8s complexity. Still had all the horizontal scaling benefits etc.
At work, we have had multiple big and bigger incidents inflicted by the "share everything" nature of kubernetes. Sadly, it never occurred to the devops team that something like this is possible. Same goes for the HN crowds here who just assume that a container scheduler is a must, be it kubernetes or nomad or docker swarm.
I blame this on Google Cloud brainwashing. Let's see how long it will take for Thomas Kurian to sunset https://cloud.google.com/compute/docs/containers/deploying-c...
Only thing that's somewhat good is Go lang.
If you're going from "our app is running directly in tomcat in our on-prem data center" to "we want to move to a mirco-services architecture with a service mesh, canary deployments, multi-cluster observability, fault injection, automatic global failover, etc...", you're almost certainly going to have a bad time.
I really like the concept of "innovation tokens" [0] here. Pick one or two big innovations at a time and don't add more until you're comfortable with what you have. Get your app running in docker first, and then get it running in a basic k8s setup first instead of leapfrogging to the more advanced stuff. Chances are if you didn't need a service mesh for your original non-k8s deployment, you probably don't need one for the v1.0 first pass of your k8s rollout.
We have not had any major k8s incidents in the past 12+ months of running our own HA cluster. We even run HA Postgres using local volumes. We've pretty much moved everything over and couldn't be happier.
We have red teamed disaster scenarios, we can bring the full cluster up on another cloud provider using vanilla k8s in about an hour (including a restore from Postgres S3 backups + wals). And this is all with a very small team.
The "declarative" part of Kubernetes control buried in YAML seems to fail to live up to that label.
Is there any literature that shows that the cost savings from packing more than one service per VM are significant enough to outweigh the cost of the messes created by middling developers who have trouble reasoning about the functioning of single process application at non-Google companies create when they think they should do what Google does?
Imho it was very simple to understand all the components. But I can't deny that it does require a solid understanding of sysadmin concepts and containers to get ahead.
I was quite proud to have grasped kubernetes within a few months, without any formal training. No finished education in my country and not from an English speaking country. So I often struggle with technical docs that are using too much academic english.
Still today when new issues arise I can relatively quickly understand why they're happening.
In some cases though you must read the docs carefully.
Im still wondered how people could give a green light to something that makes their life _harder_ not _better_ just because “it’s Google”.
My guess is — you can make more money on something as overcomplicated and non-developer friendly as k8s.
Imagine how many companies would never even exist if k8s was like a bit advanced version of swarm.
How many talented developers would be free to make something valuable for the world.
Instead of just solving stupid problems created by another developers with questionable design choices.
Honestly, I'm not sure Kubernetes is it. K8s seems like a leaky implementation detail when what people want is just to virtualize a workload (with similar simplicity to Docker) and have it work wherever the standard 'workload virtualiser' runs.
Perhaps the enterprise class stuff is always just for large division of labor efforts. Like the "Liberty" stuff from WebSphere perhaps this really just saying if smaller groups and less people want to mess with it, it should have some new thought around more consolidated abstractions.
I enjoy using it and playing with it, but so many use cases can be addressed with something simpler - either just Docker / Swarm / AWS ECS etc. alternatives or just going for VMs with well defined CI/CD processes that let you tear down the infrastructure and set it up again easily.
What very interests me are the concepts K8s build on that are not usually recognized - to me K8s seems a lot like a JVM, just that it operates on infrastructural (and not runtime) level.
I enjoy experimenting with these concepts when applied back in the runtime world - it is for instance interesting to run 100s of servers with JVM and let them load/execute new dependencies and code at runtime (JVM is very well suited for that).
This is area that is not yet explored and would probably deserve more attention as it allows for distributed rapid computing that is infrastructure/platform independent (the downside is that it requires (just) JVM and the isolation is not perfect).
Now, can you get along wit K8s? Of course you can! Human beings have been adapting to harsh environments for hundreds of thousands of years. We can figure out some shitty software. So that means that probably, K8s will not actually improve, because people will probably adapt to it rather than change it.
So I think it's time to create a new distributed computing platform. Something with a simpler design, that has the necessary functionality built in rather than bolted on. Something that is very powerful but also terse. And of course, something that only gets complicated when you need to do something complicated, and allows you stay simple as long as possible. The ability to scale the complexity, essentially.
I'm learning Kubernetes and is deploying my own test cluster on my ARM-based board at this very moment, and I already spend 3 days on K3s and have to give up due to a problem (https://github.com/k3s-io/k3s/issues/2509#issuecomment-78657...). I must say, this is way way harder than Docker Swarm.
Just want to know, will Google actually make Kubernetes a bit more friendlier for general non-GKE users?
Google actually works a huge amount with the community to simplify Kubernetes across a number of SIGs. Its always a trade of increased flexibility and options as people use it for more workloads, versus simplicity.
Autopilot is just for GKE. You can use GKE on other clouds and also onprem (bare metal or VMs with Anthos)
However, I found myself spend too much time on finding out what settings should be put into which files. I really hope there is a tool that can help users to generate and configure these settings.
I think a good example of this is `npm config edit`. When invoked, it'll open the correct config file, and list all available settings with their default values. User can then enable those settings by uncomment them as needed.
Maybe add similar functionalities into those CLI tools (kubeadm, kubectl etc) could greatly improve user-friendliness?
What I miss the most is having my infrastructure defined as code, instead of via the GUI. But given that I have only four services (out of which two use preemptible VMs and only one needs to scale) it’s not really a problem — it wouldn’t take me many minutes to replicate this setup at another cloud provider.
Dedicated servers are heavily under utilized and over provisioned because once you allot a server to a team they don't want to give it back.
VM solved this problem and changed how servers were provisioned. Docker and K8s are the next progression of this. People who compare K8s to the 'next js framework' have to do some serious context alignment...
I really like "Google Cloud Run" in this sense. Just deploy your Dockerfile.. And you're done.
A single yaml file. Limited configuration. Scaling just works. No cluster management.
I just want to build a great product.
Is a tool that will allow the configuration of base images. It is actually a thin layer wrapper around packer and ansible
Of course it's complex, it's managing a complex and very deep problem space through a mostly consistent generalized interface. Whether it is better than the alternatives depends on how deep your problems in that space go.
They should remove the need for CONFIGURING ssh.
Now they have removed the entire control plane of the container. Hiw should a developer debug something then?
You can still shell into your own containers to debug workloads, we removed SSH access to the nodes in order to provide management of the nodes (we can't take on the management if people can SSH in and make unsupported changes).
If I enter the following, it seems to work for me: Replicas: 1, CPU: 0.25, Memory: 500MiB, Ephemeral Storage: 1 GiB.
The result is $9.92 per month for the us-central-1 location.
The issue here is we don't have a kernel and compiler for distributed applications. So instead we have no choice but to expose that complexity to developers. They need to specify the IP address and port of a remote service to invoke. We've solved the IP problem in part with DNS, but now what happens when you want to migrate or load balance? Kubernetes solves this with service discovery, so you just name a service and the container orchestration engine worries about resolving that to a pod at runtime. But you still need to specify a port. We haven't yet figured out how to automate allocating ports like we managed to automate register allocation on a CPU, especially when application code itself might need to know them. We still require you to know how much storage you need and specify that in the definition of a persistent volume. Ideally, it would be as easy as it is to ask for an array of integers in a programming language. The compiler will figure out how much storage an application requires and give you that much from some pre-allocated pool you as a developer don't have to worry about.
But again, there is no such pre-allocated storage pool. There is no such compiler. There is no POSIX filesystem or memory standard for multi-machine networked systems. There's a fragmented system of vendor-locked services providing storage servers, database servers, cache servers, http servers, message queues, some more open than others. Kubernetes is an attempt to provide abstractions that make it possible to define an entire network of such individual servers declaratively and it's a noble effort. But it's complex because the underlying problem space is complex. Distributed computing is at the point in its evolution right now that single-machine computing was in around 1950 or so, when you needed to tell the program exactly where in memory to find and store a variable, exactly where on disk to fetch a block of bytes. Will it ever get to where we are now with device drivers, compilers, and kernel allocators and schedulers doing all the heavy lifting for you? This is what a Kubernetes engine is trying to be, but it's early in the game. Very early. I don't see the point in writing an article implicitly shaming the developers for trying to provide higher level abstractions that make the simple cases easier, any more than criticizing a kernel developer in 1960 for inventing virtual memory as if they're admitting that symbol to address resolution in a multi-processing system is too complex. Of course it's too complex! And we're trying to make it less complex.
We are currently on AWS ECS which works very well for us but need to move to a GDPR/Privacy Shield compliant environment hosted in the EU. I was to look into Managed K8s but am turned off by this thread :-)