Unironically using Kubernetes for my personal blog
mbuffett.com
mbuffett.com
What's neat about it is, you generally solve one problem, and once it's solved, it stays solved. Your solution is in a yaml file somewhere or in a command line option that you've persisted to a script or whatever. And everything accumulates! If you manage to get a storage cluster installed, boom, all of a sudden you have a ton of options available to you.
Once the cluster is stable and you are comfortable bringing it up and tearing it down so that everything's working right, you can start pulling parts of the k8s infrastructure in, need a DNS server? Run it on your cluster! K8s wants a docker container registry, you really don't want to run that on your cluster in the beginning, but once your cluster's secure, why not! K8s starts eating everything in your life like the effing Borg and it's great!
If you're the kinda cat that will spend 80 hours on a Factorio world, then dive into the mods and make the tweaks you really wish the authors would make, and you have enough admin experience to do general troubleshooting of complex systems, I can't recommend Kubernetes enough for the sheer fun factor.
The downside if you want to do anything serious with it is, actually the same as running your own infrastructure always is and was, network connectivity. Hardware is cheap. Last mile network connectivity, isn't. The solution to that, is, of course, colocation.
Can you expand on how this is different from traditional networking? I used to run a FreeBSD router with a ZFS storage setup and various other stuff on my home network, and it wasn't like I had to keep tinkering with it. Once I had network storage, I could use it from lots of different things, etc. But eventually you need software patches, updates, etc, and that was where the ongoing pain would sometimes crop up. Is that so different here? From some coworkers who are closer to the cluster, keeping up with K8s version changes doesn't seem like a small effort.
These files can idempotently be re-applied to a cluster to restore, or upgrade, the running applications.
When you lose your cluster and you restore to a new cluster, you will of course initially deploy the same version of kubernetes as you were using previously.
Then once you choose to upgrade to a new version of kubernetes, it can happen that some of the APIs you were using have now been deprecated/removed. This is a very slow process, so you had plenty of time to upgrade ahead of time before being forced to.
But let's say you ignore all that, and have chosen to upgrade to a new kubernetes major version and are now forced to upgrade all your yamls. This happens rarely, and recently only because a few BETA APIs have become STABLE and people are now expected to use the STABLE versions. So you go ahead and make those few changes, re-run your deploy and you're done.
apiVersion: apps/Betav1
kind: Deployment
to: apiVersion: apps/v1
kind: Deployment
Because in K8s 1.16, they stopped the backwards compatibility for Beta on deployments. But that was quite some time they left it in for backwards compatibility, and it was a simple sed script to change it. We have never changed our Dockerfiles except to change the FROM line when we update a base image. They often introduce new features, but try to keep things as stable as possible for existing setups.In our case, we store our YAML in git, and can deploy a new cluster with all our microservices in about 10 min (most of that is delays in the google global load-balancer setting up a TLS certificate for us, the machines are up and running after just a few minutes).
There can be a bit more scripting needed to do updates of the nodes to newer builds with no downtime, if your pods aren't totally stateless, but compared to what it used to be at an old job, with C apps running on Linux, behind load-balancers that had to be removed, upgraded, and added, a rolling update is like black magic.
This is exactly the bit I care about: what do I do with "stateful" systems (my db storage, mail server data,...)?
If I just mount those from external volumes, I lose a lot of idempotency and I now have to worry about whether I am trying to attach that volume to PG 9.3.1 or 9.3.2 or 12.0 (some are fine, others might cause data corruption).
I know idempotent deploys are all the rage, but all those deployed apps are there to serve some data which is as stateful as it can be.
I'd like a system where those external volumes are automatic snapshots on LVM (or ZFS/btrfs) when attached, but that introduces a whole another level of what-now if you need to go back to an older, now slightly stale data set.
How does k8s solve this problem for me?
Kubernetes is still early tech and thus has more moving parts than a more well established stack. Keeping up to date is crucial and no small feat.
1. Get a new server
2. Connect ethernet and IPMI ports
3. Add node definition to XCAT
4. Set boot to network, arm node for installation, power-on
5. Profit.
6. Get a coffee.
It's 15 minutes for [1-150] servers. Slightly longer for 150+. Because, network.Any further setup is done via Salt, if necessary. With a single state file.
Fun fact: I provision K8S nodes that way too.
Not OP, but I can chime in here.
It's not that it's solved in this specific deployment, but it's solved in yaml files that you can re-deploy to any environment. Once I get my bundle of yaml files, I can just point it to any cluster (or namespace) et voila, it'll be deployed. It's "infrastructure-as-code"[0]. To be clear, Kubernetes doesn't have a monopoly on this approach, but it certainly follows it.
Another idea is treating your servers like cattle, not pet. I used to have Linux VMs in my house named after Greek myths – these are pets. I named and loved each of 'em. And yes, after being setup, they worked. In Kubernetes, they're cattle – and you kill your cattle. You don't name and grow to love and nurture a specific VM. You can wipe your deployments and reprovsion a new one in a blink.
>I used to run a FreeBSD router with a ZFS storage setup and various other stuff on my home network
>From some coworkers who are closer to the cluster, keeping up with K8s version changes doesn't seem like a small effort.
You're not wrong here though. Sometimes there are maintenance tasks you have to do on the whole cluster (make sure there's enough storage, everything is updated, certs are good, etc). But this is orthogonal to having your infrastructure-as-code in yaml files. The yaml files assume a healthy cluster/namespace, but besides that any cluster should be fungible. This is treating your servers like cattle, instead of pets.
[0] https://docs.microsoft.com/en-us/azure/devops/learn/what-is-...
Very poetic, I love it! These kind of internet conventional wisdoms are gold, I also like:
Batteries included but swappable
Free as in Beer vs. free as in speech
Are there other contenders?
Maybe, if you're lucky, you had some set of really good and reliable Sys Admins that figured out a robust way to script and configure the setup process of your original on-premise data center and they captured that in very good, well-maintained documentation. If you're even luckier, maybe those same guys still work for you.
If not, well, Kubernetes provides an open standard that a lot of different people know how to use and the documentation for how everything is setup is directly in your source control. There's nothing exactly unique about this, but it is an actual open standard. It hasn't been owned by Google for years. It's owned by a non-profit. It's open source. If your deployments work on a cluster provided by one Kubernetes engine implementation, it'll work on any of them. You can roll your own, use GKE, EKS, whatever Microsoft offers, and it'll work in exactly the same way wherever you go.
For one person, this means nothing. This guy is just doing it for fun. But for organizations that used to lose millions of dollars a year and sometimes their entire market position thanks to vendor lock in or had to experience months of downtime for a data center migration, they may find something like this useful. Or maybe you just don't want to rely too heavily on Dennis Nedry and want to be able to bring in anyone with a couple years of experience using an open source, open standard toolset who can come in and be reasonably expected to understand how your system works pretty quickly without needing ten years of tribal knowledge.
It sounds like you're making the argument that deploying k8s for a startup is an extreme case of premature optimization.
You can get tons of credits for your startup, typically hundreds of thousands from Microsoft, Amazon, etc. -- Eventually you run out of credits, so you switch providers. I did this 3 times at a startup and got three years of free infrastructure.
If you are selling enterprise software, then you use k8s so you can deploy at enterprise without having to integrate with all the wacky requirements.
You can build a saas that offers isolated service nodes on k8s infra pretty easy if you just give everyone a small vm with a k8s cluster on it.
If you use a scale to zero model, your infra costs are a lot cheaper. Simiarly the auto-scaling k8s capabilities are really nice and it's awesome to be able to easily build systems that can scale up massively. You never know when you r startup will get popular.
If you mean assembling your own k8s infrastructure by spinning up machines, then you're right.
If you mean deploying your software on managed k8s infrastructure, then this an outstanding use of your startup's time. You will be able to find plenty of developers already familiar with the workflow, and you'll have a straightforward time growing and deploying your app on different providers. It cleanly side-steps many production pitfalls that burned our time 10 years ago.
Because all that above? It comes with a system that will actively help you binpack your workloads as much as possible into the compute you give it.
Your usual "throw separate instances"? Can get expensive quick, especially if you're not aggressively modifying instance sizes. As for "lol just use Heroku"... I call that kind of company "bankrupt", but that's probably because the difference between cost of infrastructure and cost of engineer time is wildly different in my area to SV.
Many types of startups will benefit from K8s right from day one.
All startups are not the same though. If the startup has obtained any kind of reasonable Series A, it totally makes sense to invest in kubernetes, and allow your engineering team to essentially pull in literally any kind of dependency, test it quickly, and find out if it works or not. It acts as a catalyst that would let you churn out new products really quickly, and as such its an invaluable tool to allow your startup to move quickly.
If you work in enterprise environments and are able to use kubernetes, you are set to really shock your organization with how quickly you can move. I've seen this same situation play out in a few orgs and its really amusing how stupidly productive it allows engineers who learn it to be, and how quickly they get things done and get more responsibilities, promotions etc.
This is an unbelievable amount of gate-keeping hogwash. I don't know who this person is that they think they can arbitrate what is a good usage & what cases this is too-powerful too-interesting too-useful to bother using it in.
There is so so so much fear & doubt & scare in this post. Screw this gate-keeping crap.
> Maybe, if you're lucky, you had some set of really good and reliable Sys Admins that figured out a robust way to script and configure the setup process of your original on-premise data center and they captured that in very good, well-maintained documentation. If you're even luckier, maybe those same guys still work for you.
"Only us good right & virtuous & amazing engineers can handle this! This is too pure, too amazing for mortals! They're wasting their time! They'll get the configuration wrong! They're bad people. Only professionals are qualified to play with Kubernetes!"
UGH ENOUGH. Stop this terrible attitude. This is so down-talk-y.
Please don't assume, please don't dictate your limited terms to the world. Let the world try. Let us not be cowed, & afraid, to use good tech, by these scare words.
As it turns out, it's just not that hard. It's a better environment, a better world. There are lots of home users using Kubernetes. It doesn't take a colossal investment. It's fairly secure out of the box, at least if you're not trying to run a multi-tenant home. It just works. This blog post shows that! It's really simple.
See? Look. Lots of projects: https://github.com/k8s-at-home/awesome-home-kubernetes . Lots. Good people, just trying. Not taking the poison words to be afraid, that this is too hard.
"For one person, this means nothing. This guy is just doing it for fun."
What a BAD ATTITUDE. Snarking & being mean, to people out there, trying to find better ways to do things, to create shared, meaningful value. With good, autonomic systems. With reasonably competent free-to-everyone utility scale / cloud computing. Don't accept such words as these. Do not be afraid to involve yourselves. Do not be gate-kept like this. Run Kubernetes. Run good systems. Stop being sold on second, third rung systems. Believe in yourselves. Don't exclude yourself, be afraid. You are not saving yourselves any hardship by choosing lesser technology. You can run K3S in <20 minutes, and you can start loading amazing manifests & Charts seconds after. This blog write up shows that. IT MAKES PEOPLE AFRAID to think how democratic technology could be. Please allow yourself a moment of un-doubt where you consider, maybe, this has amazing value for the home, that it's already possibly incredibly robust, please consider that applying some manifests might be super easy. Please consider that blog posts are the canonical way to share work, before Kubernetes, but now I can link you to a repo full of people sharing manifests & charts & works, that stand a decent chance of running on any cloud or at home. There's a lot of sophisticated under-the-hood boons to running Kubernetes too, but as for what the home-user gains: it's amazing. It's easy to understand, eventually, and shared.
Skills learned building one thing in Kubernetes parlay much better into doing other Kubernetes tasks, related or not so much. Kubernauts are growing & learning & exploring, in a far more communal, shared operational way than what loose daemon-monkey-wrenching spread-out weirdness came before. It's not hard, it's not weird, it's not for enterprise. It runs out of the box, nicely. It's for people who want a good means to think about & control a wide variety of digital things. People who want a reasonably consistent underpinning to their technical operations. The alternative, what we had been doing, is having a lot of different things that each had their own means to be thought of & controlled; dis-unified, chaotic, piece-meal, inconsistent. Come try a better more unified way of computing. Come try utility computing. It's for people, all people. Many of us believe it serves us all better.
With config management, be it puppet, chef, salt, ansible, you start to approach that. Except these only manage the resources you define. If your config management manages apache resources, and someone wandered in and installed nginx that's out of scope. If you've got a list of machine-local usernames that you want to exist, you might not thinking about ensuring that other usernames are absent.
With kubernetes, it gets a bit closer. Let me preface that standing and running your own kubernetes cluster is a separate level of complexity, but once you have a k8s cluster, you can dump out a list of all resources (namespaces, pods, volumeclaims, ingress controllers, services...), modify that resulting file and apply it. Most people work in chunks, having a yaml file for each set of associated resources, or template those into a helm chart. With that, you can put all associated resources into a release, which you can delete and recreate as you desire.
There's still lots of gotchas, especially around persistence. It's not uncommon to find an orphaned volume claim that prevents your release upgrade from continuing until you dig in with 3 coworkers watching you as you try to remember if it's safe to delete that or not.
Kubernetes is very complicated, but the complexity can be managed if you approach it from the bottom up. Using a cloud service robs you of that ability.
Other companies hire skilled people that aren't necessarily super-smart, because super-smart people can make a lot more money at Google. So other companies use Google / Amazon, they don't really care what underlying techs it uses, so long as they know they'll be able to hire people with the skills.
If you're rolling your own Kubernetes, you're putting your faith in your own smarts, because you're largely on your own. K8s rewards a certain kind of developer, and badly punishes someone thinking running their own k8s is going to be as seamless a process as consuming GCP resources.
s/Kubernetes/Computers
To your point, when talking to your db or [insert shared infra here], then you may want to thinking about real collocation. But it very much depends on your application’s needs - and remember that cloud providers can offer guarantees about which AZ or region you’re in.
With the magic of VPN, I can put my colo box on the same k8s cluster as the rest of my home stuff, in case I want to live dangerously.
I’m really sorry to read this. Given that you’re a Comcast customer it’s probably because you have no other viable high speed choices, but please know that regular outages at home are not normal in other parts of the world.
I have never been madder at a service provider of any service. But there is nothing I can do, at least not for six more months until Starlink finally goes online in my area and sends me a dish.
This in what? The heart of the 4th largest metro area in the richest nation on the planet? Maybe this is legitimately not a problem in most of the developed world, but U.S. residential communications infrastructure is amazingly pathetic.
I use Terraform to manage the few personal cloud resources I have. It is overkill, but I already learned it on the job so it ended up being quicker than setting stuff up by hand. If not, I wouldn’t have bothered learning it. At least in this context.
Same when building an MVP, I’m not going to pick some shiny tools I have never used before, but tools I know I can get the job done with.
Sure, I do like learning new things and tinker around with them, but it is not always the right time to do so, and the list of things I’d like to get good at, engineering related or not already is long a enough for a lifetime so one needs to prioritize.
I recently saw a video via twitter that mocked using kubernetes for your blog by showing someone making a sandwich with woodworking tools.
Overkill, yeah. But if you stretch the metaphor to the way it translates in k8s, if you get good at making a sandwich in 5 minutes with the tools, you're fairly close to also stamping out dressers and end tables in 5 minutes. At the very least you can make them, which you can't say of a bread knife. And once you do it a few times you can do it with incredible speed.
k3s is specifically designed for resource constrained environments (e.g. Edge, IoT, CI/CD) and can get away with as little as 0.05 CPU and 256M RAM (on worker nodes).
I don't recommend this as a way to save money on infrastructure but it does open the door to further use cases.
Don't know if that was an intended nod but I liked it :)
https://storage.googleapis.com/pub-tools-public-publication-...
I'm betraying my ignorance here, but how does this work? If you're running the registry in your cluster, and you tear down your cluster (and the registry with it), how do you rebuild the cluster without being able to pull images?
Assuming that you rebuild your cluster and restore your registry from backup (or simply reconnect it to some object store that it was connected to before, like S3), your registry pod (in a Deployment/Stateful/DaemonSet) will come up, and then the cluster will be able to connect to it, and the pods that were failing to start will start succeeding and you're back where you started.
My home "server" runs plex, sab, mongo, unifi, and various other things in a single node k8s with a local zfs volume provisioner. The previous revision of this server I'd switched to using docker for everything and was annoyed with the upgrade process with just docker alone. With k8s, I just use :latest for most images and an upgrades happen every time I restart a pod or reboot the machine.
I've been working with k8s since petSets and so this is all NBD to me. Linux had a learning curve too and ya'll got over it.
> Jobs do not actually automatically remove themselves from the cluster.
set `ttlSecondsAfterFinished` and they will
> And as far as I know, you can't even reuse the already created Job, so if I want to migrate the database again, I need to make it from scratch.
If you need to run a job on demand, you can always 'kubectl apply -f' the metadata, or you can create a cronjob and `kubectl create job --from` your cronjob. Or, if it needs to run at a particular part of your installation/upgrade workflow, use helm hooks. Jobs are intended to be a one-time thing, so the lack of functionality to repeatedly run them is intentional.
Gets deployed as a k8s `Job` via GitLab (can be scheduled/on-demand), with a simple script echoing back the vegeta status every 1 minute. Jobs don't yet support automatic cleanup, so another simple GitLab job deletes the job from k8s upon completion. The execution log is anyway available in GitLab/Kibana.
All in all, an extremely simple way to run load tests against a service deployed in k8s.
What grew as a personal utility is now being used by many teams for quick load tests.
https://research.google/pubs/pub41684/ https://research.google/pubs/pub43438/
In other words, jobs seem like the go-to way to take advantage of unused CPU/memory.
While i can't really share the text (because it's all in Latvian), i did find that K3s is a pretty good option for when you want to use Kubernetes, but have pretty limited hardware resources. So much so, actually, that it used only slightly more resources than Docker Swarm (and, as a consequence, the deployments that were running on the nodes could serve almost equal amounts of requests under load).
Personally, i do believe that Kubernetes is sometimes overrated for simpler deployments, but if people don't mind the manifest syntax and are comfortable with the tool and its ecosystem, i think that K3s is probably the way to go in those situations (maybe with something like Portainer/dashboard running for graphical administration, if necessary). Such a good distribution, could probably even run it in prod with a HA setup.
Just wondering, how are you exposing your services from the cluster? (I assume you aren't using a load-balancer ingress)
I was immediately overcome by all the abstractions. I had no idea where to look to figure things out, and was mostly relying on advice from my friend. I dind't know what an ingress controller was, much less how to configure one--but I knew I wanted each 'service' to have its own IP I could route to from my network.
Overall it felt like I had SO MUCH TO LEARN at *every point* it was difficult to get even close to my actual goals (CI/CD for a personal site + some hosted game servers).
I eventually went with the philosophy of "build something now, and move towards perfection later" / "don't let best be the enemy of good" and begain spooling up LXCs and VMs to do the work I needed, planning to move things into k8s later when I better understood the actual things I wanted to move.
(Plus then I got some satisfaction out of actually accomplishing the goals I wanted, instead of just banging my head on k8s documentation and learning all the abstractions.)
As an example, I've not used docker in any meaningful capacity. For anything. No idea how to make a docker image. To k8s the CI for my docker site, I needed to know how to :
1. install the dependencies, which requires compiling a plugin for pandoc, which requires installing haskell and cabal. this is expensive, so I'd prefer to get the pre-res set up once... but that doesn't seem to be how docker works? Do I need an image repository? can I use DockerHub? I've seen HN talk about how docker is trying to monetize, should I run my own repository? Can I do that on my cluster? I'll need an ingress controller to route to it... I don't even know what that is. 2. I need some way to pass the built website files to the contianer I actually want to host them. I think that means I need an NFS share of somekind to store the files, so one coantainer can load them and another can read them. Do I hos tthat on my NAS? I could put an NFS share in the cluster, maybe. No idea how to get Docker to mount one, or k8s to host one. All the examples I seem to find deal more with connecting to remote services on a host than mounting local storage. is it even local storage? 3. Everyone says infrastructure as code is good, so I guess I'll follow this flux tutorial--only to find out the one I followed is out of date, and I should follow their NEW one. But they still assume I know way more baout k8s than I actually do, but still, I'll spend the few hours to get this operational, so then in theory everything else I deploy can be IaC'd, which is just good practice.
At this point I'm so many layers of abstraction deep, I have no idea what I'm actually doing or how concepts relate to each other, and I'm no closer to actually having my goal.
So last night I spent 2 hours spinning up VMs on my cluster and installing dependencies, configuring an nginx proxy, and now I actually have my personal blog self-hosted and updatable. Way more "progress" than the 10ish hours I've sunk into building a k8s cluster already.
There's something to be said for limiting the number of abstractions you're dealing with.
I've seen this comment a lot when discussion introductions to Kubernetes, and it is probably one of the biggest "first step" problems I see, and a perfectly legitimate complaint. If there was one thing that Kubernetes could address to help "onboarding", it might be thing.
Kubernetes is essentially a collection of controllers, that each control a different aspect of the system. These are all "internal", and you don't need to know/understand them in order to stand up Kubernetes in the happy-path. If you're using a managed solution, these are all managed by your provider anyways.
The ingress controllers are the exception. It brings the concept of "controller" out of the "Kubernetes administration" space, and into the "Kubernetes user" space. It opens up a whole can of worms around "What exactly is a controller?" that you shouldn't have to understand in order to get started using Kubernetes.
(of course, the whole point of the homelab endeavor I set out on was to be fully self-hosted to learn all these underlying concetps--but definitely doesn't help that I layered some extra abstractions on top of the soup that k8s already is.)
I feel like this was more of a hyperbolic rant.
Reasonable--it definitely is a bit of a rant. I'm not sure it's hyperbolic, personally, but clearly I have a bias.
Overall, my acute frustration with k8s coming from a native-destkop-development background is I simply don't have the scaffolding to be effective relatively quickly. The learning curve is steep enough that I get discouraged before I actually begin making progress on my own goals, and I just feel like I'm trying to do things The Right Way, without understanding what I'm doing, or why, or actually accomplishing my original goals.
As you said, this is true of concepts I've had to learn in the past, but I could learn all those concepts in isolation, then apply them together. K8s I feel like I have to understand a much larger chunk of before I hit critical mass and can start being effective.
i.e. python, I can start with, like, sqlite, before I move to an external hosted DB. K8s I feel like I have to understand waaaay more components--and they don't directly build. Like Docker-compose _seems like_ a stepping stone to k8s, but I've been told is a false path.
I have been using Docker for work for a while, though I still don't feel much desire to use it for any personal projects. It seems straightforward enough - basic idea is that each Docker image runs one and only one process, and has it's own directory tree. So if you want to run a Python or Ruby webapp, all of the app files plus the interpreter and all required debs/packages/gems live in that Docker image, and you don't need to have anything else special on your host. A Dockerfile is then a sort of script that holds all of the instructions for setting up an environment your application can run in from scratch.
Of course that means if you want to run a database or a cache server or something, then you'll need to run several Docker images and coordinate communications between them and launching and scaling. I gather that's where Kubernetes and Docker Swarm and other such things come in.
I wholeheartedly agree :)
The path I'm currently taking is to spin up Concourse CI which uses containers, and start using that to CI/CD my personal site. I'm likely gonna overkill and end up hosting my own image repository, but overall I think this'll teach me the necessary concepts of containers in a directly applicable way to my end goals.
From there I can start to play with k8s itself, if I so desire.
Where I tend to get hung up is things like entrypoints. My personal site is static, so the container to build it isn't a hosted service, it's just invoking Pandoc (with a bunch of dependencies installed as well). I think that makes Pandoc the 'entry point', that just seems strange since it isn't a persistent service, and I think that makes the lifetime of the container rather ephemeral.
Which... likely is a valid usecase for docker containers, I've just only interacted with them e.g. on my Unraid box as persistent services.
The Dockerfile image build process would take care of building that. It's pretty standard for your single project Dockerfile to first assemble an image with a bunch of build tools, like this Pandoc and whatever other dependencies are needed, copy the source files into that image, run the commands to generate the final site, then assemble a second image with just Nginx and copy the output files from the first image onto that to become the final output image. That way, the whole build process is scripted and automated and doesn't require any of the tools on the local system (handy if you were doing stuff like onboarding new contributors), and the final image is small and secure because it only contains a web server and the final HTML pages and no build tools or source files.
I've been quite happy with it for a few years now on a single-node cluster and a handful of services. I could easily add another node, but haven't had the need to. The setup was quite simple and it requires practically no maintenance. If I have to reboot everything starts up automatically, image upgrades are a breeze, and I really have no issues with it.
Sure it doesn't have all the bells and whistles of a k8s cluster, but it's perfectly fine for personal use.
I'm still partly annoyed that Swarm is mostly dead in this space and k8s has undoubtedly "won". It's only a matter of time before Docker Inc. fully abandons it. Such a shame.
A while back i also tried Nomad, and while it felt polished, it was also a bit lacking in features (manual encryption config, mostly read-only UI, though that was a while back).
It's a shame that Docker Swarm never got more popular, because in my mind it's basically the sweet spot between running containers directly or using Docker Compose and full blown Kubernetes distributions. The Compose manifest format that it uses is really great and clear in my mind, instead of Kubernetes with all of its selectors, object types and so on.
Even this blogpost doesn’t explain what and why’s of Kubernetes.
I have a docker container. I deploy them on vm. I use load balancer to split traffic. Could you please walk me through what problem Kubernetes would solve here?
If you run your docker app on kubernetes, you get a lot of things for free with the platform (rolling no-downtime deployments, service discovery, auto scaling) that you’d have to set up manually if you were running your service on (say) EC2 instances instead.
It can be a headache to learn sometimes, but ultimately saves a lot of effort if your use case fits!
So it solves the bin-packing problem automatically rather than having to manually map out an efficient way to use your infra. It doesn’t reduce the complexity per say - the interactions can get complicated if you’re using stuff like node taints/tolerations instead of the more computationally simple “App X needs this much RAM but beyond that I don’t care where it lives”
If I find a bunch of raspberry PIs in my basement and want to have them join my fleet, I just do that and even if my fleet varies from 256 core CPU boxes with huge raid arrays all the way down to tiny raspberry PIs, the scheduling just works. Note here the broader pattern of abstracting away the physical hardware; it’s a really important concept to grok.
But what if I told you that you needed to run that docker container on 5,000 machines with 2,500 load balancers in regions across the world? Are you going to SSH into thousands of boxes manually and run docker commands? Are you going to try setting up some monster ansible inventory to do the same? In practice the best minds in distributed computing have found those kind of practices break down at large scale--you just cannot reason or deal with individual machines when there are thousands of them.
This is where kubernetes comes in--it's an abstraction that lets you declare "here's the state I want, X machines running Y containers, all linked through Z services" and kubernetes will make it happen, period. It will take care of contacting thousands of machines, controlling the running containers, ensuring they stay running, handling failures, monitoring, load balancing, etc. You no longer think about problems in terms of low-level machines and instances, you think about the higher level objective like deploying code.
The beautiful thing is that it scales down nicely. A simple 50 line YAML file that declares running your docker container and load-balancing it with a service can easily deploy just to your local machine, or be scaled up to run on 5,000 machines by just changing a variable in the deployment scale. The same simple one-liner kubectl command kicks of either deployment and helps you monitor its progress. If you've ever worked in distributed systems it is really incredible to see this in action at scale.
Said another way, if Linux (or whatever) is the OS for your server / VM / host level / network device, k8s is the OS for your cloud application.
And, when k8s is implemented properly, it takes a lot of headaches that can come from dealing with the myriad problems that arise when your architecture goes beyond a basic handful of "tiers."
- If you wanted to scale up or down the number of container processes running on your VM, you'd need to write some code that looks at system utilization. Autoscaling k8s clusters do that for you. They can even provision additional VMs ("nodes") for you during times of heavy traffic, or scale down to save money.
- Updating your app requires either logging into each VM, or writing an Ansible playbook to do that for you. By the time you've written a zero-downtime, health-check-honoring, contextually-aware Ansible playbook, you've made your own container orchestration solution.
- If you run multiple containers that need to talk to each other, you'd need to handle their networking. K8s gives you tools for handling networking between containers in the same namespace that allows them to communicate without exposing them to the wider internet.
- The ecosystem of utilities is as good (and sometimes better) than you'd experience in your VMs setup. cert-manager makes certificate management almost as easy as LetsEncrypt does on a single machine. Prometheus and Grafana are excellent logging and monitoring solutions (and, IMO, much easier to setup on K8s than ELK is within a distributed VM setup). Cillium provides extremely powerful and useful networking and security policies that leverage eBPF
- Changes you make to the configuration of your server won't carry over if you ever need to switch hosting providers, or (more often the case for me) just want to start fresh.
It's absolutely a huge learning curve, but eventually the complexity (mostly) goes away, and you're left with a reproducible method for deploying apps. So in the same way a rails/django developer might use an overpowered solution for their blog API, or a React developer may build a custom frontend when wordpress would also do... someone who's taken the time to familiarize themselves with K8s might find the familiarity and consistency of the interface enjoyable, even if it is clearly killing a fly with a sledgehammer.
For better or for worse, you need something outside the Dockerfile that can run it. That can be you, if you want to type out 'docker run -p8080:80' etc. You could probably script it, but does your script do restarts, failover, etc?
It does all of this by allowing you to specify a service architecture in configuration files and then actively ensures that this configuration is maintained even as the underlying state of containers change. You can specify things such as the minimum number of backend containers needed to provide a service, scaling parameters to add more backend containers as load increases, and you can tag nodes with different attributes so that containers are distributed and maintained with the appropriate amount of resources.
Kubernetes also provides automating various aspects of networking such as provisioning and configuring load balancers for service ingress as needed. It provides an internal DNS service which automatically registers names for deployments so that linked services can just refer to each other by name without any additional configuration. It can also manage things like SSL certificates which can be shared across multiple services.
Lastly, it provides you with a single place where you can store secrets and configuration values that these services require and again, you wire all of this up with configuration files which can be stored in a git (or other VCS).
But that is what it means. K8s, ECS, even docker swarm are ways to orchestrate containers to do something useful.
Take your lb example. What happens when one of the containers you deployed or VMs you deployed to dies? How does it get restarted? Where does the lb send traffic?
What happens when your VM dies? Kubernetes would automatically bring it back up. It has health checks, and knows when containers die/crash.
What happens when you need another docker container due to traffic? Again, Kubernetes fixes situations like this. Kubernetes has a lot of built in support around scaling etc.
Also, what if your docker container doesn't need a whole VM? Say you've got 5 different docker containers (which all scale independently), and lets say 3 VMs. Kubernetes will distribute them across those VMs based on there resource needs.
There's a lot more, but that's kind of what I think of when you say 'Orchestration of containers'.
If you start a deployment ("your container", roughly speaking), the age of that deployment will keep counting up even if the container exits-- k8s restarts it of course-- and even if the k8s scheduler goes down-- since we want to be able to restart the scheduler without unnecessary service restarts.
At home, it's all on one computer in my basement, so when I reboot that box, the k8s reports come back and keep telling me the deployment has been there for xx days (just with some availability hiccups).
Even with a reboot, most OCR implementations (docker daemon, eg) will keep a pod around for 10 minutes until they reschedule it, so you'd see a restart count for the pod but probably not even a different age. It's tunable but IIRC 10 minutes is the default.
I use Kubernetes because it makes my application, its configuration, and its operations portable.
That's it.
A startup that grows to have hundreds of developers might transition from running managed VMs to "the cloud". One team sets up the network (virtual networks and some subnets).
As new employees join, they have no reason to interact with those teams who are effectively "hidden" so they deploy their stuff and perhaps wrangle with subnets and what not. Someone tells you that you need to attach subnet-a20w88vhuh4fuih to your resource and it will magically be accessible in the office.
Nothing is in charge of VM sizing so you've got people blowing hundreds or thousands on massive VMs when they're only using 10% of it and vice versa, teams whose application is choking but they don't really have a good mental model of say general purpose VMs vs memory optimised so they just bump up the SKU instead of being more efficient. This is happening everywhere as the company accelerates more and more.
It gets worse when you have a shared cluster say; for an entire team that is globally distributed and the new intern application is doing some weird O(n)6 computations and absolutely blowing the side out of every other resource you've got.
Now at this point, it's effectively a communication/culture problem but Kubernetes can "fix" some of these issues in a sense.
Network for the most part becomes abstracted away and what you're left with is defining security (what ports and protocols should I expect) on an application level, rather than on a security group level. It's kinda neat because these rules are localised to your application whereas they might have been configured manually in a cloud portal or via some terraform config owned by some team in the shadows.
Each of your deployed applications become their own isolated units called pods. A pod could be one or more containers but it's effectively a standalone slice of an application (ie the web frontend while a redis instance might be another pod). There are bigger abstractions to group application pods together but that's besides the point.
These pods get deployed to a cluster (a bunch of VMs) and cough "orchestrated" but the value here is that your containers might be running right next to some containers for the business team or the machine learning team and you would never know. You don't need to know either. The value, as foreshadowed above, is that if you're being a noisy neighbour, your container will either get rebalanced somewhere else or just shut down for exceeding memory usage.
I'm a bit flakey on this point but since each node in a cluster is a massive VM, there's no need to worry about over or underspending based on your computational use as well. You define the amount of memory you want to allow and you get matched to a relevant node based on how much capacity is available. As you gain more users, you just add more nodes. Before that, you might have been "reserving" say X thousand compute hours of certain VM SKUs or whatever. You might still do that but you could feasibly just have whatever your node sizes pre purchased making capacity planning pretty straight forward.
Generally, there'll be some team whose purpose is to manage said cluster so in a funny way, it somewhat revives the whole dev/ops split in that your compute team generally know the nitty gritty of networking and what not while your developers just deploy an application and it "lives on Kubes".
I may have a missed a bunch of stuff but hopefully this outlines some of the more "people" issues a bit? It's half and half useful but also it can be used as a technical fix to a social issue.
If I add more physical hosts or scale the pods, k8s does everything for me like moving them around across the available resources
Seems to me like it's helpful to do an oil change yourself before you take it to the dealer, then understand what they do before you take it to jiffy lube the next time (or vice versa). You keep abstracting it away until you just give your credit card. :)
Deployments, scaling, logging, etc are some of the patterns it provides and the consistency matters. How many of us have worked at companies where deploying two services have been completely different? One team runs the jenkins pipeline while another team FTPS the files over. Now multiply that by several services and several tasks (logging, scaling, etc).
The benefit is in the patterns and standards it provides.
Again, totally not hating on this, we all have our hobbies and I love that the author is super into this.
Simple as sending an email!
I’m asking because I too made an email->blog post service for myself and I want to see if this is productizable.
Very simple. Although I still haven't found out how to embed links into text.
kubectl apply -f .K3s claims to be production ready.
- `jekyll build` to build `_site`
- `neocities push _site` to recursively upload modified files in _site
`docker build [...]` the image and then
`kubectl set image [...]` to update the image that kubernetes is using
Or, even better, you can just set up CI to do everything on commit/push.
For the last static site I deployed, I just tossed Caddy on k8s and set up the git module. I commit, push, hit f5, and my site is already there. There's a ton of ways you can use kubernetes, which is probably part of the adoption problem.
With k8s and a proper build system, you can roll back easily if you introduce a bug or something doesn't work. More importantly if your site doesn't start up in any specific way you deploy it, k8s won't even start a new pod.
Everything sounds all fine and dandy if all you think about is the happy path.
Seems like overkill. :P
I have a git repo containing all my helm charts & docker files, testing & deploying changes is absolutely trivial now. And it's great to have everything version controlled.
Previously I used to use Ansible, but you quickly run into issues which make you want containerization: Conflicting library/tool versions, packages that pollute too much of the system, port conflicts, hassle of keeping the playbook idempotent, etc.
So while docker-compose would also do fine, having kubernetes to manage the ingress' routing system is rather practical. And the same goes for the other bits and bobs of infrastructure offers you if you're already using it. It's just very convenient.
I've been doing this for a few years now, and am now up to 14 different apps running on my single home machine in Kubernetes, ranging from Home Assistant to PostgreSQL to Plex.
Also it's just good experience. I also use Kubernetes for work, and this has made me noticeably more proficient.
Single node Kubernetes is pretty silly on the surface, but it is useful in some cases.
So if you ever have more than one application e.g. a search engine for that blog then it's much easier to add. SSL in particular is often a real pain to setup and varies in capability wildly between applications.
- Expensive in money: Pay Amazon or Google on the order of $100 a month for reasonably reliable managed k8s, with the option always available of a goofed deployment increasing your bill by an order of magnitude
- Expensive in time and money: Set up and run your own k8s cluster, at a monthly VM hosting cost not too far from what you'd pay a managed k8s provider, and spend a ton of time dealing with 100% of the ops burden
It's a shame. I use Kubernetes at work and I like it a lot, because it does a great deal to make complex tasks simple. It would be really nice to have that same kind of fungible resource pool available for personal stuff too, and be able to just knock out a few lines of YAML and deploy a new project and have the orchestrator take it from there to running without any further effort on my part. But the economics just don't seem to be there.
(I did try DO's managed k8s service, which is considerably cheaper than the big players. It was also very new, and very flaky - not a knock on DO, it's reasonable that a new kind of service would have some teething problems. This was also a year ago, so I wouldn't be surprised if they've gotten it considerably more stable since then.)
fear is the mind killer.
Managed k8s or a simple installation in a home lab behind a NAT make it much more tolerable.
Using klipper-lb (from k3s) would be great, but then you are basically tied to k3s.
So here's a way to host your personal blog if you don't want to over engineer it:
1. Have a git repo with your nginx/whatever config files.
2. Have a VPS running debian.
3.
apt-get install nginx git ...
git clone ...
ln ... # create symbolic links to your nginx/whatever config in your git repo
systemctl restart nginx ...
4. You're done. Create a cron job to automatically pull the latest changes from your git repo if you want.The above steps should take most people around 10 minutes.
If you need to actually pivot into something that scales easier from there, I recommend following these steps/levels as your scale increases:
1. Create an automatic install script.
2. Use that script to create a .deb package instead of installing directly (optionally create a repository for this).
3. If you want to move to docker or what have you, it's trivial to install a debian package in a container.
But let's be real, you never even will have to do any of the above because it's a personal blog, and it'll probably scale to the world population on a $5 VPS, especially if you slap cloudflare in front.
As for updates, caddy v1 used to support pulling in from git, but I don't think that got ported to v2. So what I do is build+push a docker image, and have a cron job on my vps to pull+restart my website's container.
My preferences go Traefik > Caddy > Nginx, but traefik definitely has a bit of a learning curve.
https://github.com/andrewzah/andrewzah-com-source/blob/maste...
https://github.com/andrewzah/andrewzah.com-docker/tree/maste...
https://github.com/containrrr/watchtower
It would automatically pull the new image and restart your containers.
The simplest, most comprehensive cloud-native stack to help enterprises manage their entire network across data centers, on-premises servers and public clouds all the way out to the edge.
Sounds ideal for a personal blog! /s
In my particular case, I run ~5 services on that VPS (used to be ~13), and I run about ~30 containers on my home server for https://zah.rocks. Having to manually update caddy or nginx and restart every time I added a dns entry would be a huge pain.
In addition, Traefik's middlewares made it relatively simple to add in SSO with External Auth Server/Authelia + Keycloak/OpenLDAP/LDAP Account Manager (LAM).
What happens if you get a spike in usage and it kills your little server? What if you want to run on a bigger host/smaller host/your image is killed by ec2?
An automatic install script with a deb is _way_ more work and way less portable than writing:
FROM nginx COPY . .
and
docker build -t <appname>.azurecr.io/myapp
az webapp restart
I suppose you mean automatically update? If so, this has been solved decades ago. Unattended upgrades is just another debian package you can install.
If you meant just update in general, that's a pretty silly question.
> An automatic install script with a deb is _way_ more work
You misunderstand. You can use the automatic install script to create the .deb - or a script based on it. Basically instead of on a live system, you place the same files in a tree you create the deb package from.
Even if you use docker, I still recommend learning about how debian packages work. Installing your app in a docker container using a .deb is strictly nicer than having all those steps in a Dockerfile. Plus this way you can specifiy your secondary dependencies declaratively.
> more work and way less portable than writing: FROM nginx COPY . .
Now that is just flat-out wrong. The steps you'd execute after "FROM nginx" are the same steps you'd execute after apt-get install nginx. It's not more work, it's the same or less - because you don't need to deal with docker on top of everything.
You have the added benefit that afterwards you can install the .deb in a debian-based docker container, but you don't have to. You're not reliant on docker.
Oh great example btw. Because the debian-based nginx docker container is essentially created with "apt-get install nginx" in its Dockerfile. Funny how you're trying to argue against the setup I proposed by citing a docker package that does things precisely that way.
> What happens if you get a spike in usage and it kills your little server?
Yeah that's not gonna happen to a static blog behind CF. Also at least that way you're always only paying $5, regardless of what happens.
The sort of traffic that kills a $5 VPS serving only a basic static blog can bankrupt you on an automatically scaling infra.
I disagree. How do you update manually? If your answer is ssh in and manually update, that's how you end up with a pile of different versions, and no reference what they are? You've now got random packages at random versions. If you automatically update, they're not infalliable, e.g. here's [0] a post from 12 months ago about someone having an issue with nginx not restarting.
> You misunderstand. You can use the automatic install script to create the .deb So now you have a deb package, and an install script; that's more complex and more work than a dockerfile (which is standard at this point. How do you update your deb package? With docker, you use the same commands above.
> You have the added benefit that afterwards you can install the .deb in a debian-based docker container, but you don't have to. You're not reliant on docker
You're reliant on debian, and the configuration of the OS underneath it, including the version of nginx, python, ruby, etc. With a docker image, that stuff is all pinned unless you engage with updating it.
> Yeah that's not gonna happen to a static blog behind CF. Also at least that way you're always only paying $5, regardless of what happens. > The sort of traffic that kills a $5 VPS serving only a basic static blog can bankrupt you on an automatically scaling infra.
Nobody is talking about auto scaling infra here, except you. You can run docker on a $5 VPS, or on a $5 DO app platform. If your blog dies, e.g. due to aws performing maintenance[1], you need to bring it back online. If you migrate to a smaller host or larger host, you need to rebuild it.
[0] https://www.digitalocean.com/community/questions/problems-in... [1] https://www.quora.com/How-often-do-EC2-instances-fail-Why
On one server? What?
> With a docker image, that stuff is all pinned unless you engage with updating it.
Yeah now we're entering super silly territory. You can pin versions[1] in your debian package control file and/or on debian directly.
But in general automatic updates are preferable.
Unattended upgrades is often configured to only install security updates. Better to have stuff potentially going offline due to an update than running insecure versions of software.
What, do you subscribe to every single mailing list of every bit of software you're running in your docker containers, so you can manually update them when there's security issues with one?
You better have some automatism here or your setup is strictly worse than what Linux distributions could do decades ago. Like not just worse, it's broken.
> e.g. here's [0] a post from 12 months ago about someone having an issue with nginx not restarting.
That's ubuntu, not debian. Very different approach to non-security updates.
[1]: Here's how the depends line looks in one of my debian packages' control file:
Depends: python (<< 3.0.0), ffmpeg (>= 10), imagemagick (>=8), webp, clamav-daemon(>=0.98.0), cron, haproxy (>= 1.7.0), psmisc, ntp, file, certbot, firejail, exiftoolRe; version pinning that's cool - both methods work.
For updates, of course it's automated that's the whole point.
> That's ubuntu, not debian. Very different approach to non-security updates.
Actually this is a really interesting point; I don't want to be a sysadmin, I want to run my applications. I don't want to be aware of the platform differences. I actually didn't know that ubuntu and debian handled updates differently. They both expose the same interface to package management. Features like that are exactly why developers like me should be deploying docker containers to DO's app platforms for $5 a month, and not running VPS's
Just don't do any manual poking. You can't (really) do that with docker containers, and you shouldn't do it on a random server.
On a virgin server I generally do this to install my application, and I do it with a simple script (you could also use orchestration tools, but that's overkill for <10 servers) because I am even too lazy to enter five commands:
1. Add my repository
2. apt-get install my-app unattended-upgrades sshguard ...
3. config like setting the db server
4. reboot
That server will now automatically install security updates, but I need to manually tell it to install/update anything else (because I like it that way). My application, which sometimes is just a package that contains some haproxy/nginx config (if I run the server as a reverse proxy) is installed from my own repository, so I can update the way I'd update anything else on that server. You really shouldn't do anything on that server besides telling it to update (if you haven't set that to be fully automatic).
There's time-tested orchestration tools if you want to automate anything of the above. For some applications I have my build server poke a server that'll make the application on each other server update.
> I don't want to be aware of the platform differences.
I can understand that position, to a degree. But even with docker you'll be using some OS, or package manager repos, (generally) in your container.
If you use the docker repositories however, you're pretty much at the mercy of the update practices of individual packages, whereas with systems like debian you know what you're getting - which for debian is stable, but sometimes not up-to-date packages. And if you build debian-based containers, you should still be aware of this.
In any case it's really worth it to understand how dpkg/apt works and how debian packages work. It's not complicated and can be done in a day or two. You're already using them anyways.
It may also help you better understand exactly what niche docker is filling and you won't have smartasses like myself tilting their heads at you when you list something a typical linux system does (and sometimes does better) as a value proposition for docker.
Basically docker replaces and improves upon what people used to do with VM images. That's what it's competing with in the sysadmin world. It's not directly competing with debian/apt/dpkg/whatever, even though it makes some of their features redundant. It may make working with the latter more forgiving, because you can fiddle with a Dockerfile until it creates something that works, but it doesn't replace them.
Obviously protection from this problem isn't something Kube alone provides but it's still a problem with the container-less setup you're describing. Flatpack could be an alternative.
While we're on security - thinking simply installing your program via apt and running it in your VM is in any way compatible to containers in terms of the security benefit is simply wrong. In a container you'd getting filesystem isolation and capability restrictions OOTB. Via some program in apt you're going to have to deal with something like apparmor and a variety of other tools to get something even compatible to a Kube pod with default configs.
> Basically docker replaces and improves upon what people used to do with VM images.
This trivializes the feature set docker actually has.
If it's for just the blog, and the only goal is to run the blog, then it is (almost definitely) overkill.
If you already have a cluster up, or have a bunch of other projects already running on the cluster, then Kubernetes is likely the easiest way to run a blog.
> and the only goal is to run the blog
If part of the end of of "run by blog on Kubernetes" is "I want to use this as a learning experience to run other things on Kubernetes afterwards" it's definitely a worth-while exercise.
> you'll find that it is now trivial to deploy a hundred other things to it.
I mean, sort of. Starting with a blog is a good introduction and gives an on ramp to running other services, but I wouldn't say it is sufficient experience to call running hundreds of other services "trivial".
That only has 600MB of memory and the minimum memory requirements for a master node is 2GB[1].
[1] https://docs.kublr.com/installation/hardware-recommendation/
The f1-micro would be the worker node. It's still a bit of a squeeze, because GKE has system workloads that need to run on user nodes.
Honestly, it's a super comfy setup, and very little maintenance. The one thing I have is a master update script that does the docker image upgrading and rollouts. I make updates all the time and pretty much don't think about it. More energy activation than a PaaS like Dokku, but worth in long run, I think.
There are other solutions that are potentially better (e.g. Dhall) but cdk8s seems to have momentum and sense for tackling the practical stuff (integrates easily with cdk, library with simplified constructs cdk8s-plus, import and convert existing stuff easily etc)
He's on AWS, he could have gotten all those benefits with ElasticBeanstalk or ECS even. Plus no yaml files but actual IaaS (terraform etc.) and much better integrations with the other AWS services.
I'd personally still use EC2 for a blog, but if you're looking for convenience/battery-included type of thing...yes I'd argue you're better off just doing X instead.
Especially when it's just a personal project.
What matters is the declarative approach of Kubernetes. It's like the first time I learnt about React: declarative renderering/deployment based on state !
I need 2GB minimum for a K8s master node, 2GB minimum for monitoring, and then 700MB minimum for each worker node[1].
I just can't justify running K8s for a blog when there are less expensive and complex offerings out there.
[1] https://docs.kublr.com/installation/hardware-recommendation/
I just run K*3*s when I'm not using a managed cluster, and it works great and is lightweight. I've never noticed whatever is missing from it.
I don't even know what Kubernetes is (because I refuse to look it up) but somehow all these people have heard of it and apparently not any of the previous simpler solutions.
With Ansible, you can easily shoot yourself in the foot with imperative shell command.
The point is, Kubernetes is very opinionated (for good reason) on how it's hard for you to break your system.
Imagine Puppet but instead of saying "run service X and Y on host A, run service Z and W and the hot spare for U on host B, ...", you just say "deploy at least 3 instances of service X with this much CPU/memory, deploy at least 2 instance of service Y with at least this memory, these are the hosts you have available, you figure it out".
I can also confidently say that having something approximating a stable web app demands doing a lot of serious thinking, and "a single server running Apache on Digital Ocean" does not cover that case sufficiently. You need to tolerate failure, failover, load balancing, bin-packing, etc. I used to run a small autoscaling group on EC2 for my own systems; the dang thing would fail to come up on one node very frequently and so a number of the queries would fail. I eventually burnt it to the ground and redid it. I've never had that hassle in k8s. Its designed to succeed, in a way the "box of parts" approach doesn't.
Boxes of parts are useful. For a complexity-sensitive & thoughtful infrastructure engineer, having something like the old Synapse/Nerve[1] system with your apps distributed across some 5-20 machines with a monitor lease to spawn new ones on failure would probably approximate Kubernetes for a few years, until you have to do something fancypants. You've still reimplemented part of Kubernetes, though... The other angle is, boxes of parts can go in wildly weird directions.... if you need it.
Looking at some infrastructure these days professionally, the question is - when do we move to Kubernetes. It's not interesting or useful to the company to be maintaining our own thing or own strange path. The only questions are around the path - how much rework needs to happen and how much building in k8s needs to happen to get there.
GKE is a very good starting point for k8s. Strong recommend.
n.b. With respect to the cost. I consider this a professional investment / professional development expense. Spending $100-$200/month on a software engineer's salary is a reasonable return for being able to readily say I have experience in a current topic. Also I can run my own apps. :)
Isn't that not quite correct?
If I'm reading this correct (which I may very well not be), isn't this a reference to a Dockerfile:
image: marcusbuffett/blog:latest
Which then might have lots of other complexity contained within. I wouldn't call that self-contained, I think that's overselling it.> A cron job creates a job object about once per execution time of its schedule. We say "about" because there are certain circumstances where two jobs might be created, or no job might be created. We attempt to make these rare, but do not completely prevent them. Therefore, jobs should be idempotent
Is there a way to make a cron job idempotent?
Out of curiosity, what do most people here who have done such a thing (with k8s or docker swarm or otherwise) use for storage? I tried briefly to use Gluster-fs but had really bad problems with performance (most likely because I set it up wrong). Right now, I just have a ZFS Raid on one of the nodes, and have an NFS share on that.
Rook (https://rook.io/) is also cool.
There was a period where Git was at this point as well, but it seems as though everyone's got over it and decided the learning curve is always worth it.
Some of us think editing YAML files is complex enough.
Is there an alternative that lets me orchestrate deploying a bunch of processes on a bunch of servers without having to interpose Docker in between? Anyone using Nomad?
So to approximate this with Kubernetes you have to avoid the big cloud providers and things like kubeadm as that would mean at least two cores and the resulting price tag (even if you use things like burstable). Using k3s is a nice option. Must avoid load balancer and often static IPs.
I wonder why there aren't vendors who are selling k8s namespaces directly with a quota (reinvented shared hosting in k8s).
There are a bunch of blogs/content out there that have been abandoned for a long while, but you can still read the information that was put there in the past. The simpler the system serving the pages is (down to static HTML files), the more likely it is that the website stays up functional.
A blog that is built on Kubernetes sounds like it will be dropped pretty quickly once interesting in writing for it wanes.
> I think Kubernetes has fallen into the Vim/Haskell trap. People will try it for 10 minutes to an hour then get fed up. The point where you start to grok stuff just happens too late for most people to stick with it. Those become a vocal minority proclaiming it as too complex for humans to understand, and scare off people that haven’t tried it at all.
I've had good luck with the tutorials and various help articles in the official docs[0].
Otherwise the official docs are quite good and worth starting there. I'd be wary of buying books that are more than a year or two old--the k8s space moves fast and they regularly have major updates once or twice a year. Books and docs can get a bit stale or out of date with current best practices.
Then there are the specific k8s concepts: pod (containers that share a network and PID namespace), service (a name, a TCP and/or UDP port number, and the name/selector of the target pod), ingress (HTTP routing, basically haproxy/nginx/envoy or other layer 7 stuff), LoadBalancer (the magic API that connects your cluster to the outside world, allocates IP addresses for Services/Ingresses).
Then there are the background parts that take the YAML and convert it into actual running stuff: apiserver (this is the central info hub, kubectl and kubelets and controllers all talk to this basically), kubelet is the actual agent on each node, controllers have control loops and those actually calculate the necessary actions/operations to achieve what the YAML declared.
Then there are the subsystems: storage (persistent volume claims and PVs), settings (secrets - see vault too - and configmaps, and various special/custom resources like a TLS certificate, these are stored in etcd, through the apiserver of course), CNI networking flavors and knobs, an endless list of kubelet and apiserver command line arguments, PCI passthrough for GPUs or HSMs or whatever. And of course there's RBAC, security, admission webhooks, see how nowadays Authz and Authn is done through a "lookaside" OpenId Connect service (some Ingress Controllers support these out of the box).
Azure App Services also fill this niche, but only works for web-based services. If you have a job queue and a worker reading jobs off the queue but otherwise not exposing a web endpoint, it doesn't seem to work quite as well and my (limited) attempts at faking an endpoint haven't appeared to pan out :d
Running it on your own metal is supposedly what requires that team of engineers.
Do note that you need a DNS provider with a good API to programmatically control it (the DNS challenge requires setting a special txt record that Lets Encrypt verifies). Read the docs to learn more, but most of the big ones (AWS Route 53, Cloudflare, etc.) should be fine.
There is! I use external-dns. [1]
I haven't actually set up a Let's Encrypt wildcard cert, but I'm pretty certain cert-manager [2] supports it. I don't think you need a proxy if you use the DNS01 challenges.
That, and the ongoing complexity you're buying into, is exactly the problem.
In particular, how is the blog updated? I don't see how replacing and/or backing up a Docker image could be more convenient than copying files through a SFTP client and using other kinds of real clients (administrative web interfaces, SSH, etc.).
For one or two services on a machine, it's pretty good. Past that... not so much.
[program:some-docker-project]
command=docker-compose up
directory=/opt/some-docker-project
...
[program:some-other-docker-project]
command=docker-compose up
directory=/opt/some-other-docker-project
...
This style of setup has served many millions of visitors across dozens of projects in production at our company over the last ~3 years. It's like a poor-mans kubernetes, it handles autorestarting failed projects, we handle ingress with cloudflare argo tunnels in docker, and deployments are easy (just git pull && supervisorctl restart all).1) Want Ci where push to github builds the image and deploys.
2) Binpacking time. We're at the point where that's an issue
3) Host machines -- There's too much on them and we need to standardize and streamline.
If you'd like to try kubernetes and get to understand it and feel comfortable with it I'd recommend:
- Work through kubernetes the hard way[1] (ignore/replace the GCP-specific things, if you don't know enough about what the non-GCP corrolary to a GCP-thing is, that's a bit of knowledge you need to fill)
- Set up your own cluster from scratch (no kubeadm, no alternate distros, etc), use the simplest options you can find at first (ex. flannel for CNI)
- Set up ingress (NGINX ingress is a good place to start)
- Run some simple unsecured but diverse workloads (static sites on NGINX, Wordpress, etc), figure out why you might pick a StatefulSet versus DaemonSet. At this point, use the hostPath/local volumes just to avoid trying to grok volume complexity.
- Install a useful cluster tool ("addon") like cert-manager[2] from scratch so you can see Kubernetes manage something you'd normally solve with systemd timers and `certbot` (or your reverse proxy would do for you if you're running caddy or traefik)
- Start putting your YAML in source control get familiar with either kustomize, helm, or both (you could also just use Make + envsubst like I did for a while[3]). It's at this point that it should click that all you need to get back to a certain state of your cluster is to get a machine, do basic hardening (ufw, etc), install kubernetes, and run "make" in this repo (excluding things like DNS entries, etc). Now things are probably getting fun, because you can have the distant cousin of immutable infrastructure; repeatable infrastructure.
- Tear down your cluster, set it up again with kubeadm (note that kubeadm actually has a file-driven configuration option[3]), run your yaml from source control and confirm that all the workloads you had in place are back up and secured.
- (optional) Tear down your cluster, try rebuilding it with k3s[4] or k0s[5]
- Start looking around and seeing what your options are and the ecosystem that exists -- digging deeper into the interfaces that make Kubernetes tick, for example volume management (Container Storage Interface) by deploying Rook[6] or OpenEBS[7].
Note that one of the best things that can happen to you during this process is something going wrong. Every time something goes wrong, and you go back to fundamentals to fix it. Being able to reason about kubernetes ad hoc (ex. if DNS is correct but you can't reach port 80/443, what should you check first?) is the key to feeling comfortable with it, and while failures and downtime will happen, you shouldn't feel just complete despair/confusion. Normally, once things are stable kubernetes is quite worry-free.
I've made a guide like this before, I'll see if I can find it.
[EDIT] Found only one post[8]
[0]: https://vadosware.io/post/fresh-dedicated-server-to-single-n...
[1]: https://github.com/kelseyhightower/kubernetes-the-hard-way
[2]: https://github.com/jetstack/cert-manager
[3]: https://www.vadosware.io/post/using-makefiles-and-envsubst-a...
[4]: https://k3s.io/
[5]: https://docs.k0sproject.io
[6]: https://rook.io/docs