Bare Metal K8s Clustering at Scale
medium.com
medium.com
I saw these guys talk at QCon. It was a fascinating talk, and an excellent example of SRE adaptability and nonstandard, uncommon innovation given unusual constraints.
Not speaking for them, just from my memories of the talk and the following Q/A, but their reasons for this stack were primarily:
- They couldn't run in the cloud because connectivity to their sites is often terrible.
- They mostly ran IoT stuff from the k8s clusters--automated kitchen equipment like fryers and fridges, order tracking/status screens, building control systems, and metrics aggregation so they can see how businesses are doing.
- Because of the bad connectivity, a "fetch"/"push" (from the k8s clusters at the edge) model was needed for deployments/logging/administration/getting business data back up to the cloud.
- They explicitly did not process payments.
- k8s was used primarily for ease of deployment and providing a base layer of clustered reliability for pretty simple services. Since the boxes in the cluster were running in often-unventillated racks/closets full of junk in random restaurants, having that base layer was very important to them. Other solutions were evaluated and they chose k8s after consideration.
- Unlike typical IoT/automation setups here, they wanted to be able to experiment, monitor, and deploy software without the traditional industrial control practice of "take shit down, flash your controller (call a tech if you don't understand that), spin it up, and if it breaks you're down until we ship a new control unit or you manually fail over to a backup".
- However, they didn't want to fall into the IoT over-the-air update security pitfalls (it would really suck if someone hacked your fridge's temperature control system and gave a week's worth of customers salmonella). As a result they spent a ton of time making very good (and simultaneously very simple) deployment/update authorization and tracking tools. They chose the "pull" model and keying/security layers explicitly to avoid having to think about tons of open remote-access vectors and/or site hijacking.
- The k8s tooling (and some of their own) allowed easy, remote rollbacks to "default/clean state" in case something went wrong, which was critical given that downtime might compromise a restaurant and having a "reset button" automated in was important for ease-of-use by nontechnical, overworked site managers.
- The clustering allowed individual nodes to fail (which they will, because unreliable environments), and people to manually yank ones with confidence.
- While, as some commenters pointed out, the leader (re)election system chosen might be unacceptably slow/randomized for, say, a cloud database, it is perfectly sufficient for failing over a control system in a restaurant. A few seconds of delay on an order tracking screen, or a system reboot/state-loss of in-flight orders is vastly preferable than some split-brain situation making the restaurant accidentally cook 1.25x the correct number of sandwiches for hours, to go to waste.
It's important to understand their use case: they needed to basically ship something with the reliability equivalent of a Comcast modem (totally nontechnical users unboxed it, plugged it in, turned it on, and their restaurant worked) to extremely poorly-provisioned spaces (not server rooms) in very unreliable network environments. For them, k8s is an (important) implementation detail. It lets them get close to the substrate-level reliability of a much more expensive industrial control system in their sites (with clustering/reset/making sure everything is containerized and therefore less likely to totally break a host), while also letting them deploy/iterate/manage/experiment with much more confidence and flexibility than such systems provides.
I think this is a great story of using new tools for a novel (or at least unusual) purpose, and getting big benefits from it.
Brian, Caleb: great talk, great writeup. Sorry HN is . . . being HN. Keep at it.
Edit: QCon talk summary is here: https://www.infoq.com/news/2017/07/iot-edge-compute-chick-fi.... If you have any employees/friends that went, they should have access to the video. It may be made public at some point, too.
I don't think HN was super vicious. They presented an out-of-the-box solution to a problem but they didn't define the problem fully. Based on what we saw, their solution seemed way overkill.
Glad to hear that there was a solid reason behind it, not just hype and recruiting buzz.
These days it feels like everybody needs to throw in Kubernetes at everything introducing complexity for the sake of being cool.
I guess those of us that likes to run non-distributed software for small scale applications are the new grumpy grey beards....
And no, there is not a full POS server on them at this time... it will take us a few years to decompose that monolith, but this is a natural place to put it as we do so.
I am a grumpy grey beard no doubt but I still maintain: most websites do not need more than a single server -- certainly not more than a single database server. And, for most, a few hundred dollar dedidcated server is aplenty. Apply YAGNI until blue in the face.
I run a Dokku instance on a Hetzner server and it's been fantastic, I host 10-20 of my projects with thousands of daily users there without it even breaking a sweat. My only regret is that I should have used a single database instead of one DB per project, but I was lazy.
True, most websites do not have this problem, because most websites do not drive revenue like that. There are plenty of use cases where you need five nines, but only within limited not-24/7 time windows.
But for non-distributed software with only several clients, the traditional model is still fine. E.g. we still run gitlab as a pet to serve our cattle infrastructure.
> Edit 7/2/18 — since writing this, many readers have asked “why not just use the cloud? Why computing at the Edge?”. We have realized that we did not provide much context about why we’re doing what we’re doing, so we will follow up with an post about that soon.
Hey! I'm Caleb, the SRE that helped build this solution...
Sorry about the lack of context in the article, it was intended for a specific audience (QCon) where we gave a lot more context to the problem at hand.
What we were trying to solve for was; 1) Low latency 2) High Availability 3) Container based, zero-downtime deployments 4) Continued operations even in an internet-down event
If you're running applications in a few locations, in small scale environments, our approach would be way overkill for that problem set.
Setting up a single-node kubernetes is basically adding one line to the system config:
services.kubernetes.roles = ["master" "node"];The only reason i can think of that is they get to push point of sale software out by using K8s from some central system. I cant think of a worse use/abuse of k8 as a software updafe system if that's what they are doing.
The other reason is they distributed their compute and resturants pay the power bill but that sounds just as silly.
Curious to know why you would use k8s at the edge
It absolutely smells of over-engineering, though. There are a lot easier ways of pushing software out than maintaining k8s locally; and they're almost certainly going to need to build a system which manages and monitors all these clusters...
Trust me, it would have been way easier to just hook this stuff up to the cloud :-P . I still dream that we will be able to some day.
https://www.infoq.com/news/2017/07/iot-edge-compute-chick-fi...
This would make a lot of sense with something like CockroachDB - if the restaurant was offline, their local data would be preserved. But, as soon as it goes back online, then corporate would have access to all of the data.
>If the leader ever dies, a new leader will be elected
>through a simple protocol that uses random sleeps and
>leader declarations.
Why not have each node self-generate a UUID and engage in some gossip process that ends with the cluster becoming aware that some node's corresponding UUID is uniquely significant, therefore recognizing that node as a leader?I have some really bad memories of "random sleeps" at scale.
I suppose the question might be why not use VRRP itself, but if this works for them and has conflict resolution I don't think it's all that troubling.
https://www.mirantis.com/blog/how-install-kubernetes-kubeadm... has a decent step-by-step. It's mostly just the standard install but they have details on the untaint bit too.
I used flannel as pod networking, as it's really simple. If you want to run app pods on your master node, remember to untaint it. ingress-nginx is probably your best bet as an ingress controller, especially because of the amount of support given it by the k8s Slack.
It is a non-zero amount of work. If it sounds like too much work for something you're throwing together, it probably is. It is generally unnecessary.
There is also "./cluster/get-kube-local.sh" that is supposed to give you a working local cluster. But it appears to be broken right now. Might be worth opening a GH issue for that.
We're replacing a lot of these solutions with "better ways" over the next weeks and months, but I'd be happy to share how we went about it. You can contact me on LinkedIn: https://www.linkedin.com/in/calebrhurd/
The biggest key was that we use RKE for the clustering/certs on bare metal. That's definitely our secret sauce (pun intended).
What challenge is this addressing, what problem does this solve? Is there a problem to solve here?
I do assume there's a good reason for this, but as presented it seems like a very stupid waste of money.
Hah... no, but seriously... we wrote this article for QCon attendees, and gave a lot more context during our talk at that conference. We didn't realize it was going to be on here, otherwise we would have explained the "why" and not just dived in.
What we were trying to solve for was; 1) Low latency 2) High Availability 3) Container based, zero-downtime deployments 4) Continued operations even in an internet-down event
Also, as an interesting side note, the equivalent hardware has about a 6 month ROI if we put the entire load on AWS... granted it would be more efficient, so that's not an entirely fair comparison, but the hardware is unbelievably inexpensive from a cost perspective.
Bear in mind that in terms of cost, this is competing with a person driving to each restaurant and fiddling around with computers for an hour, which is a very expensive process.
1. Why not?
2. Who cares if it's not modern if it does the job?
And they wouldn't even need to make a special app, they could just make it a webapp ergo make a 1-time image with a browser...
It's becoming more common to distribute applications with orchestration software like Kubernetes. The technology around PXE booting is quite old, and mired in enterprise cruft.
> 2. Who cares if it's not modern if it does the job?
Developers love new tech, especially if they can get a Medium post out of it. This doesn't make it a good reason of course, but if this is the tech that more developers are familiar with, that's a good reason.
I personally wouldn't want to learn how to boot 6000 remote machines off built disk images over the internet, I'd rather use the skills I already have around Ansible or learn Kubernetes.
> And they wouldn't even need to make a special app, they could just make it a webapp ergo make a 1-time image with a browser...
I've never been to a Chick-fil-a, but if the setups are anything like my local McDonalds, that's a complex 5 screen setup showing a fluid mix of static images, videos, animations, and applications, not to mention that other stores have different setups/layouts/display types/etc - I don't think you'd be able to _reliably_ do this in a browser. My guess is that it's a multi-screen aware wrapper around video components and web views. That will need re-deploying regularly I would imagine. And that's not to mention the kitchen ordering system, the self-service machines, the tills, etc.
On-site machines totally make sense, smart applications deployed locally, frequently, make sense.
They'd need connectivity to the server, true, but doesn't Kubernetes also need connectivity to the cluster manager?
Consider - how would you handle order taking if the network dropped?
A buddy did IT for a theater for a while - they had a similar problem where they’d lose access to their payment processor regularly. No one ever noticed since the system queued ops locally until the network came back up.
If anything it brings in more components/complexity/headache...
If you could run your pods offline, you could run your software app offline. If you could trust your Kube repository, you could trust Se repository for your App..
What does K8s bring to the mix that is actually useful in solving a problem than the superficial ones?
I have enough problems getting kubernetes to run on a single node with hyperkube that I’m not entirely sure that I want to deal with it in the future.
but if you want to be safe, have a cellular backup network.
seems way simpler than pushing out a k8s infrastructure to every store.
And surveillance cameras, and smart locks on any safes, etc.
https://thinkprogress.org/chick-fil-a-still-anti-gay-970f079...
/s