HNHacker News
TopNewBestAskShowJobs

coffeesn0b

121 karma · joined July 3, 2018

submissionscomments
coffeesn0b··on AWS Web Console Down?
Is this the apocalypse?
coffeesn0b··on Site Reliability Engineering at Chik-Fil-A
Woot woot! (Biased author here, sorry... can’t help it)
coffeesn0b··on Edge Computing at Chick-fil-A
We will open source highlander eventually, but it’s not quite there yet.

We use a VIP that the NUCs share... ie; one of the three will always have a VIP, and if it dies another NUC grabs it. This is a poor man’s load balancer in that sense, because we only have the NUC hardware onsite;

https://github.com/kubernetes/contrib/tree/master/keepalived...

We are also looking at metallb

coffeesn0b··on Edge Computing at Chick-fil-A
Have you worked with k8s yet? A lot of the things you bring up are addressed if k8s is used properly.
coffeesn0b··on Edge Computing at Chick-fil-A
Iot device control, such as timers for food, and cameras tracking food (to name a few). These also output streams of data that is interesting to us, such that we want to exfiltrate them up to the cloud.

I’m not a fan of hype... k8s is one of the few hyped technologies that has delivered on its promises.

coffeesn0b··on Edge Computing at Chick-fil-A
You’re correct, training is centralized but the devices that use it to function are operated at the edge.
coffeesn0b··on Edge Computing at Chick-fil-A
They’re really reasonable about it... you do what you need to do.

That being said, it’s really REALLY nice to mute pagerduty for a whole day every week.

coffeesn0b··on Edge Computing at Chick-fil-A
Credit card transactions, mobile orders, timer synchronization, order receiving (tablets etc), iot devices (cameras, cooking devices) and other things planned for the future.
coffeesn0b··on Edge Computing at Chick-fil-A
Our tolerance requirements for synchronicity are broad enough that we can tolerate blips like this... at the end of the day we are automating away some simple human interactions (ie; fry the fries or track the food), we aren’t performing surgery with these systems.

So far we’ve been satisfied with Linux (Ubuntu 18.04 to be exact) and it’s overall capabilities.

Also, don’t forget that cost at scale is a big factor... doing things perfectly but expensively is not profitable at 2k locations.

coffeesn0b··on Edge Computing at Chick-fil-A
We have frequent internet outages and still need to run compute loads at the edge.
coffeesn0b··on Edge Computing at Chick-fil-A
It’s more about network outages than latency specifically... although in several remote locations latency is permanently low due to slow providers.
coffeesn0b··on Edge Computing at Chick-fil-A
Could you define what is not deterministic about containerized computing?

Examples of things we are controlling are timers for food, cameras that can recognize and track food items, drying machines that will automatically trigger and fry fries.

coffeesn0b··on Edge Computing at Chick-fil-A
If you decide to play around with it, we use RKE to build and manage the k8s environments locally.
coffeesn0b··on Edge Computing at Chick-fil-A
That is epic, and i am stealing that name.
coffeesn0b··on Edge Computing at Chick-fil-A
My coworkers found this one quite amusing.
coffeesn0b··on Edge Computing at Chick-fil-A
Well, it’s a 6 man team so most of our trade offs were for expedience, not the worlds greatest architecture. We cheated and sync’d data using a HA MongoDB setup amongst the cluster. RKE for K8s clustering was a life saver (nothing easier at this point on bare metal IMO), although RKE can be brittle at times.
coffeesn0b··on Edge Computing at Chick-fil-A
The primary challenge was just reasoning with the template and using Helm at scale... ie; what exactly did we deploy on those hundreds of varying clusters?

Other issues included; tiller would sometimes become unstable... version mismatch issues between helm local and roller... lack of a clear, outage free canary deployment... we even found cases where helm would not cleanup after itself during a deployment and retain previous config settings within k8s.

coffeesn0b··on Edge Computing at Chick-fil-A
Ultimately “the edge” will just do it for him :-P
coffeesn0b··on Edge Computing at Chick-fil-A
This is Caleb (SRE on this project). Here’s a similar NUC we use ... this one is like $300... just an example of what we use; https://www.bhphotovideo.com/c/product/1316113-REG/intel_box...
coffeesn0b··on Edge Computing at Chick-fil-A
Here's the abbreviated reason;

Internet connectivity for restaurants stink.

So we put our compute at the edge.

Trust me... as an SRE I would much prefer we shove it in the cloud and call it a day, but alas... no such luck.

coffeesn0b··on Edge Computing at Chick-fil-A
I'd send that feedback through our feedback form... our product teams definitely need to hear that; https://www.chick-fil-a.com/Customer-Service/Contact
coffeesn0b··on Edge Computing at Chick-fil-A
Thanks! We appreciate the compliment! :)
coffeesn0b··on Edge Computing at Chick-fil-A
Interesting - why? (this is Caleb, one of the authors and lead SRE for the project)
coffeesn0b··on Edge Computing at Chick-fil-A
We will have an order of magnitude more cloud datacenters than AWS! :)
coffeesn0b··on Edge Computing at Chick-fil-A
Hey! Caleb here... I'm the lead SRE building the clusters on the restaurants.

Each individual restaurant gets it's own cluster (we aren't fully deployed yet). There are too many network latency challenges and too much immaturity around federated clusters to take the route (unfortunately).

We currently use gitops, which houses all the configs per cluster in a single repo per restaurant (CRAZY amount of corgis!)... we call that Atlas (the repos) and we made a little pod called Vessel that polls and applies configs in those repos.

We're almost done building something called Fleet that will generated and manage those repos (Atlas) at scale. Ie; a UI where we can say "send this version to 10% of restaurants" and it will regenerate and deploy those configs to all the appropriate Atlas repos, which will then get pulled down by Vessel.

We tried doing all this with Helm but failed miserably. Maybe it was us? But templating vs gitops... the choice seemed obvious.

coffeesn0b··on Bare Metal K8s Clustering at Scale
We're still in the process of rolling things out, although I'm not sure I can give a percentage... I think it's safe to say that in the next 6-12 months most (if not all) of them should have an active cluster.

And no, there is not a full POS server on them at this time... it will take us a few years to decompose that monolith, but this is a natural place to put it as we do so.

coffeesn0b··on Bare Metal K8s Clustering at Scale
We deal with a challenge of frequent network outages or latency issues... keep in mind that we have locations out in the middle of nowhere with no QoS. There are a variety of loads at the restaurant that require low latency and high up time. On site K8s clusters were a natural fit for that solution.

Trust me, it would have been way easier to just hook this stuff up to the cloud :-P . I still dream that we will be able to some day.

coffeesn0b··on Bare Metal K8s Clustering at Scale
We haven't open sourced how we do this yet... we have an MVP way of doing it by using Ansible to provision the NUCs, and nmap (please don't laugh!) so that they can find each other on a specific virtual network at the restaurants.

We're replacing a lot of these solutions with "better ways" over the next weeks and months, but I'd be happy to share how we went about it. You can contact me on LinkedIn: https://www.linkedin.com/in/calebrhurd/

The biggest key was that we use RKE for the clustering/certs on bare metal. That's definitely our secret sauce (pun intended).

coffeesn0b··on Bare Metal K8s Clustering at Scale
Caleb with CFA here - random sleeps aren't exactly the most elegant solution... this was our MVP, and we'll definitely be improving on it.
coffeesn0b··on Bare Metal K8s Clustering at Scale
Caleb here (SRE at Chick-Fil-A)... we actually just did it because we thought it would look cool :-P .

Hah... no, but seriously... we wrote this article for QCon attendees, and gave a lot more context during our talk at that conference. We didn't realize it was going to be on here, otherwise we would have explained the "why" and not just dived in.

What we were trying to solve for was; 1) Low latency 2) High Availability 3) Container based, zero-downtime deployments 4) Continued operations even in an internet-down event

Also, as an interesting side note, the equivalent hardware has about a 6 month ROI if we put the entire load on AWS... granted it would be more efficient, so that's not an entirely fair comparison, but the hardware is unbelievably inexpensive from a cost perspective.

Page 1 of 2Next →