McDonalds Event Driven Architecture
medium.com
medium.com
The app is also wildly inaccurate with store open/close times (in both directions -- sometimes it tells you a 24-hour store is "currently closed"), and the in-store employees are often confused how to ring up or serve a mobile order.
But the deals are good. BOGO QPCheese just about every day.
After you order and drive away, I’ll find the app hours later still tracking me and using GPS. I almost always have to force close the app.
Then the next time I use it the previous order won’t have cleared out despite me picking it up, and I’ll have to go in and “cancel” the order.
The coupons are good though. For a week or two it was giving 30% off, and almost always has a 20% off coupon.
I will also say, despite the issues, it's better than most competitors. BK and Wendy's apps are both full of problems as well. I have had multiple accepted and paid orders on the Burger King app, when I got to the store they didn't have the item, weren't accepting mobile orders, or were totally closed.
Yes, I have noticed that here in the UK. I have to keep cancelling the previous order even though I picked it up and paid for it.
You can do that?? I got an order stuck and did delete and reinstall (clear local data probably would have worked too).
Since then, I try to avoid it getting stuck by making sure the app is frontmost with the screen on while I drive through and away.
Seriously, are you iOS or Android, and where do you cancel an order either way?
I've got a similar annoyance with basically every app or website that offers to find things within N miles of me. I live in Western Washington on the other side of Puget Sound from Seattle.
Say the first two options they give me for N are 10 miles and 25 miles.
10 miles is too small. I need more than that to include Bremerton, Silverdale, and Poulsbo, which are the three most likely places that I'll find whatever I'm looking for if it exists over here.
But 25 miles means the circle reaches past Puget Sound and includes Seattle, Bellevue, and even some of Redmond.
For example if I'm searching for Target stores and enter my zip code, their search lists my nearest store first, then 13 stores on the Seattle side (5 in Seattle, 2 in Bellevue, and 1 each in Lynnwood, Everett, Redmond, Renton, Tukwila, and Woodinville) before the store that is actually second closest to me as far as actually driving goes (Gig Harbor).
The article lacks any details material to the event streaming architecture instantiation, and defers the details to a follow-up blog post, so it is hard to draw meaningful conclusions from it.
I always find this phrase funny. Isn't buying one and getting one the normal state of affairs?
Buy implies buying, get implies buying/receiving, possibly at a cost, possibly an offer, all we know is you have it. Buy one get one free is standard here, BOGO would cause customer service staff to go mad from "do you mean 'Buy one, get two?'" inquiries.
That's a feature, not a bug.
So many apps do this and it's really annoying. There are quite a lot of commercial activity on the other side of the river from where I live, and every store's site I visit, or app I use, always recommends a couple of locations that are closest "as the bird flies", but require me to take a ferry, or drive an hour and a half, whereas the store that's actually the quickest to get to for me is just 15 minutes away from me in the next town over on my side of the river.
It's understandable though, because it's so much easier to calculate straight line distance, and I guess in most cases it gives the right, or close to the right answer. I've just accepted that I live in an edge case area for how the majority of commerces calculates these things.
> Hi do you have a mobile order today?
< Yes.
> Inexplicable silence for 10-15 seconds
> (New voice) How can I help you?
< Hi I’m picking up a mobile order (??)
> What’s the code?
Their process of using the app at the drive through is missing a step and it’s annoying.
Sometimes the cashier was too busy to catch anything you said to the auto greeting.
I refuse to play this game. When I go to McDonalds I pay in cash and don’t use the app, it’s insulting to be offered 20% off to allow McDonalds to track my precise location, no thanks.
Anyway, I only skimmed the article, but I had a chuckle seeing the title of this article pop up on HN at all.
not sure what their architecture is
It's a form. What's the problem?
Hopefully further instalments might actually talk about the problems they faced building this out, and their unique challenges.
This is part of their technical blog ( https://medium.com/mcdonalds-technical-blog ), which looks like an outreach attempt to help with recruitment (hackathons etc). Probably someone went to a PM saying "we need to talk about our stack, how do we make it sound sexy and cool?" and this is what they got back.
on one hand aws msk is good enough for an enormous application like mcdonald’s. on the other they need a backup database just to get around it not being available? what’s the real story here. interested to see where this goes
If those messages are discarded because the store can't talk to MSK (or MSK is unavailable), then things like automatic replenishment based on order volume couldn't happen. The store manager would have to do a daily physical inventory count to know how many bags of fries, boxes of drink straws, etc. to reorder.
Kafka is often used for financial applications that must not miss events, so having a backup buffer is a reasonable strategy for those use cases. Things like tracking data is likely not worth backing up due to high data volume and low external visibility when data is dropped.
Guess which services fared better during the last kinesis outage?
First, if you have lots of databases and other applications, then you are talking about a mesh of event busses - which defeats the purpose. Pushing the messages out of the various databases and into a central message bus makes the messages more easily consumable without having to know where they come from.
Second, by writing messages to a PG table first, those messages become part of the update transaction. This means you can post messages at any time during your business logic processing, but if you hit an error and roll back, those messages (which would, presumably, no longer be valid) are also rolled back.
Combine this with message idempotence and you get a very reliable messaging environment.
Just yesterday I was going to redeem a deal and the store was closed to walk-ins, so I had to use the drive thru. I couldn't use the deal anymore because I had used the code already, and you can't apparently switch between walk in/drive thru mode. Luckily the drive-thru was so slow I was able to use the code (there's a timer on the code).
That said, the backend worked great; I ordered on the mobile in-line and the order was in-store once I got to the drive thru order speaker.
People forget how hard and expensive it was 10 years ago to do a realtime architecture. Today, McD gets information from your phone to wherever and down to the stores with maybe a few seconds of delay, so you can use your code at the in-store kiosk. And it has to integrate with their in-store order and payment systems.
I'm disappointed at the article, because it doesn't talk about any of this stuff; it's just a laundry list of AWS services. That isn't the important stuff; the important stuff is really how they got all this legacy (in-store) and new stuff to work together.
Is this a recruitment tactic somehow for such companies?
But I’m always shocked for some inexplicable reason to hear that places like McDonalds and Walmart Labs are interested in solving tech problems. But I mean obviously, of course they are.
Good to be reminded now and then that there are alternatives to contributing to the ad/social media panopticon, right?
Beyond that, with a handful of exceptions (I think mostly airports and college campuses), CFA owns all their restaurants, unlike the typical franchise arrangement that McDonald's popularized. You still have independent operators, but they don’t own the restaurant and have to go through a rigorous selection process to even get one. And then, the franchisee only pays like $10,000 upfront, and CFA pays the rest of the costs to build the restaurant. Individual restaurants control their own hiring and marketing and whatnot, but this is all a very top-down approach that needs to align with corporate.
McDonald's outright owns some of its restaurants, but most are owned and run by franchisees, and usually a franchisee owns more than one location. They license the brand, food, and process, but a lot of the day to day standards are decided by the franchise groups. Whereas CFA operators have less leeway for such things.
That's a significant difference from how MCD stores are operated and it shows.
This is 2022. Businesses already have queueing software that automatically takes images of the vehicle to match with your order. I get that some customers like having the "human factor," but there's definitely room for improvement.
Chick-fil-A is the only chain where you will reliably see a massive line (often backing up to the freeway) on any given lunch hour (or a massive in-person line in New York City) that will still get you out and get you served more efficiently than another drive-thru (or walk up shop) that is 1/4 as busy.
MCD did the double drive thru ordering first, but my local CFA now has two pickup lines as well, with a heated covered roof. The window has been replaced with a wide open door and employees walk the orders to each car. Simple and it works great.
I remember CFA posting some articles a few years back on their k8s setup they use for store and inventory management.
From what I can tell the career path is into Corporate hell.
https://www.uopeople.edu/blog/hamburger-education-inside-mcd...
(Found by googling "AOC Hamburger University -cortez")
So it doesn't matter what architecture is behind McD systems if customer facing software doesn't work correctly.
> It's possible that either word could be used in the context you've provided. Both "discordant" and "dissonant" can refer to things that are unpleasant or conflicting, so either word could be used to describe the feeling of seeing a large company using Medium for its technical blog.
> However, there is a subtle difference between the two words. "Dissonant" typically refers to things that are in conflict because of their individual qualities, while "discordant" typically refers to things that are in conflict because of their relationship to each other. In the context of your sentence, "dissonant" might be a slightly better fit because it emphasizes the individual qualities of the company (i.e. its size) and the platform (i.e. Medium) that are in conflict.
Still worth it though. Their machines are awful - way too big and bad at input
What's impressive to me is that they need all that architecture. Mcdonalds sells under 100 burgers a second from what I can find, their order load is probably bursty, so assume maybe all the orders come in the same third of the day, so 300 per second, and every burger is it's own order... that's still not that much.
One order is more than one operation when you're dealing with everything a place at McDonald's scale, but even if you multiply by a factor of 10 to account for analytics, compliance, etc. 3,000 operations per second? Does that really require an entire Kafka-driven event architecture?
I'm not defending McDonald's or it's architecture -- as I stated elsewhere, the app is far from perfect, and an entirely different architecture could very easily work much better. But I do think you are severely downplaying the number of interactions or transactions required to run an app like theirs.
Taking something like Postgres and sprinkling in some strategic use of Redis would handle their usecase with horizontal scaling pretty reliably...
What it wouldn't do is let you add Kafka to your resume.
I was pointing out, however, that, as is often the case, the initial estimates in a typical "why do they need all this stuff" post, likely underestimated the transaction volume by possibly 10x. Perhaps 3000 or 30000 transactions per second could run on the same system -- I'm not an expert at that scale. But I doubt you'd find any Fortune 100 company relying solely on Postgres and Redis.
> But I doubt you'd find any Fortune 100 company relying solely on Postgres and Redis.
I mean, yeah?
Across every system they use of course that wouldn't be it: what would be generating the data that goes into them? Where would the data that goes in be going out?
I'm simply referring to their "glue" for day to day operations, which here is a pubsub system built on Kafka. Most organizations of a certain size start to pick up some set of technology that new efforts default to being built on top of if only to have access to what everyone else is doing... that's essentially what AWS started off as before it was spun out from internal usage
-
But more importantly, Fortune 100 is a very random pairing of problem spaces. I mean you won't find any built solely on Postgres and Redis for the very obvious reason I mentioned above... but you will find billions of dollars in revenue on even more boring stuff than that. The number of Oracle shops using repackaged technology that makes Postgres look like Cloud Spanner is staggering.
I find the opposite of what you do, that people tend to overestimate what it takes to handle large amounts of data reliably. And I think it's because you need some experience with this stuff to understand why you can't just think in terms of "underestimated the transaction volume by possibly 10x"(hint: 10x can mean 3 million => 30 million).
What happens is people hear that system A is going to need to go from 3,000 to 30,000, then start to architecture the way someone going from 3 million to 30 million should have, and suddenly you're building out a system that's less reliable, more expensive, and just generally worse except for what shows up on resumes.
If anything, if you're at McDonalds scale and still can't find the engineering skill to build a monolith that can handle 30k operations per second, you're playing with fire building a distributed system.
(if you're a nascent startup, then by all means stand on the shoulder of giants and don't sweat that you don't have a full blown cloud engineering org, but that's definitely not where McDonalds should be...)
I am on team monolith, but I also don't see any issue with this approach if you are happy accepting it's caveats and vendor lock in, which they it seems they were.
also, when are those ice cream machines going to get fixed?
Thread: https://threadreaderapp.com/thread/1597983918900510720.html
It's a bit like saying Magento served 1M rps on Black Friday leaving out the small but important detail that each individual store has separate infrastructure and manageable load.
Divide and conquer works, congrats to the Shopify team that their design decisions worked out for their use case. And obviously some parts of the system are still shared but my guess is that they are not part of the monolith.
I’d hate to get that AWS bandwidth bill.
Go (and now Rust) is really only used for very low level services with a high SLA (Like Infrastructure). Almost all business logic is Ruby + Rails.
I don’t think any other company could do that.
Unless you're confounding microservices w/ async architectures and saying to drop asynchronous patterns like this all together.
If the teams you are talking about never heard of threads and only know about microservices, then there is something seriously wrong with their CS education. Maybe they all were hired via leetcode. That could explain it.
I'm not confounding anything. Distributed programming has its applications and uses, but if you don't have a good reason to use it, then don't, and use a thread in a single process for background processing.
What makes it a service soup is the soup of services on the other side of the queue processor. If those components weren't services, the application would be a monolith.
Anyway, there many reasons to organize your code on services, and McDonalds is large enough for them to be perfectly valid. But if you take a closer look, those components on the article aren't the ones that do anything, they are just new queue processors that may or may not finally deliver your messages to the destination. That's an irksome architecture.
The last place I worked with a monolith (~100 developers) put quite a bit of work into making sure everyone didn't step on everyone else's toes. This mostly propagated as optimizing CI and improving test quality (since a single flakey test could derail everyone's build)
As to why "microservices" versus a few "normal" sized services
I'm not sure why it's always "monolith" or "microservices"
Hopefully by redundancies, you dont mean sharing data access logic
You need to lock when you write a shared area from multiple sources with no opinion on write ordering.
But say your pipeline is client-> decorator -> processor -> observer with client publication -> external partner, each input will go into a set of instances different from the previous and next one, and rejoin at the output who will queue and order them. You have parallel heavy work and sequential light result publication. Your simple output must be as fast as the sum of your parallel routes to minimize queueing.
Ofc it s more complex, and I prefer 0 network hop myself, but I work on a large investment bank micro service system and we do not lock, and the component are both simple and complex enough that when one disappear, everything else waits or rebalances, and when it reappears it can catch up automatically, and go on. It consumes large amount of memory to keep a duplicated state in each component and persistence is not guaranteed to be on time (in fact, our persistence layer was 30 minutes behind by mid day, for years, until we dug into the 30yo sql)
... which perhaps just shifts the question to "why do you need multiple independent teams and not just use a monolithic development team?"
POST /lock
This is only a slight exaggeration over some of the stuff I've seen people try to pull.