How we built Uber Engineering's highest query-per-second service using Go (2016)
eng.uber.com
eng.uber.com
Also I'm looking for work so if anyone is interested snag my email from my profile.
Hundreds of thousands of queries a second per cpu core on millions of data points.
-----
Please omit swipes like "Really, just poor quality work" and "anyone with even a basic understanding" from your posts to HN. It's great to add relevant information, such as an applicable data structure and a link to a good article on the same topic. But it's not great to put others and their work down, and HN has at least two guidelines that ask you not to:
When disagreeing, please reply to the argument instead of calling names. "That is idiotic; 1 + 1 is 2, not 3" can be shortened to "1 + 1 is 2, not 3."
Please don't post shallow dismissals, especially of other people's work. A good critical comment teaches us something.
Actually, a third guideline is relevant too:
Please respond to the strongest plausible interpretation of what someone says, not a weaker one that's easier to criticize.
The strongest plausible interpretation is not that the engineers lacked a basic understanding of relevant CS—or, for that matter, how to use a search engine, since R-trees would pop up nearly any place you searched about this stuff. The strongest plausible interpretation is that they had some other reason for choosing the implementation they did. For example, perhaps it was efficient enough and they were smartly choosing not to build something more complicated than needed.
https://news.ycombinator.com/newsguidelines.html
Edit: and it turns out that the article explains why they didn't use R-trees.
>Instead of indexing the geofences using R-tree or the complicated S2, we chose a simpler route based on the observation that Uber’s business model is city-centric; the business rules and the geofences used to define them are typically associated with a city. This allows us to organize the geofences into a two-level hierarchy where the first level is the city geofences (geofences defining city boundaries), and the second level is the geofences within each city.
Now there might be a valid argument for this approach if the geofences are changing constantly and reindexing the R-trees would take too long, but in the end they synchronize everything anyways, and the R-tree could easily be generated on another node, serialized and then unserialized asynchronously before swapping to an updated tree.
> Geofence lookups are required on every request from Uber’s mobile apps and must quickly (99th percentile < 100 milliseconds) answer
For a 100ms total budget, a ~50ms 99th percentile for a single microservice doesn't sound like something to boast about.
Also keep in mind how percentiles compound when you have more than one service involved in serving a customer request. For example, let's say it takes 5 internal requests to serve an external customer request and each of those services measures latency SLAs at the 99th percentile. The customer request may only finish inside the SLA 95% of the time (99%^5)
Agreed, which is why the most surprising part of this article for me was this phrase: "In our main data center serving non-China traffic".
The transit latency to the one main datacenter in the (non-China) world is a lot more significant than the server time here. (Over 50 ms just to cross the US; compare with their server-side latency numbers of 95%ile < 5 ms, 99%ile < 50 ms.) If they're serious about latency, their deployment is holding them back much more than their choice of programming language or algorithm.
Or maybe the client is not the user's phone, but some other service running within the same datacenter?
Given that the service probably did not have a 100% success rate, and almost certainly had timeouts, the "max" would also likely be at the timeout.
You filter out the 504s, just the same way you would analyze it on a per-route basis for those metrics.
Quantifying things like service latency isn't a one-size fits all thing. Every service has its nuances and use cases that make it more meaningful to measure 99%, 99.9%, 99.99% or something else.
.... just don't measure average like I've seen naive junior devs do. Average latency is the worst of all metrics to use as it will include all the outliers at the very tippy top of the spectrum and basically render the metric meaningless.
Personally I tried to push it for several years, but most Java developers (and I mean 95% in my particular case) seem to start resenting it over time.
It's hard to write good services in Vertx, mostly due to its asynchronous model combined with Java's verbosity and boiler plate.
Many teams have junior developers and IMHO it's simply not safe expecting them to write production grade services (albeit simple) mostly on their own. It can be done, but it drains the rest of the team. A more productive Java team would have used something like Spring Boot.
Also, with any technology, there are some gotchas .. and with Vertx they are much harder to figure out. In the end we changed to another Java framework.
As an interesting addendum, I found this to be a really interesting resource when comparing languages at a very base level, as it is very well written and you can actually look at the source code/ thesis papers for all of the implementations: https://github.com/ixy-languages/ixy-languages
I don't think thats true.
That said, Go's focus may be the right one for many use cases. On a high throughput service where you would assume you want to optimize for throughput, 100 ms pause times can wreak havoc because they're unpredictable and can cause work queues to explode and such. This isn't easily mitigated by load balancing. Whereas "less efficient" GC is at least predictable and you can just add a server to balance that extra work.
Every object in Java has something like three words of overhead; also since it doesn't have value semantics (yet) objects are typically allocated out-of-line, so an ArrayList makes a linear number of allocations, whereas a Go slice makes a constant number. Plus the binary sizes are typically much smaller; not really an expert but our observataion at work is non-trivial Java services allocate a lot of memory through classloading / JITing whereas a Go binary will typically be very small.
Basically, agreed on the GC tradeoff, but the higher memory footprint of Java mostly comes from other areas.
> Better memory usage always implies worse performance
That isn't really true. Better memory usage implies better cache-friendliness (and possibly better locality too).
C, Java, C# and Go all appear at the top there so it's basically a wash, +/-5%. I still think the main perception of Go as "the fastest option" on places like HN is because people are coming from some of the slower languages (Javascript, Ruby, etc).
Footnote: I don't think there's anything wrong with using a slower language, a lot of them have higher productivity overall. I also just happen to prefer programming in Java, so it's overall the best choice for most jobs for me.
Which would be Java in this case.
So few application I own have about 200 LOC of business logic and 10K LOC Spring Boot fluff along with some 60 jar files which I am not sure what they really do.
But enterprise programmers making internal software have little incentive to make their programs sleek and fast because they 1) have a captive user base who are forced to use their bad UIs, and 2) there aren't many users so you don't really need to optimize for minimal server use.
From the eyes of management, a good enterprise programmer can take a request from start to beginning and fulfill a business requirement in a reasonable amount of time. Architecture, UI, future maintainability, speed, security... that's not really on their radar at all.
There are plenty of managers who do not understand or respect feature creep and technical debt, so spending time refactoring is useless and opaque to them. Developers have no incentive to write good code, and bad developers who can talk a good game get hired on, so the codebase spirals into an unmaintainable sea of crap.
According to management, if you solve business problems, you're valuable to them. So if you care about good code, it's an uphill battle fighting for it against the other entrenched developers who really don't give a shit.
These programs usually happen to be written in Java because there are tons of Java developers out there, which is great if you don't live in a tech hub.
> High developer productivity. Go typically takes just a few days for a C++, Java or Node.js developer to learn, and the code is easy to maintain. (Thanks to static typing, no more guessing and unpleasant surprises).
This is why the rest of the world doesn't use JS for everything.
> There is a lot of momentum behind Go at Uber, so if you’re passionate about Go as an expert or a beginner, we are hiring Go developers. Oh, the places you’ll Go!
If they did it in Java, they wouldn't have to recruit for programmers quite so hard -- they could pull them out of a hat, and fairly easily get a few wizards with two decades' worth of experience in it.
Another thing of note: the python side of the team could regularly move across and help with bugs or issues in our services and jobs, but the reverse was rarely true. In the 2+ years I had that team there was one person hired for the services side who could do it, and he preferred Go to either Java or Python.
As someone who very much did not develop (as a person/programmer) that way, it seems baffling from the outside, but there's this whole world of programmers who work like that until they retire or are promoted into management. It seems bizarre to me but it must be working for them. Hell, they might even be the majority of all programmers. Bigcos, particularly the non-tech ones that nonetheless employ lots of developers, are full of such people.
I try to learn new things, but I don't do much programming outside of work. I could learn Node.js, for instance... but I've got other things I want to do.
Is that kind of what you mean?
Some folks, you say "do you think you can do [thing you're basically familiar with] in [language and platform you're not]?" and get a "yeah, probably lemme check it out... cool, compiler and language support's installed, see a couple tutorials here, I'll poke around and get something doing [subset of thing you need] then get back to you on timeline" and very likely it works out fine.
Others, you get "uh, I do (Java, .net), I don't... I don't understand what this is, is it a JVM language? Is there a jar for it?" and it's not really worth pushing any harder. And maybe some of them are deflecting because they just can't be bothered (mad respect), but most seem genuinely nervous and out-of-their-element at the mere suggestion of doing anything but their Java or .net thing (usually it's one of those) they're used to.
Though, again, the latter seem to get along just fine, career-wise. It's just a different sort of path, I guess. Seems really weird to me, but there it is.
It's been this way ever since I can recall. "Java's so fast you probably can't tell the difference most of the time". Well, OK, but, I can. So something's going on here.
FWIW, the JavaScript samples were also pretty bad because it was before async/await, and they'd all get twisted into callback hell waiting for IO.
From pg: "...you could get smarter programmers to work on a Python project than you could to work on a Java project." [1]
I feel this article is about as exciting those Java article describing 'Hello World' http server running under 256MB memory as earth shattering.
What?
like the switch from postgres to mysql.
I'm still clueless how you can have so much money and one of the biggest engineering team, but still can't correctly engineer your stuff. I mean, everybody makes wrong decisions or errors in production code.
One of the best ways to create yourself some work is to not choose an already-proven path to a solution, but invent a new one just for the sake of inventing a new one. Of course that's not how this kind of doing is justified - the justification is usually "the proven path does not scale to our needs" or "by using a special approach adapted to our needs we can be more efficient" or "the proven path is too complex, we can get by with something simpler and easier to maintain". Which might actually all be proper justifications, it's just that you should have some hard proof for these statements, like benchmark results of a comparison of different approaches. That part often gets skipped, which is actually ironic, because doing extensive evaluation and benchmarking and implementing different approaches first before choosing one for production actually serves quite well to create even more work to do.
More seriously: they raised tens of billions of dollars to make a ride hailing app and research self-driving cars. Throwing more money at a problem doesn't solve it faster, but it does pay for hiring people, and headcount is seen as a proxy for doing stuff.
When a measure becomes a target, it ceases to be a good measure.
https://en.wikipedia.org/wiki/Goodhart%27s_law
Some big start-ups[0] solve this conundrum by investing in adjacent companies to find the solutions they're paid to find. Outsourcing! This still has limits because there's so much money flowing around right now. Tossing a few million at a company with hundreds still won't solve the problem faster.
[0] Start-up definition for this post: a company that has taken funding but hasn't yet found a sustainable business model. That's how Uber is still a start-up with more money than most companies make in decades.
We do have ES but it's operated for log search, not production critical paths.
https://www.cybertec-postgresql.com/en/beating-uber-with-a-p...
For instance their demo doesn't update anything, Uber updates location every seconds, I'd like to see how PG behaves when you rebuild index thousand time per seconds.
They used 40 machines to serve New Year's Eve traffic. Perhaps with a better algorithm, they could have got away with one.
(Possibly not, because there's still a lot of HTTP and JSON munging work to be done, and the network card becomes a bottleneck at some point)
Discussed at the time: https://news.ycombinator.com/item?id=11205776.
Why wouldn't this be dogshit slow if they used all of the city geofences at once? I would think that first they would scan the country geofences, then the province geofences, then the city geofences, etc...
However, and I believe this is where the animosity in the comments is coming from - given the elitist (for lack of a better term) attitude of these engineering types at these orgs (think of the poster children of the Valley), this is pretty...lacking. I mean, the part about using the Read/Write lock on the second attempt and instead trying to go with an installed package just screams Node.js, and honestly made me chuckle. I guess Leetcoding and Production Engineering really are different things. I genuinely expected more.
This is key point.
They delivered value to the business. That's the only thing that matters.
But this bit isn't true during the interview process.
Except this is an engineering blog post, so the engineering part actually matters.
And, as demonstrated in the article [1] linked around this discussion, there were better engineering approaches that would have been even "better for business", as being more efficient means lower costs per transaction and/or higher throughput.
[1] https://medium.com/@buckhx/unwinding-uber-s-most-efficient-s...
Should they have searched for a better solution instead of implementing the one they found? They could have spent some time researching, but you can always miss something. It's better to err on the side of delivering something now with a not-so-good solution than constantly searching for a better one.
I say this as software developer who is obsessed with efficiency. I'm starting to turn around and focus more on just delivering.
The entire reason they posted the article is to brag about having accomplished intelligently. It's relevant if their approach was actually not so intelligent.
"We encountered a standard problem and applied standard solutions" is like "dog bites man". It's not what they were trying to say with the blog post.
>Should they have searched for a better solution instead of implementing the one they found? They could have spent some time researching, but you can always miss something. [...] I'm starting to turn around and focus more on just delivering.
I think the critics point is that the efficient way was probably also cheaper than what they did, and would take the same time to implement, and have lower recurring costs. It would have just been a matter of using off-the-shelf tools and not reinventing the wheel because that wheel is "complicated" and "obviously our case is special". (Someone did benchmarks, and their case is not special.)
You're right, there is a danger to what-if-ing everything and being stuck in decision paralysis. But the clear subtext is that they merit some kind of admiration for how well they did. If that subtext is wrong, it is worth pointing out.
Although hardware is relatively cheap, 170 QPS on 40 servers for this type of query is astonishingly horrible even from a Business/Product Engineering standpoint.
Boasting this as "highest QPS* engineering achievement is just awkward. It may be better suited for an article on how throwing hardware at problems is cheaper than hiring engineers.
Sure, but could there have been ways to deliver that value with lower maintenance costs and shorter lead times?
If all you care about is gross revenue and never focus on margins and cost of goods sold then you can justify any project as "adding value to the business".
If the CPU is the hugest bottleneck, the best answer is not to optimize the algorithm by going to a lower level language, but rather to invest in a different architecture like GPGPU or FPGA.
For example, this paper shows a significant speed-up for PIP( Polygon in Point) algorithm, going from 15hs (CPU) to a mere 11sec (GPU) in task load-time. https://pdfs.semanticscholar.org/1e51/e3c681e1afc908a41ac253...
Why not do it in a different language? Having Go and Node around is awesome because we can change the language. Don't you get bored to use just a single language? I'd get so so so bored if I didn't change languages every year. Glad that uber offers Go jobs.
It's a balance for sure. I know it's cool to hate on Java, but I actually like Java for it's predictability, and my deep knowledge of it, and when using Intellij, it's a breeze for me to program in.