1,396 karma · joined April 7, 2017
www.benoitrostykus.com
With a small bounded compute budget, you're going to sometimes make mistakes with your router/thinking switch. Same with speculative decoding, branch predictors etc.
Having worked close to the recsys folks at Netflix, I can tell you that this statement couldn't be further from the truth.
> (I'm guessing this is a CM4/CM5) is a disaster for a camera board. Nobody wants a 20s boot every time you want to take a picture, cameras need to be near instantaneous.
You can boot an RPI in a couple hundred milliseconds.
Truly mind-boggling times where "here is the empirical proof" means "here is what chatGPT says" to some people.
Isn't competition in free markets something Republicans believe in anymore? Because forcing Americans to buy inferior locally-made products at a premium through artificial restrictions surely isn't that.
Free trade and globalization are also a pacifying force, by creating mutual dependencies between countries.
Protectionism doesn't work.
One thing I was pitching internally when advocating for this platform is that when you have the scale to run it for the economics to make sense, you can reclaim some of AWS margins instead of having your cold tiny VMs subsidize other AWS customers higher perf. If you run the multi-tenant platform yourself, you can oversubscribe every app in a way that makes sense for your business and trade latency or throughput of software for $ on a per-container basis, so you can make much more granular and optimal decisions globally. VS having each team individually right-size their own app deployed on VMs and sharing CPU caches with randos.
I remember once at Netflix we investigated a weird latency issue on a random load balancer instance and got AWS involved: it turned out to be a noisy-neighbor on the underlying VM that gets chopped up into multiple customer-facing LB instances.
While this might be true for the governments you have personally experienced, this is far from being an aphorism.
[1]: https://justingarrison.com/blog/2024-02-08-fargate-is-not-fi...
Beyond "kernel programming is hard", there are a few other reasons why it made sense for us:
- observability & maintenance: much easier to implement and ship this type of changes in userspace than rolling out a kernel fork. We also built custom AB infra to be able to evaluate these optimizations.
- the kernel is really good at making reasonable decisions at high-frequency based on a limited amount of data and heuristics. But these decisions are far from optimal in all scenarios. In contrast in user-space we can make better decisions based on more data (or ML predictions), but do so less frequently.
> DBRX uses only 36 billion parameters at any given time. But the model itself is 132 billion parameters, letting you have your cake and eat it too in terms of speed (tokens/second) vs performance (quality).
Thierry Breton is another name that comes to mind.
Andrey (who works for Netflix) drove the effort. Chatted with him about it.
I have nothing against the EEVDF algorithm itself (in fact I like it) and I dislike CFS very much. But I dislike the current development process of the Linux scheduler even more. Proper quantitative benchmarks of CPU schedulers are missing, which is why CFS ended in the sad state it did, where hundreds of patches were submitted to fix random edge cases over the years. What makes you confident that the initial EEVDF Linux implementation won't suffer the same fate, given that the development process hasn't changed (single kernel dev implementing it and running micro benchmarks)?
Most real-world combinatorial problems can be solved relatively well. This is why Gurobi is in business, Alpha Go exists, or why Amazon and United Airlines still manage to practically solve their resource allocation problems.
[1]: https://netflixtechblog.com/predictive-cpu-isolation-of-cont...
I'd take a patch of CFS and its millions of broken knobs from Google over newly released EEVDF any day, because I trust scheduler AB testing by Google over millions of machines and every single scheduling pattern under the sun way more than whatever synthetic micro-benchmark a single kernel dev (as competent as they might be) ran.
If you're interested in quantitative analysis of schedulers & tooling around it, these 2 projects are very interesting:
https://github.com/google/schedviz
https://fuchsia.dev/fuchsia-src/concepts/kernel/fair_schedul...
I'm surprised you've made a general statement out of a company you've never actually worked for.
For engineering roles, work life balance at Netflix is mostly what you want it to be.
I refuse to interview at places that are Leetcode-heavy. Never prevented me from getting 7 figures offers at FAANG-type companies.
Similarly when I interview candidates, I don't ask them Leetcode questions either. I like to get a sense of whether or not a candidate can reason in semi-unknown territories and has good intuition, not if they can parrot a textbook.