[1] https://lintool.github.io/bigdata-2018w/slides/didp-part02b....
[2] https://lintool.github.io/bigdata-2018w/slides/didp-part09b....
[1] https://lintool.github.io/bigdata-2018w/slides/didp-part02b....
[2] https://lintool.github.io/bigdata-2018w/slides/didp-part09b....
Another example is lexicographic range sharding, which uses reservoir sampling to compute optimal tablet key split points by doing a constant-space-and-time heuristic sampling over the keyspace.
I used to think MR was just brute force, but it has many levels of algorithms. Probably too many- at some point it because hard to analyze how the system worked because of the various kinds of hedging and recovery strategies.
Secondly, you are not inventing any algorithm there. You are only using algorithm invented by others.
Thirdly, you are only deciding what solution works better.
Lastly, in an interview you have to invent this algorithm in 45 minutes.
None of this involves you to invent a new algorithm. At least not in 45 minutes. I doubt if the person giving that talk himself did it so quickly.