I tend to burn through these interviews so quickly that we have time leftover. It was always the easiest part of the FAANG interview process for me.
I tend to burn through these interviews so quickly that we have time leftover. It was always the easiest part of the FAANG interview process for me.
The other complementary skill is resource accounting -- knowing what everything costs in terms of bandwidth, latency, storage, compute, etc and how they interact. This allows you to look at any combination of hardware system and software workload, and quickly identify the resource bottlenecks. Again, if you do it long enough it becomes very intuitive. Old hands can accurately predict the performance characteristics of a software design on given hardware before a single line of code has been written, even if the design is novel, just through resource accounting. The application of these resources involves tradeoffs i.e. you can trade an excess of one kind of resource for another resource that is scarce with clever algorithms and architecture (e.g. classic space/time tradeoffs but more so).
Unfortunately, I don't have any reference material. I learned by doing over a very long time.
A concrete example is cache replacement algorithms for storage. The ideal metric is cache hit rate i.e. the percentage of the time that the storage you were looking for is in the cache -- higher the better. Literature is full of academic algorithms that focus on improving cache hit rates under a variety of workloads.
In many real systems, the cache replacement algorithm is in the hot path. Most "efficient" algorithms in literature either have poor CPU cache locality or thrash the CPU cache, causing significantly worse performance than is offset by the marginal gains in cache hit rates. Knowing this, good cache replacement algorithms are explicitly designed to minimize CPU cache locality/thrashing problems, at the cost of slightly lower cache hit rates because it has higher throughput in real systems. In this case, allowing the CPU cache to work efficiently is more important than a better disk cache hit rate.
In complex systems, the optimal algorithm in isolation is almost never the optimal algorithm in a system. There is a global resource budget. By "second-order effects", I mostly meant understanding how subtle design choices for one part of the system will tacitly interact with the performance of other parts of the system by virtue of how they use global resources. More broadly, this is the discipline of accounting for the hardware resources available to the software at any point in time and identifying conflicts for those resources.
Hopefully that gives you a sense of what I meant.