The EverInitializedPlaces example really stands out. Going from ~1.5M to ~90K apply_effects_in_block calls by changing the CFG traversal is a reminder that the biggest compiler optimizations often come from changing the algorithm, not optimizing the hot loop itself.
It also seems like the new Polonius/trait-solver work is pushing compiler performance toward a more interesting problem: doing expensive analysis only when it is actually needed.
4.57% mean wall-time reduction across 629 benchmarks in two months is pretty remarkable. Great progress.