Some 15 minutes of work improved CPU usage of my team's biggest fleet by ~40%. Considering we scaled up to 1500 c3.4xlarge hosts at peak in NA alone on that fleet, those 15 minutes kinda made my month :)
One thing to note once you eliminate the easy pickings is that as you go higher up the call graph, the profiler visualization is often misleading. There may be sections of code without safe-points, and stuff that appears wide on the flame graph may just be getting blamed for adjacent code that doesn't have safe points.