HNHacker News
TopNewBestAskShowJobs

RiotTony

36 karma · joined October 29, 2015

submissionscomments
RiotTony··on Random Acts of Optimization
I think that the Identification and iteration stages can be automated, but the comprehension stage is still a while off. No compiler is going to understand more about the what the context of the program is than a human. For example, knowing that a particle system is transparent enough (or somehow insignificant enough) that it doesn't need to be rendered in the shadow pass is a difficult thing to automatically optimise for.

Compilers are generic beasts, they need to work accurately for all valid combinations of code and data (which is a complex problem in itself) and optimisation is another layer of complexity on top of that. Current compiler optimisations work on predominantly local data and code, as that is all the state that the compiler can guarantee is accurate. If a function called from a parent is optimised for that parent, then it is conceivable that this same optimisation could be suboptimal when called from another parent (especially if the compiler was able to modify data layout).

The other issue is iteration time. This is a crucial part of software development - lowering iteration time boosts productivity immensely. If, as a programmer, you no longer care about performance due to a compiler that can optimise your code to make it run twice as fast but the compile time is hours, you are rarely going to run the optimised build. And you are going to end up hand optimising the debug/test builds yourself to make them run fast enough.

I do think that we could have better tools to help us understand where our code bottlenecks could be - I would love a plugin that somehow coloured my code's variables by cache locality.

RiotTony··on Random Acts of Optimization
You could, but I doubt it would have much of an effect on performance. Modern CPUs are ridiculously good at looking ahead and executing paths pre-emptively, meaning that branching is far less of an issue. On the older consoles, branching was a big issue, so that would have helped there.
RiotTony··on Random Acts of Optimization
Because we are interpolating between keys which aren't evenly distributed in time. What you have there is pretty much what we do when looking up the precomputed values from the table.
RiotTony··on Random Acts of Optimization
Xperf or WPA is an excellent tool for holistically analysing your application (and all the other applications running on your machine at the same time). We do use that tool for looking at lots of different things: file IO, thread contention, server performance, etc. But for a single client running, I find VTune to be excellent. Its very well integrated into Visual Studio, and provides a number of different perf experiments that you can run to isolate the causes of your bottlenecks. VTune is commercial, but you can use Very Sleepy for a free alternative.
RiotTony··on Random Acts of Optimization
Waffles is more than just a profiler. It is a nice, high level interface into our (non-public) debugging API. You are correct that the profiling info there is real-time, while Chrome is post. They all use the same buffers, but just interpret the data a little differently.

There is also other profile info which is gathered by Waffles, stuff like number of visible particles, texture calls, GPU cost per emitter, amongst others, and that information is gathered through another interface.

RiotTony··on Random Acts of Optimization
Hey everyone, I'm the author of this article and I'm glad you've found it interesting. I'll be keeping an eye on this thread, so if you have any questions or comments I'll address them as soon as I can. I can already see some awesome questions here - looking forward to the discussion.