I can sell a 2x slowdown to my boss if I can show that that's the only cost of a significantly more elegant and productive language, but at 100x that's a much harder sell.
I can sell a 2x slowdown to my boss if I can show that that's the only cost of a significantly more elegant and productive language, but at 100x that's a much harder sell.
1. Cache performance and locality may differ considerably for large applications. For example, a microbenchmark may perform well because it can monopolize the L1 cache in a way that is not possible for larger applications.
2. Aggressive inlining, a common factor in microbenchmark performance, may not scale well for large code bases. The tradeoffs here can be rather complex: https://webkit.org/blog/2826/unusual-speed-boost-size-matter...
3. Microbenchmarks often avoid abstraction in order to achieve high performance. This can create unacceptable software engineering costs.
My one complaint is that there's no benchmark that measures FFI performance. Realistically, if you build a system in Python or Ruby - you're going to be dropping down to C for your hot spots. And so scripting language performance on all these compute-intensive tasks is somewhat irrelevant, you really want to know how much overhead you'll incur crossing the scripting/C boundary (which, in my experience can sometimes be large enough that it wipes out all the gains of coding in C in the first place).
And to make things more complicated, the performance increase is generally a function of the time spent on optimizing the problem, which is dependent not only on any innate speed, but also the expressiveness of the language and the programmer's familiarity with the language. Given that you don't have infinite time, there are further tradeoffs here.
Which programming language implementations are doing that to your knowledge?
"… Kernels … Toy programs … Synthetic benchmarks … discredited today, usually because the compiler writer and architect can conspire to make the computer appear faster on these stand-in programs than on real applications."
http://benchmarksgame.alioth.debian.org/why-measure-toy-benc...
My point here is not that this is being done (I'm generally assuming that language implementors have better things to do than to pollute their standard libraries for a benchmark game); I was making a different point, namely the inability to do algorithmic improvement, and mentioned that possibility for the sake of completeness.
I think your point is more that "Don't optimize away the work." is antithetical to what we do.
This is just not what a lot of computationally expensive problems look like in practice. As a simple example, any practical solution for an NP-hard problem will be full of tradeoffs; often you just want something that's good enough and then you get to choose and adjust an algorithm for your particular problem space to find the sweet spot between time complexity and quality of the solution.
The benchmark games also have fairly simple and obvious data structures; most of them just deal with arrays and strings. Real-world problems often require you to make difficult choices about representation (where one is optimal in some situations, another in a different set of situations, but you handle all of them). Example: adjacency matrixes can be represented as bit matrices, integer/float matrices (if edges can have weights), or associate arrays of sets, to name just a few implementation options. Depending on how your graph is structured (sparse vs. dense, connectivity) and what algorithms you require, one or the other can be optimal. Clever choices can make orders of magnitude of difference for performance.
I'll offer you a concrete example: computing the factorial of large numbers basically has three well-known algorithms: (1) naive multiplication, (2) divide-and-conquer, (3) prime decomposition. Each subsequent approach improves performance over its predecessor, but is also increasingly more difficult to implement.
Now, it so happens that this particular problem is well-researched, so you can look it up, but you encounter similar problems all the time where you can't find them on stackoverflow or in the literature. And then you have to consider tradeoffs between the time spent researching better solutions, implementing those better solutions, and the time gained from the increased performance.
Many of the pi-digits programs use GMP.
https://benchmarksgame.alioth.debian.org/u64q/performance.ph...