Intel Fortran now 20% faster than C, fastest language on Shootout
shootout.alioth.debian.org
shootout.alioth.debian.org
How about Intel Fortran vs. Intel C ? How about gcc C vs. gcc Fortran ?
Hyper-optimized compiler for Intel architecture beats compiler designed to be portable to many, many target architectures? Not very shocking.
I've been doing some tuning work on our parallel dialect of ML, and many of the C programs are fairly well-optimized, though mandelbrot still has some room for improvement.
First, even if we take the conclusion as given, "Fortran" isn't faster than "C". The median of one set of programs, compiled with one Fortran compiler with one set of options and run on one processor, is faster than the median of another set of programs compiled with one C compiler with one set of one and run on that processor. This tells us approximately nothing.
If you only expand your consideration to include the whole distribution of programs they tested, GCC actually looks better than iFort (eg. there is a single program which ran 2x faster with iFort, but multiple programs ran more than 2x slower). Median is a spectacularly poor choice of metric, especially with so few samples in the distribution.
But also: Intel also makes a C compiler. Why wasn't that used? How were these benchmark programs chosen? Etc.
This study says close to nothing about "Fortran" and "C", or about any of the other languages involved.
Sorry, but I disagree here. It's one of the best resources out there for getting numbers on the relative performance of languages, despite the pitfalls. The implementations of popular languages are well tuned (sometimes extremely so), the used compilers and settings are quite sane, and the benchmark programs are fairly well distributed in functionality.
The median is a reasonable choice because it's less sensitive to outliers and was probably chosen for that reason. "Spectacularly poor metric" seems to be an opinion, not a fact.
I admit it would have been better to compare Intel C with Intel Fortran. Why wasn't this done? Perhaps some of the submitted programs don't compile correctly with Intel C, or perhaps Intel C doesn't actually perform well on them. Who knows.
But to say that this is a "terribly done study" because of that is doing it injustice.
Adding to this is that significant slowdowns, the "outliers" that you're keen to disregard, tend to dwarf performance improvements that the median might show (because of the multiplicative nature of these values: if A/fortran is 10x faster than A/C, and B/fortran is 10x slower than B/C, then A+B/fortran is 5x slower than A+B/C).
What's really right is to show the complete distribution of speedups and slowdowns (as one of the detail pages on that site does). If you look at the actual distribution for the comparison of C and Fortran in question, C actually comes out looking much better than Fortran does (which shows precisely how misleading using the median can be).
If you must to boil it down to a single number, there are several choices better than the median. I would be happier with a mean, despite its imperfections.
Comparison of medians is nice because it shows that for approximately half of fairly diverse tasks one language is faster than another, which imo is a much more interesting statistic.
The page the OP linked DOES show all the distribution - that's the beauty of box plots.
http://shootout.alioth.debian.org/u64q/which-programming-lan...
For each programming language implementation, that web page shows 7 descriptive values - not just the median!
As I've already said - The outliers are not discarded. The outliers are shown both in the chart and in the table.
The worst value for Fortran, one of the included values, is more than five times worse than that for C. This suggests (not proves) that the Intel Fortran compiler has worse performance holes than gcc C that show up for some, but not all, of these benchmarks.
You criticise but don't even suggest an alternative, let alone evaluate how well or badly that alternative compares to the median.
That's empty criticism.
I think the distance-weighted estimator would serve the purpose of the ranking better.
http://en.wikipedia.org/wiki/Distance-weighted_estimator
This measure finds a measure of the centre of the distribution that gives less weight to items further away from that centre without discarding them. In this way, the measure is not dominated by outliers, but they still contribute to the final result.
I don't say that this is the best measure, only one that is better. A factor analysis might get at the extent to which the different benchmarks are doing the same thing. Some attempt to find measures of how these activities are represented in larger pieces of code might suggest weights for the benchmarks (there's a literature on analysing loads that might be relevant here). But this would make their analysis more complex and more prone to bias leaking into their analysis, so I don't say they ought to do this; my point is that median is simply a bad population measure for the purposes of their survey.
You have argued against me without trying to defend the choice of median as the measure used for the ranking. Is that denialism?
It isn't any kind of ism.
> What is bad about median is that...
Again, that's what's bad about the median as the sole characterisation of the measurements - but the median is not presented as the sole characterisation of the measurements.
The outlier Fortran measurements are there for all to see.
Some variation on the distance-weighted estimator might work here -
http://shootout.alioth.debian.org/u64q/which-language-is-bes...
Can you find the statement "Intel Fortran now 20% faster than C" anywhere on the linked website?
> the authors acknowledge that it's a "game"
See http://shootout.alioth.debian.org/dont-jump-to-conclusions.p...
Feel free to point out the source code in the shootout where this would somehow not be the case. Shouldn't be hard to find if you focus on the benchmarks where Fortran beats C heavily.
Additionally, the argument about restrict is flawed because it can be added to the C code too, so it can't be an advantage of Fortran.
The impossibility of solving aliasing in the general case doesn't stop C compiler writers from implementing alias analysis.
p2 = function(p1);
Do p1 and p2 alias? Note: function() is in a *.so file, no source available.
How many times do you have all of the source code available for analysis?
Hint: If you want to check for some text on a web page, use page search.
http://shootout.alioth.debian.org/demo2/which-programming-la...