My experience with the computer language shootout
alexgaynor.net
alexgaynor.net
Unless you're going to allow every language to converge on simply blasting some bytes into RAM and executing optimized machine-language code and watch all the languages cluster to the exact same location give or take how long it takes to load the machine code, you have to draw some lines about what's a "real" implementation and what isn't. If I were running it, I'd be even more strict and insist that the solutions must be "idiomatic"... and what's idiomatic and what isn't would also be arbitrary, because there is no escape from the "arbitrary". Yet the end result is useful. Not totally determinative, but it's not anywhere near "useless" either.
You're right that there's a lot of nuance here, and perhaps there's no better solution. It's not even clear to me if there's a meaningful question to be answered. "This ocean's warmer than that one? What part? In summer or winter? Did you measure near a blue whale's fart?"
I can work on only one thing at a time, but I can complain about many things at once. I think both functions are useful.
I wasn't even aware of this particular problem. Now this thread's taught me (and others) to utterly ignore the shootout. That's useful even if nobody builds a replacement.
I'm also disappointed that you respond to a rebuttal by switching your argument, without acknowledging whether the rebuttal is correct or not. I'm going to stop responding to this thread.
>> "It's also not possible to send any messages once your ticket has been marked as closed, meaning to dispute a decision you basically need to pray the maintainer reopens it for some reason."
Truth: There's a public discussion forum!
I'm disappointed that Alex Gaynor never mentioned that the first problem with his program was a bug in CPython, but instead kept that to himself for his blog.
I'm disappointed that Alex Gaynor never mentions in his blog that the next version of his program didn't work on x64 - it hung and timed-out after 1 hour.
Joseph La Fata contributed a Python pi-digits program the same week - his program worked first time on x86 and x64, on PyPy and CPython and Python 3 and only used ctypes to get to GMP.
2 days ago Joseph La Fata contributed a Python spectral-norm program - his program worked first time on x86 and x64 on PyPy and CPython and Python 3.
Do you see the difference yet?
What do you think the blog entry "My experience with Alex Gaynor" would be like?
I'm not going to fact check every story I read. When I actually do something that requires choosing a platform I may look at the shootout. Or I may just try a few little programs myself.
You may think I'm being unfair, or moronic. But I suspect most people are like me. Even you, when you don't notice. PR feeds on the interested-but-uninvolved.
I find this thread tragic. People could have seen the shootout's side, but now most of them will remain uninformed because the herd passed through while you were asking rhetorical questions. It's the shootout's loss.
Because more than one language implementation was measured for that language but only one language implementation was measured for the other languages.
If only one language implementation was shown for Ruby it would be Ruby 1.9 - not JRuby
If only one language implementation was shown for Lua it would be Lua 5.1.4 - not LuaJIT
If only one language implementation was shown for Python it would be Python 3 - not PyPy
PyPy and LuaJIT are being treated more favourably than other language implementations.
A lot of the problem is in the questions they're designed to answer versus the questions people use them to answer.
For example, if I'm comparing Python and C, I typically want to know "how much slower would my program be in Python?", not "how much slower is my program in Python if I spent so much time hyper-optimizing it that I might as well have written it in C?"
But the test cases usually try to answer the latter, not the former.
http://shootout.alioth.debian.org/u64q/program.php?test=spec...
http://shootout.alioth.debian.org/u64q/program.php?test=spec...
My experience is that this is the sort of thing where everyone rags on it, but no one actually attempts to provide something "better".
Former cop Rory Miller writes about this in his book, the police experimented with BJJ and found it useless. Why? Because in BJJ you pin your opponent on his back because it makes a better show for the audience, but as a cop you always pin your opponent on his front so you can handcuff him!
Also, I think that we all know enough about programming and languages and their many uses that we can talk directly about it, rather than about an analogy.
And they're going to do benchmarks.
So you can either complain that they're not good, or you can try and improve them.
Yes, that sentence is literally correct. But it sounds like it's saying one option is not useful. And I still haven't heard a single reason why reviews of benchmarks are bad.
Responding to a criticism with "those who can't do criticize" is super, super boring. It's been done to death. You're just tarring all criticism with an overly broad brush. If it's bad criticism why is it worth responding to? And if it's plausible criticism why aren't you focusing on the actual details?
You'd think there'd be some kind of statement about that?
http://shootout.alioth.debian.org/help.php#why
> the questions people use them to answer
You'd think there'd be some kind of advice about that?
http://shootout.alioth.debian.org/dont-jump-to-conclusions.p...
> I typically want to know "how much slower would my program be in Python?"
And we should all know the answer - It depends on how you wrote your program in C and it depends on how you write your program in Python.
Something can be well designed and yet still be misused.
I'd imagine the implementation varies across the other languages by more than just syntax.
The point that worries me the most, is the amount of microoptimisation being applied to these "benchmark" programs, making the results more or less pointless for real-world use.
For the moment, they all are being measured, so it's interesting to see that programs written for CPython might perform badly with PyPy, and programs written for PyPy might perform badly with CPython.
Is libc.write a great example of programs written in "Python"?
Almost all other languages can use byte arrays, when they are the appropriate data structure for the job. The C submissions make heavy use of GCC extensions, Haskell gets to use mutable (OMG!) byte arrays and Free Pascal has about as much in common with Wirth's Pascal as the name.
But Python and Lua are not allowed to do that? Apparently not all languages are treated equally. Dismissing submissions by resorting to a flawed definition of 'standard' and then suppressing further debate is really lame.
I contributed almost all of the Lua programs to the shootout, but I do not feel particularly encouraged to continue contributing any programs.
"It's not used to do any crazy hackery like the LuaJIT one, just to access the libc write() function."
Obviously your opinion wasn't suppressed on proggit.
Obviously your opinion wasn't suppressed here.
And nothings been done to stop you posting in the discussion forum or commenting in the tracker.
'course, it'd be hard for LuaJIT to do much better in the shootout than it does now. It's beating C#.
But they are Lua programs measured on both the Lua interpreter and LuaJIT.
(Programs that rely on the LuaJIT only FFI library and won't work with the Lua interpreter are shown separately.)
Take the measurement scripts (download from the Help page), measure programs and publish your measurements.
Still, not using SSE when it's available is dumb and perhaps all the other languages need to start playing, too. A little tricker in Python, of course....
Some cases aren't hard to pick up (e.g. bulk operations on big arrays) but others require trickery of the kind that compilers usually don't (or couldn't) have.
This isn't made easier by the notoriously non-orthogonal nature of the SSE integer operations and the rather limited number of ways that you can get in and out of SSE-land (to, say, affect a conditional or get something into a GPR).
That way you can get a single comparison and can replace a pcmpgtb and pand with a single subtract. Then switch it to SSE2, unroll and you're good to go.
Alternately, http://www.azillionmonkeys.com/qed/asmexample.html in section 11 ("Converting Uppercase") contains a brainsmashing version of this entirely in SWAR ("SIMD Within a Register"), which could be adapted with a certain amount of pain (largely due to the absence of a double-quadword bitshift in SSE2, which is retardlepated).
At the same time, I can see where the guy is coming from with ctypes.