def main():
for i in xrange(10000000):
"%d %d" % (i, i)
main()Does't seem to be copying the result anywhere. Where as the C example is copying the result to memory .. which would explain why it is slower.
def main():
for i in xrange(10000000):
"%d %d" % (i, i)
main()Does't seem to be copying the result anywhere. Where as the C example is copying the result to memory .. which would explain why it is slower.
Replace the constant string with argv[1], and run the code: for Pypy, there is no real difference, but the C/C++/D compilers are unable to optimize.
Compiling of printf strings isn't done, though, because nobody cares about string performance unless they're writing UNIX command-line utils. If you're writing C you're probably only dealing with strings for (infrequent) IO, spending the vast majority of your time crunching away on pointers and integer types (floats in niche cases). In the end, you're probably only printing something out because someone needs to read it, and how fast can humans read, anyway?
This goes just as much for the printing of floating point numbers. (http://www.serpentine.com/blog/2011/06/29/here-be-dragons-ad...)
#include <stdio.h>
#include <stdlib.h>
int main() {
int i = 0;
char *x = malloc(44 * sizeof(char));
for (i = 0; i < 10000000; i++) {
sprintf(x, "%d %d", i, i);
}
free(x);
}Their version was intentionally biased, which is a shame.
It's also worth noting that calling malloc and free in a tight loop where you're always requesting the same amount of memory will be pretty fast. Good implementations of malloc - of which glibc certainly is - will consistently return the exact same chunk of memory to you, and you will be on the fast-path of the allocation algorithm.
What you need to do in this case is look at it and say, "How can I optimize this and still retain the essence of what I want to test?" Your optimizations remove that essence - if you're calling a function that is a part of an API, it will have to allocate and free its own memory. That the code is in a loop is an artifact of the experiment.
I know the aim is to show off the string operation in and of itself so why not leave it there, why the need to bring garbage collection into the mix?
Edit: it's the way the garbage collection has been thrown into the mix I'm having difficulty understanding the need for.
1. The C code is stack allocating a pointer in every loop iteration.
2. Cleaning up the memory without(potentially) triggering Python's GC.
Is the Python GC getting triggered in this scenario? If not, then this is what's actually happening, with the cleanup happening automatically when the process exits:
char x[10000000];
int i;
for (i = 0; i < 10000000; i++) {
x[i] = malloc(44 * sizeof(char));
sprintf(x, "%d %d", i, i);
}
If GC is thrown into the mix, the C code is really: char x[1000];
int i;
int j=0;
for (i = 0; i < 10000000; i++) {
x[j] = malloc(44 * sizeof(char));
sprintf(x, "%d %d", i, i);
j++;
if(j == 1000) {
for(j = 0; j < 1000; ++j)
free(x[j]);
}
}You are arguing that they should have compared against a less efficient C example, which honestly boggles my mind.
There may be other Python-specific concerns I'm missing - my work in this area is in Ruby static analysis - but one other thing is that allocating memory is a side-effect itself. The main side-effect visible to Python is that it might raise an exception for being out of memory - this would have to be special-cased as an acceptably ignored side-effect by any purity analyzer.