a2, c = a2+c, c+2
ran faster than the fragment: a2 += c
c += 2
Context and timings in the post. Speculation welcome as to why this might be true, and your predictions for Python 3 invited. a2, c = a2+c, c+2
ran faster than the fragment: a2 += c
c += 2
Context and timings in the post. Speculation welcome as to why this might be true, and your predictions for Python 3 invited.Usually, in an interpreted environment, I reason about other things: number of database calls, HTTP fetches, that kind of thing. If I need CPU performance, I'll fire up something faster.
Finding the meatiest primitive that does the job in one go is often best - batching, where the batch operation is implemented in native code - but not always, you're at the mercy of the implementer.
You are mistaken - this investigation does make sense in my context.
> If you're having Python bottlenecks, it is the wrong language.
But you are missing the point. The object of the underlying exercise was not to try to make this specific instance of the algorithm run fast, Python was not the bottleneck. The activity was exploring the effects of algorithmic changes and investigating the Big-O performance of the different variants and changes. As a part of that the underlying code was being reworked and modified to allow the larger algorithmic changes to be made, and this was one of those changes. It says so in the post:
> Then I made a small change to the code layout ready for another shot at an optimisation.
As always I ran the timing tests (in addition to the functional tests) to make sure nothing substantial had changed, and there was an anomaly which attracted my attention and piqued my curiosity. Hence the write up.
And I believe you are wrong when you say that this sort of thing is, or should be, only of interest to Python VM Devs. I think that curiosity about this sort of thing is, or should be, important to everyone. Perhaps you will then decide you don't have time to pursue it, or that your time is better spent elsewhere, and that's fair enough, but curiosity and a desire to learn has been a pervasive theme among the best devs I've ever worked with.
So again, you are wrong, this investigation does make sense in my context. Your mistake is in assuming the post is about a Python bottleneck. As I say, you've missed the point.
Languages don't get slower on a single dimension, and computers have multiple dimensions of performance. Hitting a performance bottleneck in one dimension on one language doesn't mean you'll hit the same bottleneck on a different language. Different scales might have different bottlenecks as well, think exceeding L2 cache or finally using all your GPU cores on problems that parallelize with scale. If you're using Python you'll never hit those other bottlenecks because the VM is always 40x slower than a naive native implementation.
Unless, of course, your bottleneck is in native code, which is where the easy prototyping of Python really shines.
That would require using multiple threads, which I'm pretty sure python won't do automagically. And even if it did, trying to do those calculations concurrently would almost certainly be slower due to the additional synchronization overhead.
It would be interesting to know how different python implementations react to that change.
Instruction level parallelism (on common platforms) is an optimization a processor applies, based on the conditions when an instruction is decided, but it isn't visible directly in the code. You would really need to inspect the low level interpretter loop, with some idea of the specific bytecode it was processing and specifics about the processor you're using. Setting up the order of instructions to reduce data dependencies is one of the things an optimizing c compiler does, to help the processor be able to utilize instruction level parallelism, but to compute things with thread level parallelism would require too much coordination -- starting a thread is a many stepped thing, and then transferring the values is too, even if writing to memory addresses, there's hidden coordination there.
$ for STMT in "a2, c = a2+c, c+2" "a2 += c; c += 2" "a2 = a2+c; c = c+2"; do
echo \[$(python3 -m timeit "a2 = 1; c = 1; $STMT")\] $STMT
done | sort -n
[... 0.0547 usec per loop] a2 = a2+c; c = c+2
[... 0.0569 usec per loop] a2, c = a2+c, c+2
[... 0.0608 usec per loop] a2 += c; c += 2But definitely it seemed that python2 and python3 had significantly different performance.
"a2 = a2+c; c = c+2" is still fastest on both 2 and 3, though don't know enough to say why.
Edit: did you compare the bytecode?
I'm hoping someone with a better background will find this sufficiently surprising and be provoked into doing that analysis, writing it up, and letting us know.
I've got it in my boomerang queue.
So everything that comes to hand gets a unique reference (auto-generated) and I put tags on them. I don't try to be complete, I don't try to make sure tags are unique, I just put what's on my mind and seems relevant in a meta field. The system then auto-cross-references by content and tags, so when I put tags related to stuff I'm working on it uses a sort of TF-IDF system to find relevant documents and other reference material and puts links between them.
But sometimes I don't want to deal with something just yet, so there is a tool that lets me send things into the future - I call that my "Boomerang Queue".
Edit: it's been a while since you asked the question, so I tried to find some way to let you know I've replied, but you have no contact information in your profile. I hope you see this.
I've long considered and wanted a system like you've described. It seems like it would have incredible long-term value. Is everything for yours self-hosted? What wiki do you use?
But recently I've been experimenting with zim to see if I can use that instead. It saves pages in markdown in a directory structure, so my external scripts can be adapted to run over them and create the augmenting pages. It could possibly insert the free-text links as well in the background. Just an idle thought.
It's all running on my home machine, but I can connect to it from (nearly) anywhere. It's effectively unusable on my phone, so I have a staging area in which I make notes, and then transfer things later. The act of transferring implies a review of the material, and that's of value anyway, so it's not wasted effort.
For years, decades, I've been thinking of polishing, bundling, and releasing it, but it would be one of a gazillion organisation systems and no one would find it. If they did, I'd then be committed to supporting it, so that's not going to happen. I might, one day, convert these replies into a blog post, but again, limited value.
I worry about not having the discipline to regularly curate such a system. That's probably my biggest roadblock towards just going for it.
Quite quickly the system was useful, and I had fewer decisions to make by using than by not using it, so it saved me effort. That was the key.
[0] ... or file/email/document on the screen