How to Program a Supercomputer
cray.com
cray.com
I started with the library everyone uses (g2o), attempted a vectorised c++ implementation, and ultimately decided it would be easier to develop in assembly. I've been taking measurements as I go.
C++ (g2o) ~ 10-16x
C++ (NEON intrinsics) ~ 1.5-2x
ASM 1x
In this case register layout was very important and GCC isn't able to, for example, arrange for two results to wind up in adjacent 64 bit registers and become a 128bit register. This business of combining registers is an ARMv7 wart and not an issue in Aarch64.In other cases, such as some bilinear interpolation code I wrote, GCC was able to basically generate what was more or less optimal.
It also has "item_RICname" while the Cray image has "flagCheckRicname".
RIC = "Resin identification code"? "Reuters Instrument Code"? "Routing Identifier Code"?
shutil.rmtree('Input4RTAvTEST',ignore_errors=Tr
is a reference to the shutil module[1], which I suspect is unique enough to identify Python. fName = os.path.bas
is also a pretty good hint (likely a call to os.path.basename)Interestingly, can't find this code a Google search.
[1]: https://docs.python.org/3/library/shutil.html#shutil.rmtree
Also, for the record, 1x10-9 seconds should be a nanosecond.
My point is, Freq is widely used in reporting CPU speeds. No one "knows" what 3 ns/clock speed relates to. Everyone knows 4.0 GHz. You see my point?
> for people who actually deal with timing of performance sensitive code
While it is — at least in my experience — true that code segments are usually profiled in time / operation, aren't CPUs typically measured in IPS (instructions/sec), since clock speed is now somewhat decoupled from performance given optimizations present in modern CPUs, such as instruction pipelining?
Also I think something that makes it a bit more clear the advantages of being time/ops consider optimizing a video game, and we have the task of trying to get from 60 FPS to 90 FPS and we have say 10 sub tasks that each take say 10% of the frame time, it's hard to say how much optimization I'll have to do off the top of your head. But if instead I say that you have to get from a 16.6 ms frame to an 11.1 ms frame, and you have 10 tasks that take 10% of the frame time, which is obviously about 1.66 ms each, and I also know I need to save 5.5 ms of frame, the scales of which I'm thinking in are all nice and easy to reason about.
I have a follow-up question - what's then the advantage of measuring speed as ops/unit-time (Freq)? How come we don't see 0.25 ns/op as advertised speed on a Processor?
Honestly I don't know. I suspect it's mostly historical and marketing driven. I think CPU clocks are driven by some sort of oscillator, which is likely thought of as a frequency. And ops/unit-time may be a bit ambiguous, since if we try to choose say instructions, not all instructions take the same amount of clock cycles.
Additionally from a marketing perspective, I suspect you get a bit of an hz, more is better, sort of increase of numbers.
Honestly though I really don't know, this is all just speculation. Similarly video games are often "marketed" in Frames/Second, which is exactly the same problem.
If Freq is non-linear function of period, which it is... then Period is also a non-linear function of Freq. They're an inverse of each other.
So, if we're talking about Frequency, then the same argument could be made that Period is non-linear and it cannot be used to compare speed, since Period is 1/f - a non-linear function.
wtf is going on here... my brain hurts.
Because different operations take different numbers of cycles and have different latencies. See http://www.agner.org/optimize/instruction_tables.pdf for a table of such.