No, a supercomputer won’t make your code run faster (2017)
lemire.me
lemire.me
* "I need a terabyte of RAM"
* "50k samples ran in the blink of an eye. 100k ran in an hour. 1 million took a week. I'm about to do the full 50 million observations, I expect it'll take ten days"
* "This code is always slow... that's a lot of math it does"
* "Profiler? Never heard of such a thing"
* "Big-O? What's that?"
* "I only have ten days' money, and it took me that long. Afraid I can't afford your time to dig into this"
* "AWS will fix this"
Generally I'm dealing with the intersection of two or more of these.I do think, though, that the last paragraph is a bit unfair. Literally everyone I work with is incredibly smart; just their background isn't the same as mine. Things that I find incredibly tough are idiotic to them ["Can't I just take the average and be done with it?" "no"].
So, for most of the people I work with, it's not "common sense"; they literally lack the frame of reference to even know what's awry. I never hold it against anyone, and nor should you.
My favorite example is from many years ago, a guy who had some code that took a month, "because that's just complicated math, it's how long that takes". I ran it through a profiler, and after about ten minutes work, it ran in seconds. I came out looking like a hero.
In my recent work doing financial modeling (options, cryptos, other nonsense), I pretty much just try to cache absolutely everything with some arbitrary 1 week TTL and then continue to cache new data every day. Cache hits cause a 10% bump in the current TTL time delta. Kinda heuristic but whatever, works in most cases.
Any other funny stories? I always love these engineering horror stories. Have you looked at Julia at all? Would be interesting to hear if there's been an uptick in usage from the research community from a veteran research code surgeon!
I was stupidly excited to recently do more work on that same model - I took one of its internal datastructures and turned it from a simple queue into a priority queue using a couple different queueing metrics. In both cases, I had run cases out to a scale where I was able to actually show going from O(1) to O(log(n)) performance - but got beautiful wins because the fastest work is work you never have to do.
Not often I can empirically show traditional CS101 stuff in a real-world chart.
(And you probably mean O(n) and not O(1), otherwise O(log(n)) isn't really an improvement).
I'm actually pretty excited to explore getting his problem solved; it's likely that a terabyte is about the right amount to get him over the finish line, and he's converged on this algorithm precisely because it has fairly horrifying memory needs, but trades off to give a nice big-O runtime.
I can't even imagine the power draw of such a system or the heat generation. Fucking galaxy brain engineering. At that point you need like custom motherboards and special buses, right?
I have no insight or information this; but sometimes you can throw money at a problem to make it go away, and that's literally the best deal open to you. If custom fab solves a problem more cheaply than the 25 senior software engineers it would otherwise take, then custom fab it is.
I’m unfamiliar with the machine specified in your example, but I’ve seen in person HPC clusters with more RAM than that, so I’ll use that to give you some napkin math calculations.
Dell R730 with 24x64gb ECC DIMMs = 1536gb RAM. To get to 256TB, you’d need about 170 of them (aka about 5 racks). Power draw varies depending on a number of factors, but let’s be very conservative and take a worse case and say 750W per 1U server. That’s 127,500W, which is just over 600A at 208V, so say a long block of houses with their A/Cs cranking in the summer. Now as to the heat generation, each R730 puts out up to 2850 BTU/hr, so for all 170 of them, that’s just shy of 500,000 BTU/hr. Continuing the comparison to homes, that’s 5-6 home furnaces worth of heat generation per hour.
I don’t have any photos handy of the 64gb DIMMs referenced above, but here’s a photo of 130 16gb ECC DIMMs (aka just over 2TB):
https://www.dropbox.com/s/ww1fwdn7o9se14z/2020-06-15%2020.39...
If instead of the above based on a real world HPC machine, you based it on current gen Dell R750s which can hold 32 256gb LRDIMMs totaling 8TB per machine, you could get to by with just over a single rack of machines for 256TB, albeit for a much higher price (256gb LRDIMMs cost as much each as roughly the price of a used R730 complete w/ 1.5TB of RAM). The cluster of R750s with that much RAM configured as indicated would be well into the 8 figure range.
3 racks * 40 boxes * 2 sockets * 12 DIMMs per socket * 128 GB = 368640 GB or 360 TB raw. Remove some of that for overhead (kernel, OS gunk, network buffers, whatever code you'd have to run to make it into a cluster), and there's your 256 TB with a bunch left over.
Stuff like this is why I sometimes say that a lot of these startups could be done with "three Xeon racks" if they actually cared about doing it right.
[Note: assumes 2nd generation Xeons (Skylakes), with 24 DIMMs. Ice Lake = more/bigger memory, fewer actual machines.]
- For Java, personally I like the netbeans profiler
- Octave's built in profiler gets the job done: https://octave.org/doc/v6.2.0/Profiling.html
- I use JetBrain's tooling a lot.
- I like the profiler in Rider for VB.NET and C#
- The one in PyCharm is ... fine. Python profiling is ugly, PyCharm does a clean job handling it.
- For R, Rprof is ... fine: https://stat.ethz.ch/R-manual/R-devel/library/utils/html/Rprof.html
- STATA's built in profiler does as good a job as possible given it's a weird language
For various other things, where a profiler doesn't really exist per se: - SAS can be configured to output runtime after each proc. There's no cost, you should do it
- SAS proc SQL has an undocumented _method flag: http://support.sas.com/techsup/technote/ts553.html
- Past "EXPLAIN QUERY PLAN", one of the tools I build ends up doing a simple visualisation of SQLite's explain query plan output for *every* query it runs. There's no way to *not* get that estimate
Those are the tools I can think of that I've used in the last few months.Curious if you ever find stuff written in those languages, as opposed to the gamut you just mention, which are classical "research tools" (language/IDE combinations that let you not need to think about the language specifically.
I do cross paths with a fair bit of FORTRAN, but in my experience, generally if someone's writing FORTRAN, they already know the deal with performance.
Fun fact: When I interviewed for my job [12 years ago], I interviewed with a guy using SQLite and FLTK in a C++ model. I'd been using SQLite and FLTK for some recent personal projects; in the interview, I solved a bunch of performance problems he'd been having for a couple years. I interviewed on Monday, and started working there two days later.
My group has a perpetually open posting and are always looking for people: https://rand.wd5.myworkdayjobs.com/en-US/External_Career_Sit...
[If the link doesn't work, go to rand.org, search jobs, then search for "research software engineer"]
EDIT: That link is to a specific posting. Our perpetual one is going back up any day now, once we get through the machinations.
> * "I need a terabyte of RAM"
A lot of our work involve running algorithms such as PCA and ICA on thousands of biomedical imaging data. It is totally normal for us to allocation 1TB RAM. We have 3 nodes with 1.5TB RAM in our cluster. Aapproximation algorithms do exist and are used for resource optimization though.
I took one look at the code, and changed it. It then took about literally 5 seconds to run the code. Then I added some code that would allow the entry of a page number, and the users could enter that page number and print from that page. So, it would then take always a maximum of 3 hours to print out the output. Then I told the night operator to run this program for the operations department at midnight. So, as far as the operations department was concerned, that brought the time down to zero, and the printouts were on everyone's desk when they walked in the door in the morning.
The manager of that specific department literally came up to me after it was rolled out and was crying hard, and hugged me for about a minute, crying the whole time, as her job was on the line, no fault of her own, because she couldn't get the job done, because she couldn't get the printouts.
There was no way that the most powerful computer in the world that would have got that code to run any faster, because of the horrible way it was written. Going from maybe 18 hours to zero time, on the same exact computer....what is that improvement in percentage anyways? What's 18 hours/0 hours? haha...
You solve the actual problem, but this would have been quicker.
You can do that with an actor model by giving each actor thread affinity, but it's rarely what you want since it's normally a lot more efficient to move the code to the data than move the data to the code. In a factory the machines are heavier than the products they're working on so it makes sense to move the machines to the materials; for most code tasks it's much more like harvesting a field or something where it's much more efficient to move the code to the data.
https://www.usenix.org/system/files/conference/hotos15/hotos...
http://www.frankmcsherry.org/graph/scalability/cost/2015/01/...