NNSA, Nvidia to create an open-source Fortran compiler front-end for LLVM
llnl.gov
llnl.gov
It'd be nice to see an LLVM frontend for Fortran; Intel has a stranglehold with ifort due to their excellent vectorization. Hopefully this project can provoke some competition.
Gfortran is OK and ifort isn't insurmountably more advanced, but Intel definitely has the mindshare in the HPC market locked down and they do a lot to promote their technologies in that community. That is what really gives them the vice grip on the market. IBM is really the only other company with the same cachet in that market.
I'm dabbling in satellite image processing and have it working in Python (numpy et al) on a single core, but want to try splitting the work across multiple cores as 80 minutes per image is a long time to wait!
The first step is masking out clouds (morphological erosion + more), and then applying a haze optimized transform to the image data. I'm confused about making code parallelisable when the transform of a pixel relies on neighbouring pixels - how do you remove boundary errors (take extra pixels, then crop off each strip before recombining I guess)?
There must be something in-place to deal with boundary errors round the edge of the image, but so far I haven't grasped it!
Numpy, Scipy, and particularly scikit-image are where the time is spent - my dabbling was originally to try and speed things up, but I've also become interested in understanding what is going on.
Usually the strategy is to break the domain into chunks, with each chunk being multiple cells. Inside the chunk, the processor already has the data for adjacent cells so it can just read it. On the boundaries where some of the adjacent cells are in another chunk and on a different processor, it has to ask for the data... the vast majority of scientific codes here use OpenMP (not so bad, shared memory) or MPI (absolutely miserable to use IMO, but powerful distributed memory) to communicate. There's other tools for the job that are nicer than MPI, but they don't get much attention from HPC people (who tend to be very conservative with new technology, especially "high-level" abstractions).
Inter-processor communication is orders of magnitude slower than querying local memory (especially if that processor is say on another machine), so on top of that there's lots of clever methods of what is essentially figuring out how to slice up the problem domain; examples include Red-Black schemes, domain decomposition and Wavefront propagation.
To answer your question specifically on boundary conditions (and here I'm talking about the edges of the whole image where there's literally nothing in some cells nearby, not boundaries between partitioned regions like I was above)... you just have to choose something. There are three common approaches, which I'll illustrate by imagining you're doing a blur that decides how to blur a pixel by looking at the 8 pixels that surround it and that you're on the far right edge of the image (so the right three neighboring pixels would lie outside the image):
1. Zero - treat the three pixels as black/zero when calculating. Equivalently, treat them as whatever arbitrary value you find works best (e.g. white/1).
2. Reflective - treat the three pixels as whatever value the pixels on the left side at the appropriate positions would be. That is to say, wrap around to the other side.
3. Symmetric - treat the right three pixels as if they were the same as the left three (or the center three if you like).
This is a common problem also when solving something like a differential equation on a discretized spatial grid. In this case, the edge values are determined based on your boundary conditions- method 1 corresponds to vacuum boundaries, 2. corresponds to periodic boundaries and 3. is a specific type of extrapolated boundary.
SCALE is mid-way into migrating to C++. AFAIK though MCNP isn't going to change.
Man this community and it's diverse group of amazing people. I don't there is a single place on the internet as valuable as this community.
Sometimes I just have to pinch myself...
I weep to think of of all the cycles wasted on expensive tax payer funded hardware due to using Python over much faster Intel Fortran. We are talking like orders of magnitude.
I once ported a fellow undergrads astronomy program from Python to C+CUDA and it ran on their personal workstation in less than a day, when it had been using the department cluster for a week before :/
People often hate on Fortran because its old and creeky when it often as fast or faster than C...
Even then, in a lot of cases programmer time is infinitely more expensive than CPU time, so it often still makes sense if Python is the more accessible choice.
I think this is generally true outside of scientific computing, but often not the case here. Esp. if you consider the low salaries of scientists in the public service compared to private industry.
Even in the private sector I have routinely made things an order of mag. faster saving the need to massively scale out. People also forget about total cost of ownership, 10 servers costs much less than 10x what 100 servers require operate when you consider cooling, networking, power, part replacement etc.
But I agree, if you're writing a massively parallel simulation in Python and running it on a supercomputer it's a net waste of money. There's a big trend in the government projects to build libraries like Trilinos or MOOSE, which work kind of along the same philosophy as NumPy - let specialists implement the parts that really matter (in terms of computational expense) in a compiled language and let the scientists glue together those parts in a higher level language. It's a good strategy IMO, especially for technologies like CUDA where there's a huge benefit but also a steep learning curve.
Scientist/engineer detected.
In the standard programming dialect, "code" is a mass noun. "Program" or "library" is countable.
I'm familiar with the terminology, but thanks for the lesson. I spend the vast majority of my time as a developer, actually.
As far as I hear it used, "code(s)" is the standard term for what is loosely any large simulation package amongst people in the numerical community (which includes mathematicians and computer scientists alike, so it's not just an engineering thing).
My favorite old term, which is still used at times, is 'false drop'. (Eg, see http://www.researchinformation.info/features/feature.php?fea... .) The term comes from the days of manual punched cards, using edge-notched cards, where you would stick a needle through a given position deck of cards and lift. The cards with the notch at that position would fall out. Those with a hole would stay. Hence, the cards you want "fall out".
But if you used multi-punch encoding (eg, encode two rare terms to a single location) in order to decrease the number of holes required, then you might get some "false drops".
Nowadays we mostly say 'false positives' for this case.
They had a glass room overlooking the actual computer room where you'd go after handing one of the techs your cards to be entered. They brought the guy up there and he was waiting after handing off his cards and they had pre-arranged with the technician to trip and dump the cards everywhere. He said the Army guy turned white as a sheet of paper and looked like he was going to cry (this would have necessitated many hours putting the cards back in order). Then he told the guy that it was fine and they already had a copy loaded in on the drive.
The usual solution was to take a marker and swipe the deck along the top side, in a diagonal, so they are much easier to sort.
You might recall that Fortran ignored everything beyond column 72. This was where you put the 'sequence number', aka line number, so when you drop the deck you stick the cards the sorting machine and voila, it's ordered again.
Ahh, I see Wikipedia covers this bit of arcana, at https://en.wikipedia.org/wiki/Computer_programming_in_the_pu... .
But there are many stories of the heartbreak of those who didn't take one of these precautions.
Ha! In looking for one I came across http://www.columbia.edu/cu/computinghistory/sorter.html :
> You may have heard the story of the operator who dropped a whole box of cards. Wanting to put things right as quickly as possible, he sorted the cards, without consulting the user. As it turned out, that was the worst possible response. Up until that point, the box had contained a sample of random numbers. –Ted Powell, Dec 2006
I've never seen a punch card in my life, but the term is still around.
The resulting discussion contains a mini-FAQ: http://lists.llvm.org/pipermail/llvm-dev/2015-November/09243...
(both via the LLVM Weekly newsletter: http://blog.llvm.org/2015/11/llvm-weekly-98-nov-16th-2015.ht...)
Proof of concept: https://github.com/harveywi/arpack-js
https://news.ycombinator.com/item?id=10578354
LLVM has a huge momentum now and it's negatively affecting GCC. GCC has been in trouble before but they got out of it by being the best free compiler available. I don't know how much longer they'll be able to sustain without making some changes.
> Between 2000 and 2004, [g95 fortran] front-end was coupled to the rest of the infrastructure of the GNU Compiler Collection. This was not trivial (just as it will not be trivial to couple the PGI front-end to the LLVM infrastructure).
They really don't get that LLVM's modularity is a strength...
Fortran is in a somewhat unique niche- basically anybody who uses a computer uses a library written in Fortran (via BLAS) but very few people are aware of it, especially that it has continued to develop and is actively in usd. There are plenty of places (mostly related to parallel evaluation) where Fortran is state of the art. But speaking as someone mildly competent in it, it isn't something everybody needs to care about. It has its niche and is pretty awful at everything else, best left to its little island of programmers. Haskell is my favorite language and C++ or Python are what I use in a professional setting, but I won't hesitate to break out the good old `IMPLICIT NONE` if it is the right tool.
It likely seems niche to most programmers, but array slicing is such a godsend to certain areas of software that there really is no more natural language to write a code in. There is Matlab, but its purpose is for prototyping. In fact it is a major credit to Matlab that the conversion is very natural.
This niche is why Fortran exists and has such a strong following. And the more modern features are like candy sprinkled on the core features. I could write a sonnet about modern fortran.
There's also issues with wanting to use the old Fortran code in new projects and if the new alternatives don't have a good interface for Fortran, it's a bit of a pain to integrate.
C++ really crippled itself in this regard by never making a standard. I never bought Stroustrup's claims of "but you can write your own" for these reasons.
Granted, there's plenty of fast libraries written in C/Fortran that don't have a Julia equivalent yet, and depending on the overhead of ccall(), you may wish to stick with writing the rest of your code in the language of the library.
My other problem is code generation. I can't generate Julia code yet.
Perhaps I am misunderstanding your meaning here, but Julia has a range of metaprogramming functionality including both Lisp-like macros, as well as "generated functions" that allow custom code specialization based on input type signatures (https://medium.com/@acidflask/smoothing-data-with-julia-s-ge...). These capabilities have been used for DSLs (e.g. https://github.com/JuliaOpt/JuMP.jl) and parser generation (e.g. https://github.com/abeschneider/PEGParser.jl).
Just because they're good alternatives doesn't mean they merit rewriting software that cost millions upon millions to write the first time.
Plus when you're talking about code that runs on 5000 machines, even just a 5% difference is meaningful.
That said, I believe most of the post-MPI HPC frameworks are based on C/C++ with little support for FORTRAN.
Sometimes this community is so cringeworthy.
It sounds more like a cultish anti-sentiment if anything.
We detached this subthread from https://news.ycombinator.com/item?id=10578612 and marked it off-topic.
The only possible reason for Nvidia to do this is because they want to have proprietary compiler addons someday.
Remember that this was a predictable outcome when you're inevitably burned by this.
But, no amount of technical merit can save you when you don't have a free compiler anymore. Don't say Stallman didn't warn you, because he did, just like he warned you about all the other crap.
Remember, when you're getting burned by this, that you chose to make snarky comments on the Internet instead of realizing what the moral path forward was and doing the ethical thing.