I'm at the point where I only get Raspberry PI-based hardware. This way I know that I'm able to update/or reinstall in a year or two.
358 karma · joined June 30, 2013
I'm at the point where I only get Raspberry PI-based hardware. This way I know that I'm able to update/or reinstall in a year or two.
Most of the on premise IT, with the exception of the Jenkins servers, are managed by a local devops company.
We are a small company with 7 employees.
I usually have to spend a week or so to adapt our builds once we get a system upgrade. It's mostly to hack around weird Cray setups and because we dare to link C++ and Fortran code bases on GPU.
I'm working with a weather model and it takes around 1 hour to build and test everything: We build against 2 super computers, 3 different compilers, 2 different architectures (CPU, GPU) and single and double precision. Building alone takes 30 minutes (more in some cases). For testing we reserve 1x node with 8 GPUs and 9 CPU cores on one machine and 2x nodes with 1 GPUs in a 30 minutes debug slot.
With a pull-request based workflow we are able to push to the master multiple times a day. 40 minutes might be achievable by massively revamping our build mechanism and getting jenkins to store gigabytes in build artifacts for each run. However, I do not think it is worth it, because changes can take multiple days, or in some rare cases months to implement. If people need to wait for an hour for their tests to validate that's not so bad.
Right now, we have quite a elaborate build environment where we first load the GCC build environment to compile the C++ code base. Then we load a Fortran build environment to compile the Fortran code base. I would rather have to use one environment to work in. In the most recent Cray and PGI environments (6.0.3) I have to also add the linker path to GCC version I used for C++. This is to avoid linker issues with the C++11 abi in the newer GCC compilers (somehow, the Cray/PGI environment always wants to link with old GCC libraries). Its not a big problem with our build scripts but annoying nevertheless.
Well, actually it's actually mostly Jenkins that runs the code for us on Daint. ETH, which I'm also affiliated with, is starting to use the model in GPU mode for climate runs on Daint.
What are you working on at Cray?
I meant it as a general trend. There's interest from the community to run the GPU model on their system but unfortunately the official COSMO code is not fully GPU ready yet.
The COSMO model which I'm working on is 18 years old. We ported the compute intensive part of the model to C++ using a stencil library and integrated OpenACC pragmas (think OpenMP, but for GPUs) in the rest of the code base.
I'm not a big fan of OpenACC, because it requires the user to make assumptions on the underlying hardware and quite a bit of thinking to get high-performance code. It is quite time consuming to integrate the pragmas so that the code performs. OpenACC capable compilers used to be very unstable, we regularly got compiler breakages and regressions. It got a bit better since we send the code to the vendors, but we still see regressions from time to time. All (usable) OpenACC implementations are proprietary, so we have a vendor dependency. The PGI OpenACC binaries used to be pretty slow, but it is now almost up to par with the Cray binaries.
The successor model to COSMO, ICON, developed at DWD and MPI is also written in Fortran using OpenACC pragmas (and I think OpenMP pragmas because they plan to use it on XeonPhi). The code is interesting in the sense that they are using a icosahedral grid, which stands in contrast to the square grid COSMO uses: Data accesses are not straightforward. Keep in mind that Fortran does not have easy abstractions for data accesses outside of square grids, so you have to use a couple of macros/functions/etc.. to get to the fields you need. Disclaimer: I have not seen the code myself, but the DWD certainly knows how to write high-performance code, so I'm certain most of the code is optimized.
I'm currently looking for a rust guide that shows me some programming patterns.
For example:
- How to best implement an observer pattern
- Best practices for vector code
- Best practices for tree implementations and how to implement a lambda on top of it.
I'm interested in small snippets so that I can get some initial productive code and progress from there.
Just have a look at the state of the art math libraries in rust and compare it to something like Eigen or Cgal. The C++ code is way more flexible and expressive than the rust code. If you don't believe me, check how the rust libraries handle matrix implementations. Often you will find specialized implementations of 1x1 to 4x4 matrices but no generic n-dimensional matrix code.
Edit: There's also support for youtube channels: http://blog.bazqux.com/2015/06/subscribing-to-youtube-feeds....
Here's the scientific description of the model: http://www2.cosmo-model.org/content/model/documentation/core...
The 1.1km forecast runs on 144 GPUs, the 2.2km probabilistic ensemble forecast is computed on 168 GPUs (8 GPUs or 1/2 node per ensemble member). The 7km EU forecast is run on 8 GPUs as well.