Supercomputers for Weather Forecasting Have Come a Long Way
top500.org
top500.org
Currently the most powerful weather forecasting computer is the UK Met Office
machine, a Cray XC40 machine that delivers 6.8 petaflops and sits at number 11
on the TOP500.
But this does not run the production forecast, that is done on one of two smaller machines which are capable just over 3 petaflops.Source: I work for Cray at the UK Met Office.
Edit : Answered while I was posting this
Many people think these large machines are intended for a small handful of users that run extremely large jobs. That's typically not the case; they typically have tens of users per day up to hundreds per week, depending on institution. With diverse users comes diverse software and the challenges based on that.
The COSMO model which I'm working on is 18 years old. We ported the compute intensive part of the model to C++ using a stencil library and integrated OpenACC pragmas (think OpenMP, but for GPUs) in the rest of the code base.
I'm not a big fan of OpenACC, because it requires the user to make assumptions on the underlying hardware and quite a bit of thinking to get high-performance code. It is quite time consuming to integrate the pragmas so that the code performs. OpenACC capable compilers used to be very unstable, we regularly got compiler breakages and regressions. It got a bit better since we send the code to the vendors, but we still see regressions from time to time. All (usable) OpenACC implementations are proprietary, so we have a vendor dependency. The PGI OpenACC binaries used to be pretty slow, but it is now almost up to par with the Cray binaries.
The successor model to COSMO, ICON, developed at DWD and MPI is also written in Fortran using OpenACC pragmas (and I think OpenMP pragmas because they plan to use it on XeonPhi). The code is interesting in the sense that they are using a icosahedral grid, which stands in contrast to the square grid COSMO uses: Data accesses are not straightforward. Keep in mind that Fortran does not have easy abstractions for data accesses outside of square grids, so you have to use a couple of macros/functions/etc.. to get to the fields you need. Disclaimer: I have not seen the code myself, but the DWD certainly knows how to write high-performance code, so I'm certain most of the code is optimized.
The MeteoSwiss utilizes K80 GPU's in the system hosted at CSCS though right? Or were you saying you don't expect a shift to accelerator-based weather models as a general trend?
I meant it as a general trend. There's interest from the community to run the GPU model on their system but unfortunately the official COSMO code is not fully GPU ready yet.
Well, actually it's actually mostly Jenkins that runs the code for us on Daint. ETH, which I'm also affiliated with, is starting to use the model in GPU mode for climate runs on Daint.
What are you working on at Cray?
I'm a sysadmin so I get called when things so bump in the night and maintain the running of the systems. I also get to help out with other systems in the EMEA region.
Right now, we have quite a elaborate build environment where we first load the GCC build environment to compile the C++ code base. Then we load a Fortran build environment to compile the Fortran code base. I would rather have to use one environment to work in. In the most recent Cray and PGI environments (6.0.3) I have to also add the linker path to GCC version I used for C++. This is to avoid linker issues with the C++11 abi in the newer GCC compilers (somehow, the Cray/PGI environment always wants to link with old GCC libraries). Its not a big problem with our build scripts but annoying nevertheless.
I remember when the TOP500 was being taken over by Beowulf Clusters https://en.wikipedia.org/wiki/Beowulf_cluster Linux made that happen. At the time (and today) it ran on small embedded systems, to some of the world's fastest supercomputers and practically everything in between. Incredible compared to the alternative contemporary OSs.
Now we have powerful computers and lots of data freely available (well maybe less now thanks to Trump), I've always dreamed of running an old computer model say from the 1980s at home.
However, if you want to run the kind of forecasts that a weather center would do, you big issue would be getting input data. Some of the biggest advances in forecasting have come from improving the amount of satellite data being 'assimilated' into weather forecasts - look how much better the southern hemisphere became compared to the northern hemisphere during the early satellite era (1980-2000). Getting this data in a timely fashion is more of a challenge to do at home.
[1] - http://www2.mmm.ucar.edu/wrf/users/
[2] - https://nomads.ncdc.noaa.gov/
[3]Page 3 in - http://www.ecmwf.int/sites/default/files/elibrary/2012/14553...
[1] https://www.ral.ucar.edu/projects/ncar-docker-wrf
Note that these vectors are generated not by just sampling the uncertainty in I.C.'s. Of course, this is because the space of I.C. perturbations is too high-dimensional to cover. In the method above, selection of perturbation direction is based on the adjustments implied when new data is sync'ed to the model.
There are other techniques.
These models are just computer code translations of physical equations mainly. You get an equation to model how air (fluid) moves, code it up, set up a numerical process to integrate over a small time step. Depending on how many assumptions you want to make the equation can be relatively simple or highly complex.
In short, it's fuckin insane. This has some good stuff in it https://www.ral.ucar.edu/projects/armyrange/references/forec...
Just take a moment to consider the feedback loops: huge cloud cover blocks sunlight, so the ground temperature is cooler, so ground air pressure is lower, which probably acts like a sponge to pull in any nearby clouds, which would then increase chances of precipitation. I would imagine the time of day matters too, because cloud cover at night can trap in radiant heat... fucking insane.
I wonder if they put an open bounty on the problem, say for specific geographical regions: find a better prediction scheme for city X, win $Y
It ironic because now we have to rewrite everything to use SIMD/vector accelerators (GPUs).
We need 10x-100x more stations/sensors for the data to be good. I've been thinking of ways to create a cheap solution for a simple weather station that would ping the data once a minute to the central server over something like GSM, but I'm more of a software guy. Determining wind strength/direction is not that hard, but making it all weather proof is.
There's a huge community already feeding data into weather underground. I am not sure if NOAA uses weather underground data, but I suspect they do not because the quality of the data is not very reliable (station mount, model, terrain in the immediate vicinity of the station, etc.)
Weather balloons and satellites fill in most of the rest of the data sources.
All weather models are 3D. Terrain is a major factor in weather. The weather forecast is usually pretty good in the time window of the next 24 hours around very specific places, such as the TAF report for major airports, which covers the 2 miles around the airport, and resolving weather to within a 2 hour time frame. Of course, you likely won't be windsurfing there, and I suspect you want a higher resolution weather, such as to within 1/4 mile, with 30 minutes event bracket, and 7 days out. Wouldn't we all!
2) Yes, the models are 3D, but the data is 2D and is sparse. I'm fully aware of the airport wind data, and it's not even close to being enough, because the number of airports itself is very limited, especially outside of the US.
Its much more expensive than a ground-level mechanical anemometer, but also much less expensive than a tower with an anemometer.
Ocean surface winds are retrieved on vast scales by satellite measurements, due to their effect on the surface roughness, which is measured by radar (e.g., https://manati.star.nesdis.noaa.gov/products.php). But that's an exceptional case.
http://www.atrad.com.au/products/wind-profilers/boundary-lay...
Also, SODAR: http://www.vaisala.com/en/energy/Weather-Measurement/Remote-...
The trick is to make them so cheap they're disposable. Preferably bio degradable, though that'd be tough for the electronics. Weather proof is impossible.
For example, surface temperature forecast verification for Jan 2017: http://www.mdl.nws.noaa.gov/~verification/ndfd/index.php?mo_...
I can predict Hawaii's temperatures out 20 years based on historical data much better than Chicago's. So some adjustment for area variability should be accounted for. People also care a lot more based on how off the temperatures are. So, being off by 5 degrees every day is better than being off by 1 degree for 4 days then 20 for the 5th.
I assume there is some adjustments for this stuff?
I remember flying on one day when the forecast said that I'd be able to climb to 3,000ft in one spot, and to 6,000ft in another spot only about five miles away. In the air, I tried the first spot because it looked better to me. I topped out at 3,000ft. I moved over to the second spot and sure enough, 6,000ft. It was pretty amazing how good that one was.
Due to the banding, you can get extreme variations in snowfall, leading to, "I thought we were supposed to get a foot, but we only got an inch!" type of comments. Or, bands will line up with the storm motion and just pelt an area harder than expected.
Example: http://imgur.com/i0j4I7g
http://blogs.mprnews.org/updraft/2017/02/snowcover-view-from... - under "GFS model shifted". They do similar articles semi regularly.
But in the big picture sense, and for things like hurricanes, forecasting has improved a lot. It's unlikely we'd be caught by surprise today as happened with the Blizzard of '78 when massive numbers of cars got stuck on the highway and had to be abandoned or people were stuck at their offices for a week. (The weather events still occur of course but they're less likely to catch people unprepared.)
It's probably less true than it used to be but, for years, a lot of people in Boston were really paranoid about big snow storms. Everyone knew someone or at least knew someone who knew someone who had be evacuated from Route 128 or slept in their office for a week.
About a decade later a friend of mine moved to Boston and she once told me that she had never seen a northern city where people took off home from work at the first snowflake the way they did in the Boston area.
My girlfriend is from Florida, and even after three years of living here she is still unused to the habit of keeping ice scrapers, umbrellas, rain boots and an extra jacket in the car at all times.
Upstate weather is a mystery. When I was in school we had 80 degree weather towards the end of the spring semester before a small snowfall the day before graduation.
This is a model that didn't quite make it as the replacement for the GFS in the US. It had broad backing from the research community, and if you follow their references back you'll find a wealth of research.
You can run a version of it yourself.
We can't even achieve exascale right now! (But we'll get there)
And, it's more likely that when people realize how incredibly hard it is to build exascale machines, and how unproductive they will be, they'll just end up training deep learning models which approximate the perfect simulations well enough and cost-effectively enough on commodity GPUs that nobody will buy or build supercomputers any more.
There's a reason why the big DL players hire lots of HPC guys for deep learning. And also, parallelizing SGD is a highly non-trivial task, until very recently, you had to linearly lower the learning rate as you added new workers, and keep in mind that bandwidth is far more expensive (both economically, temporally and energy-wise) than compute.
That being said, many people are already exploring deep learning for weather simulation (such as Yandex I believe) and it has worked very well, and is near, or beating SOTA iirc, so I definitely think there's a future there.
For the record, I think exacsale is achievable, but not with the current architecture (admittedly I'm biased, as a deep learning chip and supercomputing chip startup founder), but I think my objective evidence is pretty strong that with a new architecture, exascale is possible. On the other hand, zettascale may be the end of supercomputer scaling. Moore's law, Dennard Scaling, et al are no saviors either, since communication and scheduling logic now dominates the cost of computation.
FWIW, I work at Google, have a background in parallelizing simulations, and built an execution system that runs huge hyperparameter combinations. We called it "exacycle" because it exceeded 1 exaflop / second (no communication between tasks) using only idle cycles.
I think you probably missed my point: it's now pretty well established that for any physical simulation process, you can train a net using far less energy and get an equivalent quality result. The training itself doesn't require lots of communication unless the model is enormous.
You may have a background in parallelizing simulations, but I beg to differ about parallelizing deep learning training. There is a lot of communication involved. For one, your model is very likely to be enormous (perhaps even necessitating model parallelism) and second, there is a LOT of communication involved even with data parallelism.
Most parallelism in DL uses coarse exchange, although I agree wrt to large models. That's similarly true for every supercomputer application (it's the only reason people use supercomputers now; it's much easier and faster if your system fits on a single system).