Supercomputer quietly puts U.S. weather resources back on top
usatoday.com
usatoday.com
Weather forecasting accuracy is still constrained by both the available observations and the ability to run the forecast models. This new supercomputer won't be nearly enough to process all the new data that will arrive in the next few years, so the pendulum will swing back to the bottleneck being computing power soon enough. I think the most exciting things will happen over the next 5-10 years, as we end up with another 2-3 order of magnitude increase in the available observations to the models and we'll build new supercomputers to process them. I bet we can get incredible improvements, things like reliable 2-week weather forecasts.
> The reason we encode velocity data as an image is so we can pass it off to the GPU on the iPhone and iPad. Both the storm prediction and the smooth animations are calculated on the device itself, rather than the server, and all the magic happens directly on the GPU.
Sadly, until we have a breakthrough in battery technology, all of those will be parallel processors programmed to be asleep as often as they can.
I don't know anything about it, but I can think of numerous ways to mitigate the issues you mention. E.g.:
* Does it actually matter that the absolute value of temperature sensors is inaccurate, or is it enough that they have good relative accuracy (do they?)
* Do you really need to tell if someone is indoors or outdoors given a large enough sample of users in a given area? That also seems like something you could detect with heuristics (work hours, jumps up/down in temperature).
There's a lot of ways to tease useful signals from noisy data. Is it really the case that smart phone temperature sensors in the aggregate are useless as an input into weather forecasting models?
I can't answer much to your second question but I imagine it's easier in colder environments, especially if the phone is not in a pocket. In addition, here's a publication that claims to be able to get useful signals from noisy data, so I don't know if I would say the measurements are useless [2].
[1] Fig. 10 in https://www.sensirion.com/fileadmin/user_upload/customers/se...
[2] http://onlinelibrary.wiley.com/doi/10.1002/grl.50786/full
Fixed weather stations have corrective mappings associated with them. For example the station at the airport might be assumed to be 1 degree warmer than the generalised temperature for your city. It might be known to collect 5% more rain when the wind is northerly, and 8% less for southerly. These mappings are very important, because they let you shut-down a weather station, and build a new one in a different location, and compare the data from the two.
if the weather stations aren't in consistent locations, with discoverable local behaviors, then converting the raw data into generalised data about the area is going to be very difficult.
What I do know is that whilst teasing signals out of noisy data is achievable when the data are unbiased, the possibility of biases makes it significantly harder, especially when you don't know what the biases might be. This can be a significant problem with all sorts of data that you might think were fairly good, including both in-situ and satellite measurements.
i'd been browsing through some WMO reports and other associated NWP lit recently and i'd seen stuff on GPS-RO, but nothing on anyone assimilating "smart" devices.
There are researchers who use this data - we have collected about 4 billion atmospheric pressure measurements that we have distributed for academic and government research. The primary researchers are Cliff Mass and his lab at the University of Washington. There are also groups in Canada and the US that are using the data. IBM is now also collecting and using smartphone pressure data through their mobile apps. [2]
Generally speaking, the current trend is to take the live data stream, run it through a quality-control algorithm and then use kalman filters in the WRF data assimilation package.
There are some papers published, but it is still early. I will find some links to papers if you'd like to read them. [3]
[2] http://www.nytimes.com/2015/10/29/technology/ibm-to-acquire-...
[3] Utility of Dense Pressure Observations for Improving Mesoscale Analyses and Forecasts: http://www.atmos.washington.edu/~hakim/papers/madaus_hakim_m...
"It also learns every time you actively report to the community on sky conditions and hazards, translating this information into weather predictions."
does sunshine actively run a numerical weather prediction model like WRF?
it'd be great to assimilate all these extra sensors into the NWP centre's models, but i imagine things like cal/val and WMO agreements to share data might make things difficult for commercial companies?
And yes it would be great to have all these sensor readings available to NOAA, Environment Canada, ECMWF, and everywhere. I have made lots of progress in getting them to talk about it, but it's a long road before any government starts using this data in its own NWP models.
https://www.congress.gov/bill/114th-congress/house-bill/1561...
there's some stuff being written this year as well for the side that i work (space). with all the upcoming sensor gaps, they're looking at alternative ways to cover.
There is no reason to look at this in any way adversarially. I guess it touches a nerve, because of the similar silly comments elsewhere, worrying about how the USA is "losing" if, say, the Chinese economy's size exceeds the size of our own. Other folks doing well doesn't make us worse. When it happens we should be happy: other people are doing well, and we've got an opportunity to learn and improve too!
Does your software become better if you run it on a faster computer? It will certainly help, but he instilled on me the impression that the US had a bit more catching up to do than just buying a new supercomputer.
Another important topic is coupling ocean and atmosphere models, and how you handle the data assimilation across your coupling - again, this is an active area of research, with lots of subtlety.
Of course, the model dynamics and physics are also very important, as is the resolution. This last issue is one of the places where bigger computers have really direct benefits, the other being increasing the size of an ensemble.
It's also important to realise that the various systems tend to be better suited to certain things, so whilst the ECMWF system is better in key global metrics, it doesn't (on average) provide better forecasts of all quantities in all situations.
[0] http://onlinelibrary.wiley.com/doi/10.1034/j.1600-0870.2001....
http://www.uvm.edu/~cdanfort/research/danforth-bates-thesis....
Here is the UK met office. "Most of the time the atmosphere behaves rather like the lower-left picture where we can predict with confidence for a few days and have to use probabilities thereafter. " http://research.metoffice.gov.uk/research/nwp/ensemble/conce...
Meanwhile, Orlando itself is relatively safe when it comes to hurricanes. Few make it far enough inland to cause significant damage there. (Though, obviously, some do. I was there during the 2004 season.)
Here in central Texas, we have the same thing, but for the opposite reasons. Our water table is pretty deep - typical wells in my town are 900ft deep. But rather than sand, we've got limestone. Typical land has inches of topsoil sitting on top of solid limestone, so excavating a basement is prohibitively expensive.
So you're telling me one could blast out a cave in the backyard? Spare no expense!
http://www.newyorker.com/magazine/2015/12/21/the-siege-of-mi...
If you're worried about it flooding, you should probably spend less time building stilts and more time dealing with the fact that Delaware no longer exists, nor does most of New Jersey, New York City, large swathes of the Eastern seaboard, 75% of Louisiana, the Texas coast all the way in to Houston, much of Los Angeles, and all of Sacramento.
On the other hand, it will be nice to splash in the ocean at Joshua Tree National Park, and I'm sure Arkansas residents will be happy to finally be able to own beachfront property.
http://www.jessstryker.com/national-parks/everglades/rock-re...
Thanks!
The main differentiation vs a standard datacenter is low-latency high-throughput interconnects, often with a specific or reconfigurable topology. Something like InfiniBand, where you can ensure very consistent, very favorable bounds on your latency and throughput. So think of it as being a "custom installation" or "custom cluster" instead.
Nowadays there really are only a few real choices left in high-performance computing hardware: CPU vs GPU, processor architecture (eg POWER vs x86), and interconnection. Everything else is more or less irrelevant.
I have to admit though, I've always wanted to see what you'd get if you built a Cray-2 with GPU chips (eg Pascal) or Xeon Phi as processing modules. The sheer density of chips Cray managed to cram in that case is impressive.
Glenn Lockwood has a nice blogpost with much more detail: http://glennklockwood.blogspot.com/2013/05/fdr-infiniband-vs...
Depending on topology, you can get ~1 µs latency with Aries. It's based on PCIe-3, so it's different from InfiniBand.
"Chaos monkey" works fine when you're dealing with random cloudstuff and you are in a position to duck-punch your production code. When you have a large dataset that you need to process reliably and predictably, the mantra shifts to "fail never", and that requires engineering up front in combination with carefully-planned maintenance.
Full disclosure: I've worked on this specific supercomputer. While I'm not on the admin team in question, I do work with them occasionally and I have a lot of respect for them -- their job is not an easy one.
Bottom line: If you are using your system to get work done, the trend is to spend the money for reliability and availability. If you want to save money, it comes at a price (but is certainly do-able)
Related: Look at Apple, Google, Facebook, Amazon, Microsoft... they are large enough to build their custom systems themselves. Not everyone operates at this scale. Individuals can buy a cheap PC and build a NAS... or they can buy one from Netgear, Buffalo, QNAP that "jsut works" at a price premium.
The point I wanted to add is that ECMWF, generally acknowledged to be the leader in NWP, also uses a pair of large Cray supercomputers (http://www.ecmwf.int/en/computing/our-facilities/supercomput...).
Note that the Cray XC30 uses Xeon processors and a custom interconnect. The old ECMWF system used IBM Power chips.
[0] http://www.metoffice.gov.uk/news/releases/archive/2014/new-h...
high fidelity weather modelling algorithms are highly constrained by node-to-node communication latency. So their supercomputer probably has some networked-direct-memory-access or another non-standard setup tuned toward passing data between adjacent nodes quickly.
[1] http://cliffmass.blogspot.com/2016/02/the-national-weather-s...
[2] https://www.google.com/webhp?#q=supercomputer+site:+cliffmas...
http://mathsci.ucd.ie/~plynch/eniac/CFvN-1950.pdf
It's something you can solve in a few lines of python and run with minimal compute resources. A nice thing to try is to feed in some data from somewhere like below and watch what happens as you integrate it forwards in time.
http://www.esrl.noaa.gov/psd/data/gridded/data.ncep.reanalys...
There are some available weather modeling packages, such as WRF[0], as a mesoscale (few km to 1000 kms) model, that one can download and run. You'd obviously need to find a resource for the appropriate input data. I don't know how easy it is to get up and running, though.
http://www.elsevierscitech.com/emails/physics/climate/the_or...
> In fact, the spurious tendencies are due to an imbalance between the pressure and wind fields resulting in large amplitude high frequency gravity wave oscillations.
Since this paper was published within the last 20 years, I can't imagine what they were referring to by 'gravity waves'
Do you know what is meant by this statement?
Some examples of gravity waves in the atmosphere: [1] [2]
[0] http://glossary.ametsoc.org/wiki/Gravity_wave [1] http://cimss.ssec.wisc.edu/goes/blog/archives/2051 [2] https://www.youtube.com/watch?v=yXnkzeCU3bE
Basically, gravity waves (or g-waves) are a type of perturbation in stratified media where the restoring force on sound waves is buoyancy. The other scenario is where the restoring force is pressure, which are pressure waves (or p-waves).
Gravity waves are important in not just planetary atmospheres but stellar media as well, particularly in the outer region of stars. A popular candidate for /gravitational/ wave sources is binary neutron star systems, so the two words aren't interchangeable there since they refer to two very different phenomena!
Perhaps he should have been having a chat with the NSA boys...
I know it's big.. no idea how big.. it's not tangible to me.. never visited.. I'm sure most americans have not visited.
"The Library of Congress is the largest library in the world, with more than 162 million items on approximately 838 miles of bookshelves. The collections include more than 38 million books and other print materials, 3.6 million recordings, 14 million photographs, 5.5 million maps, 7.1 million pieces of sheet music and 70 million manuscripts."[1]