Tracking Heat Records in 400 U.S. Cities
pudding.cool
pudding.cool
Block maxima (e.g daily max temperature) follow a generalized extreme value (GEV) distribution.
One can try to see whether this distribution is stationary or not.
Things can get much fancier by e.g. modeling correlations across cities but the above is pretty basic.
A great practical intro is: https://link.springer.com/book/10.1007/978-1-4471-3675-0
0 is really cold. bellow freezing. Negative temps are pretty rare. freezing is 32 which you just have to memorize.
70 is room temperature, and 68-75 is comfortable for most
100 is really hot, temperatures above this are also fairly uncommon for most folks.
I'm really curious why the site needs such invasive analytics. It's not by accident - it costs a minimum $1200/year.
And $100/mo might not seem so much if it can make debunking interest-spoofing attempts more trivial, especially if there is ever any kind of bonus to the visualization authors for having a big hit.
They have long time series: https://climatereanalyzer.org/reanalysis/monthly_tseries/
They have wonderful daily difference-from-average "anomoly" maps: https://climatereanalyzer.org/wx/DailySummary/#t2anom
I really dislike this site you've linked. The whole middle of the chart is completely overwhelmed with data points unless you zoom way in, until you can only see a decade or so. There's the extreme data points of each year that stick out, but what's actually happening most of the time is utterly occluded & unclear. The data masks rather than informs.
https://climate-explorer.oikolab.com - I created this to show how temperature has changed for any arbitrary location of interest. Working on adding CMIP6 projection too.
http://myweb.wwu.edu/dbunny/pdfs/Evid_Based_Climate_Sci/Ev_B...
- Reanalysis data is generated from all available data source - land station, satellite, weather balloon, airplanes, etc. They are corrected from observational biases, using the similar approach as Kalman filter, combining data with known laws of physics.
- I do have confidence and use these data for analyzing weather-driven utility load data with excellent result.
- The data assimilation process and methodologies are all well documented - https://rmets.onlinelibrary.wiley.com/doi/epdf/10.1002/qj.38...
I'm sure we'll return to the regularly scheduled inferno soon.
What was the hottest year? What was the year with the most broken records? What is the overall trend? It actually told me what the hottest day on record was, but there is so much to read I completely glazed over and missed it the first time.
Limiting data points on a zoomable map is a great reasonable default, but it degrades pretty terribly when it matters.
It also ignores heat islands as cities get larger.
Isn't it the case that as geographies get smaller, the count of them gets higher, and you'll always get on square inch of the US that's breaking the time-of-year temperature record right this second!
I'm all for data, but the statistics here seem selected to drive towards the conclusion "global warming is a huge problem". Climate change is so real that we don't need to resort to this sort of trickery -- it harms rather than hurts the message.
For Los Angeles (my city) the data goes back 145 days a year. If temperature has not been increasing (and therefore each individual day follows an identical distribution) that would still mean that 1/145 days we'd see a record high in Los Angeles - so about two highs a year. Highs in Los Angeles are obviously going to be correlated to highs in San Diego - but not exactly because of cloud cover and stuff so its reasonable to think that even if we were in a boring unchanging climate we'd still see hundreds if not thousands of records set every day across the country.
With a normal distribution you wouldn't expect even one record to be set every year. The expected case would be zero records in a year.
A record in that data is defined as a particular high in a particular city on a particular day. The data goes back 145 years in Los Angeles.
So if it is the hottest May 24th ever that would be a record.
If everyday followed an identical and uncorrelated distribution then we would expect that 1/145 days would be record.
In your IQ example imagine if you had a school with 365 separate classrooms with 144 students in each one. A new 145th student then enters each classroom. The chance that the new student is the one with the highest IQ is 1/145. So in the universe of the 365 classrooms you'd expect 2 new IQ records to be set.
Let me ask this, why in your school with 365 classrooms would you expect 2 new IQ records to be set with 145 new students. With a normal distribution, you'd expect all of those new students to be within one standard deviation of 100.
> The chance that the new student is the one with the highest IQ is 1/145
Why do you think this is the case?
Here is some python code you can run. You can change around the distribution however you want and you'd get the same exact results. The chance that a particular record is the highest in a set of IID variables is not dependent on the distribution itself.
from numpy import random
RecordsSet=0
TotalSimulations=365000
for i in range(0,TotalSimulations):
ClassRoomSample = random.normal(100, 15, 144)
RecordIQ = max(ClassRoomSample)
NewStudentIQ = random.normal(100, 15, 1)
if NewStudentIQ>RecordIQ:
RecordsSet=RecordsSet+1
print(100*float(1)/145)
print(100*float(RecordsSet)/TotalSimulations)
A record is set 1 out of every 145 days/classes. Which means in a year/school of 365 days/classes you'd expect a bit more than 2 records on average.However, we should be somewhat regularly setting new highs just because we have on the order of just 200 years of data (so just 200 data points for each daily high); if the planet were not warming, and the high for any given day will be normally distributed around some average for that day, then as time marches on the odds of getting an extreme outlier increases. Some cities on some days will have a high that is very improbable, while others will not have gotten "lucky" yet.
Graphs such as https://www.aos.wisc.edu/~sco/clim-history/stations/msn/MSN-... show the average high, mean, and low temperature.
This is even more pronounced if you look at the summer (June, July, August) temperatures - https://www.aos.wisc.edu/~sco/clim-history/stations/msn/MSN-...
You can also look and, while it isn't the daily records, you can see the daily temperatures - https://www.aos.wisc.edu/~sco/clim-history/stations/msn/msn-... for this year so far.
I’d be interested to see extremes on both ends.
The all time record high in Seattle (per the site’s data) was 107, last June. Other data I looked up shows 108. There are of course discrepancies between different data sources, my localized weather showed 112 that day. “Officially”, several days I’ve experienced >100 temps in my 20 years here are in the records as lower.
We typically have little to no precipitation during the summer. And of course every weather statistic varies by year, persistently higher summer temps have arrived (albeit lagging behind many other places) along with a much longer and more severe wildfire season over the last few years.
I’m crabby about the rain right now, but I fully expect to be relieved by it in a few short months.
But in fact this April and May are unusually cold and wet, very annoyingly so.
What went into animating between these two views of the data?
Anyone else reminded of the bad-old-days when Flash websites were common and broke all UI norms?
Their website, their decision on how to “force” you to experience it. Remember, you can always close the tab if it’s not compelling or interesting to you personally.
Imagine an unholy text book where once you turn the page, the pages before become inaccessible. Its your own fault for not understanding everything on the previous page and you have to read the entire thing before you can start again to find that page.
Yes its the users fault.
Don't hide compelling or interesting information behind bullshit ui restraints. Particularly egregious was the "Click here for more information, click two pixels lower to permanently leave this screen."
I'm tired of the same discussion again and again on HN. Javascript bad. Anything but #000 text on #fff bad. CSS and animations?! Fuggedaboutit. If it's not free range, handcrafted HTML, I don't want it anywhere near my 5950X.
After all, your goal is to make sure the viewer understands the information being shown, not impressing them with moving text.
You should open source the data.
The data isn't theirs to open source (or not).
> Temperature records are collected from ACIS, which tracks weather for approximately 400 US cities. ACIS has data from the ThreadEx project.
The ThreadEx project is http://threadex.rcc-acis.org
This suggests its using the data from a service that is collected from http://www.ncei.noaa.gov which right on the front has: Looking for Data?
Maybe people who have more expertise in the field or the data collection can chime in.
Why? A thermometer works exactly the same it did 100 years ago. For such a simple measurement like temperature I can't think of many reasons why measurements from the 30s could be significantly different (as in more than fractions).
I also guess that if the data was considered unreliable for any reason it would not be used.
- Most cities grow and develop over time and the parts of cities with more buildings and especially with more cement will radiate more heat. If an observation station was built somewhere less developed then city growth might artificially push up temperatures around the station. This effect might be strongest somewhere like Phoenix or LA.
- This one seems unlikely (well, I'm sure it happens but it's very likely to have already been corrected for) but aggregation methods might have changed over time. I could imagine in the 30s someone might check the thermometer every hour and record the current temperature in a logbook. Modern thermometers are almost certainly automated. If the "highest temperature" is measured by recording the temperature each minute and taking the max then you're going to get a higher number than the manual method.
- site selection might have changed: observation posts probably used to be buildings which needed to be large enough to fit equipment and the people to operate them. City real-estate is expensive and meteorology used to be a cost-center so these buildings were probably placed at the edge of town. I'm sure modern stations are much smaller, they don't need a full-time staff to operate them, so they can be placed closer to the (hotter) centers of cities. Alternatively: I'm sure many weather stations are now situated at airports, an option which was not available in the 30s, and airports are particularly pavement-heavy (hotter) parts of cities.
The temperature record - at least, what's used for generating long-term trends - is homogenized to remove discontinuities (step-wise or linear) associated with this sort of drift or other significant changes. Regardless, this effect likely wouldn't matter for the stats the website illustrates.
> This one seems unlikely (well, I'm sure it happens but it's very likely to have already been corrected for) but aggregation standards might have changed over time. I could imagine in the 30s someone might check the thermometer every hour and record the current temperature in a logbook. Modern thermometers are almost certainly automated. If the "highest temperature" is measured by recording the temperature each minute and taking the max then you're going to get a higher number than the manual method.
When manual reading of the temperatures was the norm, you'd have a mercury and an alcohol thermometer to record the daily max and min, respectively. For a mercury thermometer, it would have a stopper that rises and records the maximum level the liquid reaches. A human needs to invert the thermometer to reset it. Same for an alcohol thermometer, albeit in reverse (records the minimum). So granularity of observation doesn't matter - the daily max, historically, is truly the maximum temperature taken since the last time a human reset the stopper.
> site selection might have changed: observation posts probably used to be buildings which needed to be large enough to fit equipment and the people to operate them. City real-estate is expensive and meteorology used to be a cost-center so these buildings were probably placed at the edge of town. I'm sure modern stations are much smaller, they don't need a full-time staff to operate them, so they can be placed closer to the (hotter) centers of cities. Alternatively: I'm sure many weather stations are now situated at airports, an option which was not available in the 30s, and airports are particularly pavement-heavy (hotter) parts of cities.
Also accounted for by homogenization. In most cases when siting changes, you consider that site a new station.
For any wanting more information on one of the maximum and minimum thermometers: https://en.wikipedia.org/wiki/Six%27s_thermometer which was invented back in 1780.
This is often noted in the climate record data. If you look at https://www.aos.wisc.edu/~sco/clim-history/stations/msn/MSN-... you will see vertical dotted lines indicating location change for that measure.
And yes, it does impact it - https://www.aos.wisc.edu/~sco/clim-history/stations/msn/MSN-...
While you're here, do you happen know of a good book / other resource to read if I want to learn more about how we decided climate change was real and anthropogenic? I take it for granted; I've never actually looked at any data.
Yes, more or less, they're the same. It's simple enough to restrict measurements to just surface observations made with a thermometer of some sort if you want a wholly apples-to-apples, long-term temperature record. You don't have to add in satellite data (satellites don't directly retrieve surface air temperature so the direct comparison is fraught) although there do exist techniques which leverage it to augment the lack of spatial coverage of surface observations, especially over oceans.
Quality issues have been studied ad nauseum over the past thirty years. Nearly half a dozen independent efforts have all tackled this problem. Probably the most notable recent one is the Berkeley Earth Land/Ocean Temperature Record (Rohde and Hausfather, 2020 -https://essd.copernicus.org/articles/12/3469/2020/essd-12-34...).