Mapping Traffic Fatalities
lucaspuente.github.io
lucaspuente.github.io
As a next step, I suggest calculating the actual rate of fatalities in a particular location vs. some naive estimate for that area, based on traffic counts, car utilization, or nearby population. This will highlight places that have a higher-than-expected number of fatalities.
Some of these places can already be seen visually. For example, Florida is blotted in red even though the population is concentrated around the coasts and bays and the Orlando area; therefore US 301 and other 'interior' portions are more dangerous than population alone would explain.
The Carolinas are similarly covered, well-outside of the Crescent from Raleigh-Durham through the Triad to the Charlotte suburbs. California's Central Valley looks particularly dangerous too, despite most of the traffic sticking to the far western edge with I-5, or the eastern valley floor where the built-up areas are.
I think if this is true at all (I'm not sure it is) it's just an artifact of the large blobs used on the traffic fatalities map.
A population map, in other words. As a very basic demonstration of the capabilities of R, this is OK, but as a visualization, this is 100% useless. You would need something like "percentage of expected fatalities for a given stretch of road, given the type of road" to actually have something interesting. And that takes a lot more work.
[1] http://www.statemaster.com/graph/ene_gas_con_percap-energy-g...
However, you won't get much insight from plotting this data geographically: a table (as in that link) is enough. Basically you get exactly what is expected: Daniel Patrick Moynihan's "Law of the Canadian Border" - southern states have more social problems.
If you want high-resolution geographical data for purposes such as finding problem roads, that would be a lot more difficult - you need to have a graph of the road network (from e.g. Open Street Map), with traffic data attached to the arcs of the network (I don't know of any public access databases like that, although such data is commercially available for road planning) plus accident locations accurately mapped to the roads. I suspect it wouldn't be a quick project to do.
The crash analysis includes an analysis of every accident within a five year period, identification of high crash locations and what caused the crash (high speed, inebriation, or a list of other options), and then what strategies can be adopted for the most impact in a benefit/cost analysis.
Just looking at heat maps of crashes without inclusion of traffic volumes (one crash per million miles is nothing compared to an area of one crash per thousand miles traveled) becomes meaningless.
But then for some reason, they throw out all rationality when it comes to naming files and fields. It's as if there was employee who was old-school and was stuck in mindset of naming things under the 8-character all-caps limitation. And then someone else thought, "Fuck this", and went the entirely opposite direction by naming files with long titles and spaces, and the two employees had a passive-aggressive argument that never got resolved over the decades.
This makes scraping and collating the data a real pain in the ass. But compiling the data got a lot easier when I realized that wget works on FTP sites. Here's a mirror of the FTP site (around 1.7GB zipped) that I put up on Github (still working on getting everything together in one database): https://github.com/wgetsnaps/ftp.nhtsa.dot.gov--fars
https://www.icpsr.umich.edu/icpsrweb/NACJD/NIBRS/
Counting homicides, and even determining what exactly is a homicide, is not as clear as one might think:
http://contentsmagazine.com/articles/homicide-watch-an-inter...
Not that many agencies post the data themselves. Surprisingly, the NYPD does post it as part of their felonies dataset: https://data.cityofnewyork.us/Public-Safety/NYPD-7-Major-Fel...