Uber NYC Trip Data from April to Sept. 2014
github.com
github.com
I think it's fair to say we (Uber) didn't intend for this data to be made public, but it was produced under a Freedom of Information Law (FOIL) Request, and now it is out.
So I'm curious how we at Uber could help improve cities and society with anonymized dumps of this kind of data - if this was a formalized thing.
Let me know your thoughts here or directly (benm at the domain of uber)
The other thing I think you could do that would be very useful for drivers is show them areas that there are typical long waits / longer rides right now. I know a lot of drivers do some planning, but they seem to be in the dark. I'm sure a 'where to uber' app on android / ios for drivers telling them when to get where to take advantage of surge pricing for customers that will want to get taken to a desirable spot would be VERY popular among drivers.
Some wicked-smart former Code-For-America Fellows working on this YC-backed project.
What they do not know is where new bus services might be useful. If Uber was able to show that there is a demand for transport between X and Y not currently served by busses (e.g.: late night riders are there but the busses aren't, or there isn't even a bus that services route between A and B).
Services like taxis and Uber would be invaluable for planning mass transit systems.
Hypothetically there may be more facets of data not provided to cities, etc.
Second item is DUI rates before and after Uber comes to town. For that matter how much surge pricing happens during major events such as concerts, football games, etc where people may perceive they must drive. Combine that with congestion data and perhaps uber can propose a partnership with the local metro to pick people up outside of the deepest traffic area by doing a bus/rail hop first.
For instance Houston has light rail from our sporting venues to park and ride lots. If you can group people going to the same neighborhood it's possible to shift congestion and make the last mile cheaper for riders and more environmentally friendly to boot.
You also have major events where really no one should be driving but they are so data is gathered: hurricanes, ice storms, blizzards and so on. In Houston we have regular flooding and it would be interesting to see where reroutes happen due to deep water and it would be very useful to have drivers marking flooded spots so it can be shared back to other drivers via Waze/Google Maps and maybe even CoH fire and police.
I could easily see an autonomous car on the side of the road with an led bar in its back window or on its roof warning that water is 12" past this point. Many of the worst spots have machine vision readable flood gauges but even for those that don't it may be possible to gather surface elevation from the Google maps car since they have such accurate GPS, and that data doesn't change often.
In fact one could see Uber / Google subleasing a few of those cars to HPD just for the purpose of sending them out to gather data (including video) on flooded intersections so police could spend their attention elsewhere. And of course so they'll know if someone does flood themselves out anyway by driving past the car and then stalling. Autonomous cars already track object distance and speed of vehicles in their view.
That's enough for interesting visualizations and statistical analysis of comparisons between the two, although the original 538 article (http://fivethirtyeight.com/features/uber-is-serving-new-york... ) is pretty good. (for my own visualizations of the NYC Taxi dataset, see my blog post: http://minimaxir.com/2015/08/nyc-map/ )
The Aggregate_FHV_Data.xlsx contains data on Lyft as well. In September 2014, Lyft did 115,999 total pickups in NYC, while Uber did 1,028,136 pickups. (however, Lyft didn't have any activity in NYC until the end of July.)
I was just going to mention that if people were interested in exploring this type of data, there exists a Kaggle competition[1] for "Taxi Trajectory Prediction". The data contains the full GPS paths of the taxis and comes with pre-existing scripts you can run online, including a Python script for visualizing the city via taxi paths[2].
There's even a secondary challenge for predicting travel times which include visualizations of taxi travel speed, producing "veins" on the city streets proportional to the speed you can travel along them[3].
[1]: https://www.kaggle.com/c/pkdd-15-predict-taxi-service-trajec...
[2]: https://www.kaggle.com/mcwitt/pkdd-15-predict-taxi-service-t...
[3]: https://www.kaggle.com/wikunia/pkdd-15-taxi-trip-time-predic...
I'm a normal/infrastructure engineer who reluctantly ends up focusing on test systems every once in a while. I get all the tests running, from a single script, with all results collated and machine readable, get CI set up properly, make the system reliable enough to trust, etc. Because someone has to do it! And yes I've done this in organizations where there was a QA engineer or two, who did reasonably cheap semi-automated QA, but couldn't put together a whole system and make it reliable enough to run on its own. But there are definitely a few places where QA engineer is not the "lowest rank".
Trips outside of NYC (aka the 'burbs) have the full street address listed. Those are addresses of people's single family homes, vs. the NYC addresses which tend to be multi-family dwellings.
Uber had a pretty sharp uptick in early September (end of summer?) which Lyft didn't appear to have.
https://www.amigocloud.com/api/v1/users/22/projects/3153/dat...
By the way, here is the data in many different formats (kml, shapefile, geojson, etc) in case you want to play with it in other software: https://www.amigocloud.com/data_share/8850bab3c62141278ac5c4...