40% of NYC Taxi Trips Are Uniquely Identified by Census Tracts and Hour
gist.github.com
gist.github.com
I think an argument can still be made that even if the OP is right about the quantity that can be uniquely identified -- keeping the coordinate data still outweighs the real-life privacy risk, that is, the small number of people who want to hire a private investigator/specialist to analyze this data to catch a specific person would find it much faster to track the person the way that PI's normally do so. But the rebuttal can't simply be, "uniquely identifiable trips are probably so rare as to be inconsequential"
Hmmm... but if you already have those pieces of information (start tract, end tract, start hour) what would you want to get from the data? How much someone paid? How much they tipped? Whether they paid with cash or card?
Can anyone see an obvious nefarious use for this data?
Or, you arrive at work late, claiming that you stopped off at a client's before heading in to work, but your employer can now verify that you actually got the taxi from your home address.
I'm assuming in these examples that you just need to know pick-up OR drop-off - if you need both, then I agree with you that it's not much of a concern.
You need both pick-up and drop-off to get a unique row.
Now, I think about your examples, though, perhaps the absence of a particular row could show you didn't do something.
I'm actually working on a small project that has a much, much smaller dataset than the NYC Taxi data but some similar attributes (geographic coordinates mainly). I'd love to produce something like this with what I find (assuming I can find anything interesting).
This is not correct; you need to say "full birthday" which includes the year, otherwise the statement is nonsense.