MTA Open Data Challenge
new.mta.info
new.mta.info
FOIA is the better alternative because it gives you the original, pre-cleaned data. Open data is a lie.
Would love to read more about your experience with Open Data. Any place where I can reach out?
And this one makes some rounds: https://mchap.io/that-time-the-city-of-seattle-accidentally-...
Feel free to reach out!
I do this with my power company’s outage map: https://github.com/patricktrainer/entergy-outages
67k commits!
Sometimes what can happen is that somebody inexperienced will try to make some assessment of the data and come to the exact wrong conclusion because they didn't know what not to trust. But it gets on the news anyway and damage is done.
We can do better than that.
Open data portals generally have data is useful form. FOI probably gives you PDFs.
Having submitted thousands of FOIA requests, I get the impression that you haven't, actually, submitted many FOIA requests. I've received many, many, many, many non-PDF FOIA responses.
Share me some of the open data you've worked with and I'd love to poke at it and tell you where it's wrong and where assumptions about its data is wrong.
Have you had to fight a lot of malicious compliance which balloons up your request count? Or do they typically require an incredibly narrow request that you have up submit N entries per topic?
Even still, I challenge you to challenge yourself to understand where your blind spots are. I've done it many times and have found significant problems with the open datasets I've worked with. If you think my take is weird, it's only because you're not looking or the data you're looking at is inconsequential.
To me, this stuff is literal life and death. If we make mistakes in our analysis because of misinformation from the source, then the lives and deaths of people we're trying to understand becomes tarnished. We can treat our neighbors better than that.
There are lots of reasons someone doesn't want to be "challenged" by some blowhard on the internet. One of them, true in my case, is I don't even work in this area anymore, as I said in my original post.
I really hope you are nicer in person.
Can I ask why exactly you think my take is "very weird"?
Your original post was exceptionally dismissive, without explanation, and your comment on FOI was said so confidently probabilistic that it struck me that you misunderstood what I was suggesting. pardon my aggressive response. I get a lot of similar dismissiveness whenever I interact with government agencies, often where I'm told that something doesn't exist, or "Just look at the data portal", while the data portal is intentionally missing the information I look for. I don't expect you to answer my question, but I hope you can try to understand where I'm coming from in my thoughts and opinions on open data. My intent was only to get you to share your thoughts further.
This HN item for instance, is not about that kind of data. The datasets in question tell you about the transport network, the services, the patronage, the history, all kinds of interesting stuff.
So I find it "weird" that you would respond to a good-faith effort of sharing tons of information about a public transport network with this hostile approach of disparaging open data portals, and advocating instead an approach which is extremely resource-intensive for government bodies, when it's completely uncalled for.
Yeah, if you want to investigate a government cover-up, or shine light on some terrible mismanagement of resources, go for your life and submit FOI requests. Your mention of having filed thousands of FOI requests suggests you have consumed many tens of thousands of hours of public servants' time, and I really hope the results justify it.
Years ago during the pandemic early days, a harvard epidemiology student asked me to proof-read his paper that argued that covid-19 killed more white people than any other race. The dataset he used was the Cook County Medical Examiner dataset. There was a column in there for the race information. If you're curious how it's populated, I can share with you the information.
Previously, I'd FOIA'd the data and received many more columns of information including the names of the individuals who'd died which showed a very clear pattern that the race information on the open data portal was not always accurate for Hispanic-origin names. The details are complicated, and I'm happy to explain my fact checking methods, but the Harvard student's analysis was just flat wrong because it made assumptions that the race data was correct. It was not.
Their response was initially along the lines of, "even if it's 50% it's still going to be true". It ended up being more like 80%, showing that people with Hispanic-origin names were significantly more likely to die of COVID-19.
If you think your audience isn't academics at mega institutions who believe that open data is 100% accurate data, then you've made many incorrect assumptions and I encourage you to reconsider.
"my audience"?
What makes you think I have an audience?
So that means what you want to do is specialize in identifying bias in these datasets and finding the smoking gun. Such a task can be an ugly business but necessary for the public good, pushing data sharers to either share good data, or not share, but not share tricksy data in this unethical way.
[1] https://ny1.com/nyc/all-boroughs/news/2024/09/25/mta-board-a...
I wanted to investigate how well MTA is managing its workforce and compensation (as to require additional tax in form of Congestion Pricing to fix its budget hole), but there seems to be no dataset for that.
Does anyone have links to MTA payroll/hours/overtime related dataset?
or alternatively, I need dataset to study each and every subway improvement project, and components of each project in materials, labor and etc
https://new.mta.info/article/introducing-subway-origin-desti...
Then I think, oh, right, wrong MTA. Guess I've spent too much time dealing with email servers.
Depends what it is. Long as it’s not something you could steal yourself. Ha!
Or perhaps... a subway seat? https://new.mta.info/document/85661
Plus the MTA has a huge budget crunch. I really don’t think they could justify spending money on something with such an unclear outcome.
> 3. Eligibility: The Challenge is open to legal residents of the United States. Entrants must be 18 years of age or older as of their date of entry. The Challenge is subject to federal, state, and local laws and regulations and is void where prohibited by law. Employees and contractors of the MTA, its subsidiaries, affiliates, and directors (collectively the “Employees”), as well as members of an Employee’s immediate family and/or those living in the same household, are ineligible to participate in the Challenge.
Contrary to what you seem to believe...There were more geoblocks when the EU law went into action a couple of years ago. There are less now.
Source for that?
I mean … as I understand the Europeans' law, only if you're doing dumb things to begin with, like giving users' data away to random 3rd parties hellbent on shoving "ads" down one's throat. If you had just made this site a simple HTML page that just had the information the MTA wanted to convey on it, AIUI the EU doesn't have a problem.
Which … the MTA does appear to be, sending requests to Google, LinkedIn, and some other CDNs.
I also don't think the MTA has any EU presence, so what are they going to do?
There is a massive difference between complying with the law and proving you comply. (Think: IRS audit.)
> don't think the MTA has any EU presence, so what are they going to do?
Send letters. The MTA would be obligated to respond to them, which means legal bills.
> The MTA would be obligated to respond to them, which means legal bills.
…why would the MTA be obligated to respond to them? They've no jurisdiction/sovereignty over an American transit agency.
Why would they audit themselves against laws that don't apply to them? (Again, jurisdiction?) I've never worked for a company that audited itself against every law from every nation on Earth; we complied with the laws where we had a presence and did business.
We're talking about a US transit agency. Even thinking about whether the EU has a problem with the agency's website is sort of absurd to begin with.
Not how jurisdiction works.
As a part-time New York City taxpayer, I'd rather we not be paying EU lawyers to make sure the MTA's open data complies with European law.
You can enforce what people and companies do within your borders. You cannot enforce what companies or people outside of your borders do.
Isn’t the GDPR’s basic theory about jurisdiction that, if I’m sitting in New York City but routinely serving my web content to people in France, that service I’m providing relies on browsing intentions and tracking functions being executed by a user and on a machine in France, and therefore the meat of the “wrongdoing” is happening within their borders?
You can choose to do that the European way or not at all. And the local contests division of the NYC local transit authority is choosing “not at all.”
Isn’t this then a case of NYC complying with the EU’s express wishes for privacy by not “exporting” code they don’t want there?
I have no problem with voluntarily complying to GDPR-style privacy regulations because it's the right thing to do. Where I am able to make the decision, we store basically no user data beyond what's required to do whatever the user is trying to do.
My problem is the EU pretending that US companies must be fully GDPR-compliant because someone in France chooses to go to their website. At the end of the day, laws are only laws because you can enforce them. If I had a magic wand and could rob a bank but the police for some reason were unable to arrest me, the fact that bank robbery is illegal is merely semantic at that point. If I chose to flaunt GDPR non-compliance on a US-based website the EU would be impotent to do anything other than block the site, which wouldn't make me any more likely to suddenly become GDPR compliant.
It's a fiction and I probably wouldn't care about it nearly as much except it has essentially ruined the public internet with cookie banners everywhere.
Every time a cookie banner gets displayed on some non-EU resident's personal blog, a puppy dies.
I can access it just fine from Sweden :shrug: