This doesn't solve the problem for research where the datasets are in the many terabytes, but then again there are many papers where the datasets are well under a gigabyte.
This doesn't solve the problem for research where the datasets are in the many terabytes, but then again there are many papers where the datasets are well under a gigabyte.
No one got tenure for a well curated data set.
We've had a ton of scientists sacrificing their own individual reputation and relationship with publishers (and of their groups) in name of something everyone agrees but few really stand up for, because the personal gains are almost exclusively negative. That's textbook use case of regulation.
I don't know specifically what should be done but perhaps a requiring publishers certain obligations (e.g. responsibility of maintaining papers for a long date and turning them public afterwards); or maybe simply a universal obligation to open publications after say 5 years.
People fear this will compromise quality or sustainability of publishers. But the community need publishers. It's a tag of credibility. So if publishers are into trouble (and they're really needed) they'll find a way by e.g. demanding payments from publications from the most wealthy labs.
The problem is that whether or not a paper is Open Access, "Data is available on request from the author" may be an undocumented bit of spaghetti code, may be stored on a Zip disk around here somewhere I'm sure, or may just be lost.
Regulating "You must make your data accessible, and maintain it well" is much harder to implement, and much harder to check. Some grants now have sections describing what will happen to the data etc., but right now there really is very little reason beyond their own personal desire for researchers to maintain good quality software and data repositories.
Additionally, the discussion over shared data is usually a cry to improve reproducibility. In this publication-centric world, one can't publish a paper that is just "we reproduced an existing paper", so outsiders look to "poach" new applications or findings from the data.
The data from my dissertation is on a Zip disk. I didn't mean for it to get lost to the world, but, you know...
I'm on a project now where provision of the raw data to the granting agency for public availability is a requirement. But the metadata, database structuring, and answering questions from people who take my data and have a question about it has added noticeably to my workload.
I know this is an imposition under current incentive structures but it seems to me on a macro level that this is exactly what we the public would want to see occurring, people actively reading other people's research and data and unrelated scientists asking questions of each other and reviewing each others findings.