The incentives are definitely bad, and that's where the actual fix should be.
The code I write is not my code, it's the banks.
However I think it would be better to directly incentivize data release rather than require it, at least in biomedical sciences. Because of patient privacy issues there is no way raw data release can be required across the board. And I certainly do not trust the NIH to come up with a coherent set of rules for when it is versus isn't allowed, which would mean loopholes and more corruption.
This has not been born out in my experience, or the experience of others. Data products are chronically undervalued and undercited in science, and do not come with guarantees of authorship unless you put up barriers to access them without it.
Open data is, at this point, a decision made for either ideological reasons or as a condition of funding and/or publication, not a decision that is itself usually "worth" the investment of curating a public data set.
But yes, overall more openness is good. Still, the cost losing trust in society is very high (as you need to verify everything).
I've already heard of someone planning a product ( initially targeted at lazy^H^H^H^Hbusy high schoolers and undergrads) that will use AI to reverse discover citations that fit a predetermined narrative in a research paper. Write whatever the hell you want, and the AI will do its best to backsolve a pile of citations that support your unsourced claims and arguments. The founder, and I use that term very generously, expects the academic community to net support this because it will boost citation counts for the vast majority of low citation, low visibility works.
Did you say you'd keep it forever? Did you say 5 years? Who's in charge of making sure that this centralized repository isn't inappropriately holding and distributing data?
Funding organizations also have different requirements.
Then a script is only useful if paired with a set of libraries of a particular version, a specific compiler/interpreter, an OS, also there may be specialized hardware involved. Some of the languages used in science like SAS, Stata, SPSS, and Matlab, etc aren't free and open source so you can't always just bundle it.
And even data storage isn't trivial. For a recent small conference abstract I processed ~150GB of data. Hundreds (thousands? Tens of thousands?) of other papers have looked at that same data. You would really want some way to deduplicate that storage, but that introduces some additional complexity.
I do like this vision but I think it would be a major undertaking that would require a lot of well funded institutions coming together rather than any one in particular doing it on their own.
"Your data is open and available in perpetuity for whatever use" is in deep conflict with how we think about human subjects data, often for very good reason.
It's not just the deleted_at column being set, there's no backup, it's really gone. Every copy, forever.
I appreciate the ethics of it, and part of my reason for working in this area is because of these ethics, but even 5 years in there is so much reluctance to press that delete button.
When you are doing science, you often do not just have a single standardized data format. It's not like taking a picture with a camera, where the jpg format is a standard and the metadata is a standard. If you are, say, storing data from a radio telescope, the metadata is more like taking a snapshots of a production database. Over time you might track additional data, like how much the telescope slewed recently, how much radio interference was nearby according to some new algorithm, etc etc.
Your data formats are constantly changing. So your analysis scripts are constantly changing as well. But this sort of thing doesn't just maintain itself. You very often need to add code like, how do you handle versions of the data where column X is missing. A research project might spend a year gathering data, and change the data format a tiny bit ten times over that year. If you publish something a few times a year, there eventually are a huge number of data versions and script versions that old publications rely on.
It isn't impossible to maintain data like this. You can have code that regularly runs integration tests and reruns past analyses. But most research doesn't operate to this level of "software engineering quality". One-off Jupyter notebooks, code that the developer only got working on their local machine, and so on.
I think we could do better, but it would involve hiring more software engineers and building engineering teams to support scientific research. It is not as simple as allocating more budget towards hosting large files.