Scientific Data Repositories
nature.com
nature.com
https://febs.onlinelibrary.wiley.com/doi/10.1002/1873-3468.1...
All I have found is ScienceVerse [1] that aims to develop a syntax/"a Grammar of Science"
I recommend these article for discussions of how to share when you can't share all your raw data publicly because of participant privacy:
https://www.nature.com/articles/s41586-020-2766-y (I am a co-author)
https://www.nature.com/articles/s41576-020-0257-5
There's some more discussion in the article linked in the parent to but it's mostly not about that.
I worry a bit about people just dumping data into large repositories, without thinking much about the format or the later uses, but only focussing on a checklist that needs to be ticked off to get that precious bean (publication) for the bean-counters (deans).
(For those that haven't spotted it, these are permitted under 'Generalist repositories')
> An advantage is that there are local people who can help the researchers with the process, e.g. in setting up useful metadata and so forth.
I helped set up such a service nearly 10 years ago, and still help run it. There undoubtedly are advantages to depositing with us for the reasons you mention, plus we permit far larger publications than most services (our largest are around 1TB).
However we are a large, general university, and so have to deal with deposits ranging from theology related images to CT scans of fossils specimens to synthetic chemistry data. And all points in between.
Being general limits our capacity for detailed help concerning metadata and format standards for researchers since we just don't have enough data librarians with these specialisms. So my advice is to use a community established repository where available (UK Data Archive is a good example).
You are right about people just dumping data. Since 2015 (iirc) researchers have been expected by funders and publishers to plan their data storage and make it available ultimately. That doesn't necessarily lead to quality publications, though our reviewers try their best.
To paraphrase a researcher "I intend to give this process the minimum required". (This is not a typical response, happily)
Regarding your dumping data comment. Yes that's certainly an issue. The problem is that researchers are required to do this, but there is no real consideration of the time this takes (it does not make a difference to your career/reputation etc. if you publish good or bad data), it's really out of the expertise of most researchers and there is little help provided by universities.
Yes, but good data could help people trying to monetize your research and give you nothing immensely.
I work in research in a public university in Canada. IT is basically tech support, fix-my-email. There's no chance they would support hosting our data or any other sort of service.
The university expects that researchers self-fund their own stuff using grants. Need laptops for your grad students? Grant money. Need a server, bunch of disks and a sysadmin to care for it? Grant money. Which is only realistic if you're a huge lab with millions a year in grant money. And even then, what happens whn this grant runs out? Your new grant does not pay for hosting some 10 year old data, all money is earmarked for your new project (literally, it would be illegal to spend the money on another project). So the old hosting quietly goes away.
https://www.carl-abrc.ca/advancing-research/institutional-re...
https://www.carl-abrc.ca/advancing-research/institutional-re...
https://zenodo.org/communities/hoffmanlab/
GitHub even provides a guide to depositing a GitHub repo in Zenodo:
I ran into this when doing some OCR experiments[1], finding acquiring data and pre-trained models to be the most time-consuming part of the enterprise. This ended up adding enough additional hassle that I didn't manage to get anything really interesting going, although figuring out how to containerize other peoples' code was educational. Personally, I think I'll be relying on some combination of institutional repositories + torrents/IPFS for any large datasets/models I end up releasing in the future.
-----
But that probably won’t happen because torrents are a dirty word due to illegal activity and they also give up control of the data.
I noticed many of these repos are javascript-walled. Is there any kind of standard API through which you can search for repos and fetch datasets?
There are also services like JISC data monitor and of course the citation databases (e.g. Scopus) now contain datasets.