How to make your scientific data accessible, discoverable and useful
nature.com
nature.com
Here's what you can do. Make a github account. Make an account on arxiv.org or biorxiv.org. Publish your work onto arxiv. Put all of your data onto your github account, including copies of whatever you put onto arxiv and including extensive markdown readme file that is a synopsis of the paper. Disseminate it by announcing it on your twitter or mastodon or rss blog or substack or research community discord and even put it to a site like hacker news if it's relevant. If you are with an affiliated institution then their press office will put a small release.
Yes this is better. Like if you are a scientist and you find an amazing new cancer gene then you should probably put it in the actual gene repository and not randomly on github. Because you're this hypothetical scientist you probably know the exact niche place that is most appropriate for that already. Like GenBank or whatever is the one that there was a scandal about the pre-covid coronavirus genes were mysteriously deleted from it.
Their response in this instance was to fund SRA-in-the-cloud and other ventures, such as having PIs in well-connected locations like U.Chicago rent datacenter space in an exchange, negotiate very cheap hosting and bandwidth, and then give people access to compute and data either there, or in AWS (https://www.uchicagomedicine.org/forefront/news/university-o...)
This still doesn't address the "high quality metadata problem", which IMHO is the NP-complete problem of biology.
1. 9 Jupyter Notebooks attached (with HTML converts to the journal), all figures and statistics generated in notebooks, all commits of the entire 5 year research process versioned in a Gitlab repo, using Jupytext for clarity
2. All base data shared, using HyperLogLog to reduce privacy conflicts
3. Versioned docker image added to our registry, which includes Jupyter and the analysis environment used for the study (Carto-Lab Docker Version 0.9.0 [2])
4. Post acceptance, I published another notebook how other users can load and work with the shared data, including making (limited) additional inference [3]
5. For the peer review process, I added all (redacted) notebook HTML files to a Github repo [4]
It was a fun experiment where I tried a maximum of transparency in research. This maybe added 1 full year of additional work, but I still don't regret it. Given the quite specific audience for this paper, I doubt that anyone has ever tried to open the Jupyter Notebooks - I even doubt that reviewers had a look at them, at least by judging from the comments during peer review.[1] https://doi.org/10.1371/journal.pone.0280423
[2] https://gitlab.vgiscience.de/lbsn/tools/jupyterlab
[3] https://kartographie.geo.tu-dresden.de/ad/sunsetsunrise-demo...
Effortless reproducibly plus open source must be the gold standard, regardless of the form it takes.
As an aside, back when I used to live in a van, I wanted an app that could hyperlocally find a parkable location on a street or parking lot with the most shade for a given time of year. It seems roughly estimatable if high resolution height information were available and combined with the the sun path. It could also be useful if one wanted to locate housing with minimum insolation.
[1] https://digitalscience.figshare.com/articles/report/The_Stat...
[1] https://www.digital-science.com/tldr/article/seven-million-o...
This includes other lab members on the same project...
A facility like CERN could have an accurate model of the equipment available, you just add the description of your experiment and the results to it.
Imho there aren't enough tools to discover scientific data.
I think the publishers (or maybe the universities? or anybody at the center of a community of experts really) should host an API which maps set of CTPH hashes to URL's (or ideally, CID's for use in something like IPFS). The goal would be that anybody (author or otherwise) could attach metadata after publication.
Maybe it's criticism, maybe it's instructions on how to get the included code to run, maybe it's links to related research that occurred after the initial publication...
Suppose you have metadata to attach, you generate CTPH's for the article, pick a subset of them which corresponds with the location you want to anchor your metadata to, and upload the pair to the context aggregator (these would likely be topic-centered, so if it's a biology paper you'd find a biology aggregator).
When people view the paper, they can generate the same CTPH's and query the appropriate aggregator, and they'll get the annotations back which link locations in the article's text to metadata that, for whatever reason, was not included in the original publication.
I want to use CTPH's instead of DOI's or somesuch because they don't require a third party to index the items for you, and they still work even if you have only part of the article (like maybe the rest is hidden by pagination or a paywall). You could do a speech-to-text transcription, annotate it in this way, and somebody else who generated the same transcript could then find your annotations without ever creating an ID for the speech you're annotating.
a) Massive data volumes (~100 Gb - 1 Pb/project)
ai) This means that data is typically stored on limited access machines like HPC clusters
bi) This also means that shipping this data around is financially expensive, and cannot be supported purely by small client machines
b) A low number of seeders; scientific data is not exactly popular, and there may be network restrictions on uploads through the typically used networks;c) The requirement for a data legacy; torrents are fantastic for ephemeral data (e.g. operating system builds), but are terrible for data that must be archived and kept for potentially decades to centuries.
It would be challenging to find a solution robust enough for CERN type data but also simple enough for an n=3 undergraduate research project (that may have yielded some interesting results).
I don't know what the solution is there. My intuition is that university libraries could be involved, and that a data librarian could help you get your small study into shape or be embedded at a percentage effort on a large study.
It only adds. I don't understand how it subtracts.
Hoping IPFS makes it someday because the idea is great.
> IPFS or torrent are the best options for distributing data
And suggesting that IPFS is not a good option for distributing scientific data due to its complexity.
In this discussion it means "distributed" in the "made available" sense.
The parent did not use IPFS because it didn't work for them. So no, the data was not distributed via IPFS.