It's very much not an open source mentality. Everybody wants other scientists to share their data, but nobody wants to share their own.
It's very much not an open source mentality. Everybody wants other scientists to share their data, but nobody wants to share their own.
- Preparing protocols and getting them through ethical review committees
- Finding suitable patients and getting them on board
- 1 on 1 patient consultations, oftentimes multiple of those for each patient and spread through a few years.
I'd say the balance can easily turn to 90% data collection, 10% everything else in many cases...
Which I guess is why all efforts to try to split up the work in a more efficient manner have failed and the current model persists. This means that open academic research effectively remains a cottage industry compared to industrial, military, and other research types that can be kept secret.
I was hoping stating it all together like that would make a solution more obvious, but I don't think it did. Having researchers commission data collection and every user having to pay would create barriers to entry we don't want.
Unfortunately, it's absurdly hard to change what counts as credit towards a successful academic career, so this seems like a pipe dream. Academia is still waging a very uphill battle to get Research Engineer to be a long-term viable career path, despite the absolutely critical need for software engineering in all areas of modern science.
- fear of getting scooped: there's a general fear that sharing data makes it easier for competing labs to publish on the same topics, leading to the publishing lab "missing out" on publications
- fear of being wrong: sharing the data leads to more scrutiny, leading to a higher chance of a fatal flaw being found
- extra work: publishing data often involves jumping through extra hoops with little to no personal gain
I think there's a variant of getting scooped worth pointing out, which is that in some fields people are scared of being scooped by a low quality lab that fakes the results.
I had a very senior materials science researcher describe what he thought was a promising line of study get shut down because a low quality lab in china scooped him with results good enough for publication but not good enough that people would want to replicate. He was 100% confident their data was faked but getting grant money once that experiment was done was like pulling teeth and so he abandoned the route for other ideas.
Most problems in academia come down to bad incentives. Researchers are incentivised to publish for prestige and citations, which are poor proxies for improving our understanding of the world around us.
* The time horizon used to judge incentives vary person to person.
* Willing martyrs exist for various causes. Therefore, it is not just the incentives that matter, but the belief about those incentives.
* Self-sacrifice for others exists (e.g. a mother rushing into a burning building for the chance of saving her children). Therefore, it is not just the incentives that matter, but also the subjective weighting of conflicting incentives.
* A person may spend a great deal of time researching betting patterns for playing craps, none of which can overcome the house's advantage. Therefore, it is not just the incentives that matter, but the knowledge of the outcomes.
* Somebody may work in an oppressive environment for their entire life, rather than changing careers. Therefore, it is not just the effort required that matters, but also the relative amount of effort on different time horizons.
While I agree that incentives are the best way to move population-level behavior, and is usually a good starting point for understanding individual behavior, your categorical statement doesn't have those caveats.
The problem with that system is that the "value" would have to be determined in some independent manner, e.g. an audited cost+ basis, which wouldn't be reflective of the true value of the idea. Granting time-limited monopoly rights allows the idea to be priced via normal market mechanisms, at the cost of high transaction costs.
At any rate, you get the idea. The justification for academic research is that the private sector, supposedly, won't do long term basic science and that grant recipients will. That's a very dubious set of assumptions. We see basic science being done by private labs all over the place, most obviously in AI but also in many other areas of computer science, we see it in the biotech sector, we see it in semiconductor physics etc. The areas where there's no private investment tend also to be the most controversial areas of academia where many people have concluded that the results are systematically useless e.g. social sciences. And meanwhile what we see amongst grant recipients is systematically unscientific behaviour like knowingly publishing claims that aren't actually true, refusing to share data even after promising to do so, accepting unvalidated modelling as 'science' and a million other things.
If he truly believes the released data were fake, he can for sure publish his results. It’s not unusual for scientists to have different views on certain subjects. Some journals even have specific discussion sections for contradictory results. Finding the truth and correcting scientific records is also an achievement.
And at least, in the filed of biology/medicine, there are journals accepting “second first” papers, for example https://elifesciences.org/articles/30076 .
And quit a lot of publications are not grand breaking ideas. Incremental improvements can also be published.
I don’t understand your friend’s rationale to abandon the whole project he already started and didn’t want to find the truth and prove his theory (maybe it’s a different field thing).
You need the funding to get to that point though, and since it wasn't "new" they weren't able to get funded to research it. That's the meaning I got from the comment.
He has lots of ideas he wants to try, and he ranks them. The incentives for following this particular idea were reduced by the actions of the other lab, so it dropped in the rankings and he pursued one of the more highly ranked ones.
It's not about an "open source mentality" - almost all of my analytical code is online, and when I work with simulated data sets (the bulk of my work) it's freely available.
It's about incentives, and effort. A couple examples:
1) Primary data collection is hugely time and resource intensive. It's hard to be okay with all of that, getting a single paper out, and then immediately having that be expected to be freely open and available.
The argument that it's necessary for reproducibility is a compelling one, but also one that's much rare than what most people are looking for, which is to do new and novel research using your primary data.
2) Research has a potentially long time from data collection to publication. It's entirely possibly the reason isn't "I may want to work on that someday" but "A graduate student is working on that right now". I have been in the position of potentially having a graduate student be scooped on something, and it is terrible (it ended up being a false alarm). They're potentially also the person who put in a lot of the labor on publishing the paper in the first place.
3) Sharing comes with an unknown expectation of support, documentation, etc. wherein there's absolutely no incentive structure in place to support the work that takes. I write in time and budget to do so in my grants, but that also mean my grants have less science in them than other people's, and I have to rely on that mattering to someone enough to make up the difference.
4) Genuine data ownership issues. Health data is especially complex in this space. Do you want a simulated data set that will give you the same broad patterns? That's easy enough. But if you want the real data - that's potentially a very complex issue wherein even if I say "Sure, let's get that process started..." you're potentially looking at a year or more before you get the data.
I'd also add that primary data collection is risky. You can spend months to years working to get data, only to end up with nothing---not even a null result---if the experiment itself goes south. I lost over a year of my PhD when an animal I had spent ages training fell while playing. A classmate had a virus wipe out her colony of genetically-engineered mice. Using someone else's data obviates these sorts of risks--while doing a better job preparing you for non-academic careers too!
As a result, I think it's not totally nuts for the original data collectors to maybe get some kind of "risk premium." I can imagine a lot of ways to do it, but early/exclusive access to data has been one of them.
The answer for a lot of this is to put in the hard work of building relationships with people who collect a lot of primary data, and make sure your work amplifies their work in a way that benefits you both. But that's a lot of effort, hard, and uncertain in its own ways.
I definitely understand the worry when it comes to high profile discoveries, but surely the work a graduate student is working on, isn't usually that high profile?
And to be clear, I'm not downplaying the importance of anyones contributions, but rather why the focus isn't shifted onto where data is sourced from? If someone used your data, shouldn't it be immediately obvious if they also have data publishing requirements? And if thats true, why isn't it a bigger faux pas to use another labs data before that lab has even published their paper on it?
How much press was given to the second time the LHC discovered the Higgs, with higher precision?
Robert Dicke, in the running for "greatest experimental physicist in history", was working rapidly toward discovering the cosmic microwave background, with a keen physical insights driving the research. His team was beaten by a team at Bell Labs, who, when they discovered an irreducible isotropic source of noise emanating from the Universe, literally called Dicke [1] to help them understand the signal.
Who got the Nobel Prize? The team from Bell Labs. Never mind that Dicke literally wrote the companion paper explaining why the Bell Labs measurement was so important. Why? Because they were first.
Regarding graduate students: It is extremely common in physics to have the lead author (and coordinator) on important/cutting-edge research be a graduate student. It is only the largest collaborations where this isn't always the case.
Graduate students tend to follow a single thread of research from beginning to end, learning the ropes on the way, while faculty generally manage and mentor multiple threads (each with a student) simultaneously. When it comes time to publish, frequently it is the student who knows all of the details of the measurement the best.
For your final question, the answer goes back to "because being first is almost everything", combined with "we have to give the credit to somebody, right?". If someone notices that the eighth word on 50% of Guinness World Record citations is "person", should the credit for the discovery go to the person who noticed that interesting fact, or should it primarily go to Guinness World Records, who assembled and produced the dataset over decades of work?
Or, more-concretely, when someone applies for and receives fifteen minutes of time on the Hubble Telescope and, using a clever time-domain analysis, happens to find a repeating optical signal that blinks 13 times, then 53 times, then 13 times, then 53 times (as good a SETI candidate as you'll find), who gets the credit? Is it the researcher with the keen insight and good luck, or is it the huge amalgamation of humanity, spanning a half-century of science and engineering, that made such a measurement possible?
[1] https://theconversation.com/the-cmb-how-an-accidental-discov....
If you want a faculty position after graduation you've got one shot to create a strong research profile. Getting scooped once can close career doors forever.
People also often don't understand what grad students do. They do all of the research. 100% of it. PIs are funding the lab and providing mentorship. The first authors on major groundbreaking papers are usually grad students.
Getting scooped doesn't just matter for high profile discoveries. The example I was thinking of is the graduate student's research asking "Does this thing the field is doing make sense?"
That's a Yes or No answer. There's not a lot of interest once the question has been answered. Even if it can be published, it'll be a harder road, potentially finding itself in a less prestigious journal, etc. At worst, it's a write-off. Most graduate students are coming out with 2 or 3 paper dissertations, so the loss of one, or taking a hit on another one, can set someone behind, dramatically change the tenor of their job search, etc.
"And to be clear, I'm not downplaying the importance of anyones contributions, but rather why the focus isn't shifted onto where data is sourced from? If someone used your data, shouldn't it be immediately obvious if they also have data publishing requirements? And if thats true, why isn't it a bigger faux pas to use another labs data before that lab has even published their paper on it?"
I mean, it might not be the only thing that's been published on that data - that's the point. Once it's publicly available, it's publicly available, even if people are still working on it.
Indeed, one of the "Data is available upon reasonable request" things is "We have a student working on this, could you not?" (another being "You're a crank, and we're not sending you this to misinterpret/misuse it"), but a lot of open data folks actively dislike that for obvious reasons.
The academic researchers that generate these datasets and their funding sources such as granting agencies and donors want to see the money being used to generate something important. In this case it certainly in the researchers interest to avoid sharing the data for as long as possible to get as many publications as possible. This makes it easier to get more funding.
There is a push by journals now for researchers to include the data in publicly accessible repositories such as the NCBI SRA or European Genome-phenome Archive. These sites archive the data and make it available for researchers. However, in most cases they require a comprehensive data sharing agreement between your institution and theirs. Most of these agreements have extremely demanding requirements that make it very difficult if not impossible for a requesting institutes legal team to agree to. For example, the institute that owns the data may make demands such as "the right to approve any publications or research work created with this data", which entails sending any publication to them prior to submission and giving them veto power. I understand the reasoning for the agreements their intent was to protect private information from being shared publicly while allowing researchers to use the data. But I think the system is being abused now.
On the other hand, it can be very difficult to get past institutes to make data public. For example, sequencing data contains personally identifiable information and getting approval to share the raw data can be difficult. For prospective studies, Human patients need to consent to their data being shared and may not consent to it being shared publicly.
I can't comment on industry as much. However, I have colleagues who work at companies and publish numerous papers on the same proprietary datasets but refuse to share it with anyone. This is particularly challenging in fields such as cancer research where they may publish some fancy/superior new model for disease risk stratification on their own data but without sharing it's impossible for researchers to independently validate it.
We're experimenting with building a data collection system that uses adversarial/multiplayer game mechanics incentivizes researchers to share (gain co-authorship) and excludes them from the benefits (e.g. co-authorship) if the minimum submitted data threshold isn't met.
Still early days, but we figured this is the only way to get people to share data, even within other people within the same project / lab. Similarly, I'm getting a lot of ideas from HN on event-driven data collection and immutable data storage to help make the data more tamper proof.
The basic conflict is that it's hard/expensive/time-consuming to make data, and easier to write papers on existing data sets. So just split the work, and make it a legitimate occupation to create good data sets. As of now, generating the data by itself has no benefits for scientists, other than the papers they can subsequently write using it. Feels a bit like having to make a screwdriver every time you want to work on your car.
That needs to be tracked better, things are improving but until there are metrics and tracking as a common part of evaluation I don't think it will change.
As an experimentalist, I really want people to make use of my measurements, but I really don't want to be a co-author on a paper that makes wild and incorrect conclusions using my data.
I know people who were told by their PI to tell other researchers at conferences that they were working on the stuff that they already tried that didn't work. This would encourage competing labs to try out these techniques and waste their time.
Being scooped is a huge blow to a career, for faculty but more intensely for graduate students. With arXiv, you can be scooped at literally any moment rather than just on the known conference schedule. Being scooped was personally traumatic for me and I still have nightmares about the experience.
This leads to a massively toxic ecosystem of over-competition where individual labs win at the detriment of scientific progress.
Of course, arguing about the right answer is supposed to be what science is about. That doesn't make it any more pleasant.
Is there a way I could contact you?