Handful of Biologists Went Rogue and Published Directly to Internet
nytimes.com
nytimes.com
That didn't really happen as expected. In my own chosen field (biology) it happened much more slowly than I hoped- physics (with arxiv) was far better. However, just getting PDFs on biorxiv is only a small part of the long game. I did not appreciate the huge importance placed on publishing in high-profile journals has on one's career trajectory and how large a role that would play in slowing the transition to free publication and post-publication review.
The long game is to enable the vast existing resources, and the new resources, to be parsed semantically by artificial intelligence algorithms. We've already passed the point where individuals can understand the full literature in their chosen field and so eventually we will need AI just to make progress.
Do you realize that many researchers would be ashamed if others saw their latex code?
A friend worked for a journal, he kept a folder of the worst that he had ever seen, and could rant for hours one some gems. Including some papers that were beautifully typeset, about elegant combinatorics/type theory, and were coded on the latex equivalent of a spaceship cobbled with recycled plastic bottles, scraps and duct tape.
Do you imagine going through thousand lines macro tex file that grew over the years, when Latex is one of the worst mess ever, with conflicting packages, absurd syntactic 'features', and legacy crippling flaws that make debugging an horrible task?
Please, continue to compile your .tex.
I really don't get this argument; I see it come up regarding open data for reproducibility too. TeX/LaTeX source is just an extra piece of metadata which might help someone out in a few years time. That's nothing to feel ashamed of, and it's certainly less shameful than not providing the code. If the code allows flaws in the research to be exposed, that's either something to be proud of, or (if it was intentional) it's not the code that's shameful.
My own LaTeX is mostly an inconsistent mixture of copypasta from stackexchange. That doesn't mean I'm ashamed of it, and it's all in public Git repos at http://chriswarbo.net/git if any human or bot actually cares to look.
Speaking of which, the commit histories of those Git repositories aren't very good either. Lots of commit messages like "Fixes", or "Oops, typo in last commit", with multiple changes bundled into single commits, atomic changes spread across multiple commits, etc. But, in the same way as my LaTeX source, it's better than having no commit history, which would be the case if I set myself higher standards.
Of course, I can follow higher standards, but that's when I'm paid to and/or when collaboration is important. For personal stuff, I'd rather invest the time I save interacting with Git on writing more tests, or refactoring, etc.
Becoming more proficient with TeX saves time and makes for more re-usable typesetting structures, which allow you to accurately do what you want faster and faster.
It's similar to using emacs or vim. You could say, "Why should programming devotees expect others to adopt those editors?"
It's about productivity. I'm a productivity devotee and when I hear people use short-term thinking to discount how valuable the up-front costs for learning TeX (or emacs/vim, etc) are, it motivates me to help them see why it's not true.
Of course, I don't want to be dogmatic either. If something works for you and your personal utility function discounts the value of things like TeX or emacs, that's fine.
Not necessarily, there's an opportunity cost to all the time you spend become proficient with TeX. If there are other products that offer sufficient control over typesetting with less time and effort, and you don't anticipate needing the advanced features that are unavailable in competing products, then becoming a TeX expert is a complete waste of time. I'm happy to concede that it's probably the best tool and certainly the best value for money, but if someone's needs are quite limited then an adequate-but-easy-to-use tool may be preferable to the superior-but-user-unfriendly one.
Actually, that's a great comparison. I don't use either of those (yeah, I said it) because of their arcane interfaces. I don't doubt for a second that wizards with those programs, and LaTeX, can do great things, but IDEs exist for the same reason word processors do - to make it easier to focus on creating.
IDEs exist for the same reason 7/11 exists. Convenience. You pay a higher price (lost productivity) to experience certain comforts (clicking a button instead of issuing a keystroke command).
My experience has universally been that IDEs only get in the way of focusing on creating. They might be necessary when you are literally just beginning with software and you need a lot of scaffolding to help you -- and I don't think anyone would begrudge university students using IDEs.
But not using a power editor once you are an experienced programmer is hard to understand. I mean, it's so important that it was even included as a whole chapter in The Pragmatic Programmer.
Paying the up-front costs of proficient editor usage is like paying the up-front costs of wearing braces so you can have straight teeth or wearing correcting shoes so you can walk properly. Of course it's more convenient in an absolute sense to simply not wear braces or corrective shoes. But that's not the point. The point is that the value you get from the outcome (e.g. straight teeth, proper walking, or dramatically increased productivity when developing) is far greater than the cost.
In other words, I think you're using a hyperbolic discount function. You are placing such a high value on absolutely immediate "focus on creating" that you incorrectly discount how much more focus on creating you would be able to do with the power of a real editor framework.
For the 90%+, it's an arbitrary checkbox requirement to be checked on the path towards something else.
No it is about aesthetics, which are mostly irrelevant for 99% of all research papers.
Unless you think that reading red text on a blue background in payprus would be an equally enjoyable reading experience it more than simple aesthetics.
> Fine, TeX is an ugly language that only a programmer could love, but in over twenty years nobody has yet been able to replace its expressive power and flexibility
Which doesn't make it less ugly. If we can't cure common cold yet, that is no reason to praise it and making it sound as if it's a great thing. It is a problem and so is the need to use ugly language to produce aesthetically pleasing papers. The fact that we don't have solution for this problem yet does not turn it into non-problem. It just says nobody yet was good enough to solve it.
> It's similar to using emacs or vim. You could say, "Why should programming devotees expect others to adopt those editors?"
Exactly. There were times when I wrote emacs lisp and vim macros, and had time to dive into arcane mechanics of it. I don't have that luxury anymore. I need the editor to just work and do what I need to be done, without me investing significant time into it. Maybe if I spend couple of days on learning stuff about vim I could do in one minute when now I do in ten, but I'm not sure it pays off anymore, and frankly I have more interesting things to learn. I do use vim (not emacs anymore) but for me I came to the conclusion that added value of diving into the arcana with the goal of being more efficient later vs. brute-forcing through the immediate tasks is just not worth it. Later you'd need a different thing than you thought you'd need anyway.
And, note, I am a programmer. I can imagine how less interested would be a person who does not spend their days digging into the code anyway.
The joy of LaTeX is that it handles the aesthetics for you. You can focus entirely on structure. Or that's how it works when you have good boilerplate to start from.
One LaTeX-agnostic example is when people abuse hyphens-like this-which makes it impossible to parse the sentence without reading it several times.
I'm one of the founders, and we're also making it possible to directly submit projects from Overleaf to repositories such as the arXiv and bioRxiv, and to traditional journals, to help speed up the submission and publication process.[2]
Ultimately, this would fall over as people accidentally made concurrent changes and it would be hard if not impossible to merge between two documents. Doc files were never designed to be put into a version control system with merge semantics. Some people use shared folders with locking but locking doesn't work well over file shares.
I'm not sure but I suspect now people use dropbox instead of email; it's not clear to me how they manage conflicts and merges.
I rather my format be a text-line-based non-rendered format that I can render down if need be, but use the full power of patch and diff to manage conflicts.
I am certainly aware that many scientists don't want to publish their code out of embarrassment. It was a surprise to me, but I've heard that feedback.
I investigated various alternatives after I found LaTeX too arcane, including DocBook. It's still not entirely clear what solution exists to the collection of problems I have (want to semantically mark up my paper, want it to display nicely regardless of screen, etc).
Are you aware of pandoc?
About the 2nd part, may I rephrase it as PDFs are the docker of publication .. (Adobe would probably be very proud of that).
I disagree. The text is indexed and searchable.
> the meaning of the papers isn't addressed by that.
I agree. I wonder if starting with research papers that have lots of symbolic logic (e.g. math proofs) would be the easiest starting point for a system like this.
The exact same think could be said about HTML, but that hasn't stopped people...
I understand the words, but in this context it has me a bit confused as to what problem you see AI solving.
The semantic parsing part is an annotation process. AI algorithms are required to make accurate annotations - a huge amount of context is required.
So if authors instead marked up their papers such that each mentioned entity's semantic meaning was obvious, it would make it much easier to AI that scans all papers and generates hypotheses.
Thank you very much for unpacking it and all the best in your career in Biology! I am in tech now, but studied Biology at Texas A&M for undergrad so hearing the words in your response reminded me of the good ol' days!
I want AIs to automatically find conflicting papers/hypotheses, and propose experiments that resolve the ambiguities.
I mean, I don't know much about LaTeX, but I doubt there are Elements for "Hypthese", "Definition", "exact reference" etc. If you would have those, described in a structured, simple language - then I guess, it will be much easier to process those Information for a KI, when the context is clear.
Or something pythonlike (also supported by a IDE):
hypothesis:
(indent) blablabla link:"link_to_Element_in_paper"
When I spoke to them they said if they had an AI that could do as good a job as people, they wouldn't need contractors. I think Google's approach would be to contract a bunch of scientists, have them read and interpret the papers, then use that data to train a deep net that could do it more accurately (you need some baseline humans to act as golden standards). This worked well for Google in several publicized examples, such as discriminating house numbers from numbers on cards in Street View imagery.
Who is it primarily that is concerned with such things? (Curious.)
Is it fellow academics? Or university administrators? Or government agencies that control grants?
Whoever it is, these are the people we need to work on, it seems.
Yes.
Fellow Academics: The people who will be discussing your tenure case, writing letters of support, nominating you for prizes, etc.
University Administrators: The people who are the final word in your tenure case, and who do things like evaluate how well your department/college/etc. is performing.
Grants: Be they your fellow academics in the form of reviewers, program officers, etc., the prestige of your publications will likely matter to them.
Publications are the unit of currency in academia at the moment, and the incentive and evaluation structures at almost every step are oriented around that. Going your own way is a laudable step, perhaps, but largely the luxury of established researchers who have already had their prestige publications, or those willing to take the potential hit to their careers for the principled stand.
If you could get them to all drop caring at once, that would work. However, if any critical mass of any majority of groups sticks with it, those who try to ignore the high-profile publication wheel will just get squeeeeezed out. Simple game theory.
Yeah especially considering the roles CERN and NCSA played. Things might be different if they had published there immediately instead of staying on their journals.
However, compared to the number of potential biologists or manuscripts that could be published, it is just a HANDFUL.
Rouge? They're positively crimson, darling.
But to stay on topic, I'd say that the impact of blogs and wikis have been felt, it is just up to the gatekeepers, the powers-that-be to accept this as relevant. Like somebody said here, "publications" and Nature/Science papers shouldn't be nearly as influential (wrt career trajectory) as they were in 1986.
The system is aged and inefficient (some would even argue it's rotten) and IMO comprehensive changes are needed. Like racial or gender discrimination can't be addressed without changing the social rules people live by, the current academic system that's rather elitist, non-inclusive, discriminatory, often more biased and less fair than many think needs to change substantially.
Such change will be aided by important people setting examples (and often going back to their old ways). However more substantial change is needed on multiple levels, most importantly: academic leaders and funding agencies (run by the former) need to stop looking at who's who and how many Nature/Science/insert-your-fancy-journal papers does the person have. For instance, the culture of applying for grant money with work that's half done to maximize one's chances needs to stop and so should the over-emphasis of impressive and positive results.
Additionally, publishers exploiting everyone need to die out and as long as these researchers "go rogue" with a single paper (rather than for instance committing to publish 100% preprints and >75% open access), not much will change.
You need years of intense studying even to understand the current state of the art in a chosen scientific field.
Experiments require more money too.
It seems to me that any given field of "science" isn't any harder than dropping into a new area of software engineering. There's a lot to learn, sure, but there's a lot to do as well and there are areas where an complete rank amateur can make significant contributions.
It's still true that every day I discover something I have no idea at all about. But it seems that this is fairly normal in science - people know a lot about a single specialized area and not much outside it.
There are 3 sentences in my comment. What specific claim do you consider to be wrong and what is your evidence?
The crucial difference is that an amateur can provide a meaningful contribution to an open-source software project while the same is unlikely in (modern) science.
http://motherboard.vice.com/read/meet-the-amateur-comet-hunt...
http://io9.gizmodo.com/5841287/the-story-of-the-woman-who-di...
You need years of intense studying even to understand the current state of the art in a chosen scientific field.
From personal experience, this isn't the case. But you'll complain about citations, so: https://www.quantamagazine.org/20160313-mathematicians-disco...
Experiments require more money too.
See the examples provided above.
[0]: https://en.wikipedia.org/wiki/Citizen_science
[1]: http://journals.plos.org/plosone/article?id=10.1371/journal....
edit: clarification
The only wrinkle to this is that citations in more popular venues are more likely to be seen and re-cited.
I.e., actually getting real science done fast, taking advantage of how the long journal process was bypassed.
This whole thing needs to start at the level of the funding agency, namely NIH. Publishing in a good journal is a prerequisite to getting a grant. Try getting an R01 on a Biorxiv paper. Not gonna happen.
Awesome.
I'm always curious to know how physics, as a field, got over this and related humps. Was it just easier to get everyone on board because it's a smaller community? Were the journals not as savvy to the fact that pre-prints are not really in their interest?
It was a huge shock to me, coming from CS: knowledge sharing has such a high value in our community.
[0] Richard Posner ("pharmaceuticals are the poster child for the patent system.... Most industries could get along fine without patent protection.") http://www.theatlantic.com/business/archive/2012/07/why-ther...
[1] Notch "I would personally prefer it to have those be government funded (like with CERN or NASA) and patent free as opposed to what’s happening with medicine, but I do understand why some people thin[k] patents are good in these areas." http://notch.tumblr.com/post/27751395263/on-patents
On the one hand, I have sympathy for the very real concerns people have about bad (potentially harmful) health information going out. On the other hand, I think the primary reason it is dangerous to begin with is a lack of a culture of vigorous discussion. It is only dangerous if it cannot be thoroughly hashed out. Unfortunately, it seems there is no place where that can happen.
I have no words for my disgust at the lack of honesty that scooping would demonstrate.
Speeding up the knowledge sharing and to solve more problems more quickly is a good thing IMMO. As the same article pointed out, Physicists has been releasing research results in preprints since 1990s.
Figure 5. The statistics ....
to Figure (5) The statistics ....
because the reviewer liked () format better, although (IMO) the new format is so ugly.Another reviewer saw the comment and had a really nasty debate. The paper was published with the original format nonetheless.
Fractionalization is generally undesirable, but it's plausible that bioArxiv can make tweaks that accelerate adoption among biologists compared to arXiv.
I have never once found a paper of particular relevance to my research (geology/geophysics) on arXiv. I haven't looked very many times, but after striking out several times, why keep trying?
If there was a geoRxiv then I would probably browse it more regularly because the chances of me finding something relevant would be much higher.
For that reason, and another, I kind of disagree that fragmentation (or fractionalization) is undesirable. The second reason, which is not unrelated, is that–at least with peer-reviewed journals–the quality of the work in field- or subfield-specific venues is often far higher than in the sort of pan-scientific journals like Science or Nature. I think a lot of it has to do with the quality of the reviewing/editing, but a lot of it is that the papers have to be written for and justified to a wider audience that wants to be wowed and doesn't know the background well enough to evaluate the science for its own sake.
If I write a paper and send it to Tectonophysics, I know that the readership will understand what I'm doing and why, and I will write the paper accordingly. If I write the paper for Nature, then I have describe and justify the how and why to a wide range of people from my peers to journalists for phys.org and the NYT. Sometimes that's fine: If I find out that the Seattle fault is loaded and ready to pop, the press, policy makers and citizens need to know. But if I find out that the stress field on the Seattle fault is largely determined by the topography near the fault and that has some persistent influence on how an earthquake rupture propagates on the fault (but doesn't necessarily change the seismic hazard) then I don't need to go through the rigmarole of explaining and justifying any of that to anyone who isn't intrinsically interested, and maybe more importantly, I don't have to explain (i.e. gloss over) the subtleties, ambiguities and caveats of the work to an audience that lacks the relevant background. This simply allows me to write a more clear and more honest paper.
This brings up a tangent that is relevant to the broader topic of self-publication: You always need to write to a specific audience, and with a journal you know who that audience is. With a blog or a website, you don't necessarily. That may be fine but it can trip a lot of people up, and make the writing much worse.
Let me reply to your points in particular:
> Discoverability is a big part of it as well, at least at the user end...If there was a geoRxiv then I would probably browse it more regularly because the chances of me finding something relevant would be much higher.
Are you just talking about the individual subject areas (ecology, genetics, etc.)? The arXiv has those as well, which you can subscribe to. Few physicists subscribe to the entire thing. (Incidentally, folks may find https://scirate.com/ filters a bit better.)
Of course, the arXiv doesn't have a dedicated biology section, much less sub-divisions, but this is because there hasn't been enough interest.
> The second reason, which is not unrelated, is that–at least with peer-reviewed journals–the quality of the work in field- or subfield-specific venues is often far higher than in the sort of pan-scientific journals like Science or Nature.
The main point of the arxiv is to put everything in one place which is permanent, searchable, sortable, freely available etc. Filters generally come from elsewhere, such as the aforementioned sections or SciRate, or by simply looking at arXiv papers published in certain journals (without needing journal access).
Obviously, in the absence of additional filters, the bioArxiv won't be useful as filter either.
> If I write a paper and send it to Tectonophysics,...
You'll find that there are plenty of popular-level papers on the arXiv sharing space with highly technical ones. This is generally noted in the abstract. While the arXiv is not meant for public consumption, there are plenty of filters that try to pluck out accessible papers, e.g., the Physics ArXiv blog (which isn't as official as it sounds) https://medium.com/the-physics-arxiv-blog
So why does it seem like Greider's paper received none of that assistance? It's typeset in double-spaced Arial. Both double-spacing and Arial are immediate indications of unprofessional typesetting.
It's probably more likely due to a lack of demand. There's only about 4k articles on all of biorxiv in any case.
> If university libraries drop their costly journal subscriptions in favor of free preprints, journals may well withdraw permission to use them
withdraw permission to do what exactly, and enforced how?
If somebody is looking for an idea for a new venture, this is a problem, yet to be solved !
There definitely needs to be something that makes it easier to parse and otherwise interact with a PDF. But, for network-effect reasons, it's probably easier to introduce a parseable overlay for PDFs than to replace the format wholesale.
But on the flipside, the TeX document will often be ugly no matter how you are reading it, and the author can't easily apply tweaks for aesthetics or readability--many try, hence the TeX markup horrorshow.
One area I have never understood (apart from editor laziness) is letting the authors write the article title and abstract. So many great papers are overlooked because the authors wrote a boring or misleading title or abstract.
At least I hope humanity is getting more sophisticated. What is the median of the age of Trump supporters and the one sigma std dev? That would be an interesting statistic.