Gene name errors are widespread in the scientific literature
genomebiology.biomedcentral.com
genomebiology.biomedcentral.com
I don't know if this can be fixed and how. Lots of people seem to have their own process and they neither understand that the tools can be used more efficiently / correctly, nor question the long, manual process they follow currently. They basically require an "intermediate excel" course which ends with "if you're doing copy/paste 3 times in a row, you should look for better solution". I'm not even questioning the use of excel at this point...
While the tools often are crude, a problem is that biochemical and molecular biology researchers lack the bioinformatics skills and other IT skills that often is badly needed in modern genetic research. The core competence of these researchers is in the lab, and bioinformatics skills is outside of that. That you see effects like conversion of gene names in excel does not surprise me at all.
You also see the same problem in other areas where researchers need skills outside of their usual comfort zone. The classical example is statistics.
Part of the problem, as a sibling mentioned, is that many people get into these kinds of fields from non-programming or even computer-illiterate backgrounds. At an undergraduate level, even in heavily numerical disciplines like physics, there is relatively little coding. Even then, it's tedious crap like F90 and bulletproof C. So people get into a Masters or a PhD in a field they love and are suddenly confronted by data analysis, and they have no idea what to do.
I've spoken a lot to my girlfriend (an astrophysicist) about this, as she's in this position herself. It's not that she isn't smart, she just has little to no experience with data wrangling and she's (in my opinion) been inadequately trained. I've solved things with Python one-liners that she would have spent literally days doing manually. We've had conversations along the lines of:
"But why does this file have 700 lines of input data filenames hardcoded? You know you could write some code to grab them for you?"
"Yes, but in the time it would take me to write the code, I may as well do it by hand." [In the end I wrote a 5 line regex to do it]
So I can assure you, people question the manual process and they think you're magical when you show them a quicker way of doing things. However, often there isn't the motivation or confidence to try and be magical themselves.
She spent a while setting up a simple Excel sheet into which they enter all test results and now all the statistics are calculated automatically using proper methods.
This kind of problem is rife everywhere. People don't know and are not interested in how to use computers to help them because "it's too complicated".
Excel doesn't help.
Every job I've ever had, I always look at how my predecessors do the job, then automate the crap that I don't want to deal with. :/
https://bmcbioinformatics.biomedcentral.com/articles/10.1186...
Bonus: the number of errors is positively correlated with Impact Factor (a tool used by statistically illiterate administrative types to judge the quality of research).
Remember, these are people who don't understand data types, or if they do they are to lazy to declare them.
I don't see them formulating a correct SQL query. Or use any kind of programming language that has strict typing.
It would have to be a custom-tailored system that knows about nomenclature in the field. Sounds not very efficient.
Well, neither can programmers, since SQL injection is still the most common security vulnerability in software, so I agree, SQL is probably a bad tool for the job :)
> It would have to be a custom-tailored system that knows about nomenclature in the field. Sounds not very efficient.
Is Excel a custom-tailored system that knows about nomenclature in the field? The article seems to explicitly argue against that.
I don't think it's a failure to understand data types. It's a mismatch between what you expect the software to do, and what it does by default. Unfortunately, Microsoft has steadfastly refused to allow any way to change the auto-formatting options(check out some people really pissed off for being treated like children here - http://answers.microsoft.com/en-us/office/forum/office_2007-...).
There's two useful prongs of attack here - one is to somehow force Excel to conform to the expectations of researchers - perhaps an extension that works to prevent the most egregious cases of auto-formatting gone wrong? The alternative would be what you suggest - creating and marketing a custom solution, the problem there is that you'd need either buy-in from a significant number of researchers to spread it, or you'd need to replicate a lot of Excel features to make the transition smooth for others.
Sure, but converting "DEC1" to "December, 1st" is not egregiously wrong, it's a valuable feature and in most of the cases the expected thing.
You could argue that all you can do is mitigation since CSV files don't offer much ability to influence how Excel will load them, and all you would need is one improperly-configured Excel along the pipeline to break the data. However, this is a significant mitigation - Excel apps would be configured once, and you would deal with a situation 1% of the time, and the solution would be trivial(just configure it!). Instead now you're dealing with the problem every time, and the solution(just mark the cells!) takes a lot more effort.
I don't think Microsoft necessarily has incentive to add this configuration(the science community as a whole is probably a tiny blip on its radar), but this is why we create modular and extensible software - so others can tweak it to their liking.
Putting an option in to disable something that only a miniscule number of users would ever want is not sensible.
Note bene: I'm only talking about this DEC1-like conversions. I agree that there are lots of conversions that are annoying.
The problem isn't with SQL, it's with the idiotic idea of gluing SQL queries from strings. See also: template languages for web pages, i.e. gluing strings to build what is really a tree.
--
I think the hate against Excel is unfounded. Sure, it's not the perfectly suited software for the domain - in the same way like a dedicated fruit slicer is better than a knife at slicing fruits, a dedicated fruit peeler is better at peeling them, etc. but one knife can do all those jobs pretty well on its own, while also being able to do countless other things. Excel is a very versatile and powerful tool, and this power comes from its flexibility. The solutions often proposed, like a "proper" database-backed system, often involves a fixed workflow and having to call in IT support every time something doesn't fit that workflow (which is always).
So IMO - yes for people dealing with data using Excel. Yes for them learning a programming language, SQL and a generic database system. But strong no for forcing them to use domain-specific dedicated "tools" that impose a particular workflow on them.
And is available for free - excel came with university / corporate license to all desktops for no extra cost (in most places relevant for this discussion). And popular + easily available - you need to convince IT to allow it on the network / preinstall it on the provided systems.
Yes, I can probably code something entirely in SQL to do it, but it will not be portable across SQL implementations; and yes, I can code that as a variable in my program... but neither of them seem to be as fluent as the way Excel does it.
I wouldn't use Excel for prod, but it comes in handy for a lot of small dumb shit purely because of =.
If portability is really important, I think your best bet would be a view that adds the calculated columns.
Unfortunately, that's still not as easy as using Excel, by a long stretch.
Sure, the vast biotech community outnumbers those few people who work in some other jobs (or use Excel for private stuff) and enter dates or floating point numbers.
Excel is trying to limit the effect of sloppiness on the part of laypeople. They usually don't set cell formatting to date or number. Excel really does the right thing there.
And if you're a scientist, presumably with a respectable degree and professional responsibilities, learn to use your tool correctly.
Nice how the correct ways to use Excel are listed as cumbersome workarounds.
1. Out of a thousand people entering "SEPT2", how many mean "September 2nd"?
2. How many of them use that shortcut intentionally?
3. How many people mean some gene?
My answers are:
1. almost all of them
2. A respectable part of those in 1.
3. A dozen?
And far more than a dozen people mean the gene, come on.
And even if you want to type september second, trying to guess a year to attach is not the best idea.
Pfft, I bet you only say that because you have a low h-index.
Even counting citations is better, although if admins would actually read the papers, that would be better-er. Even NIH intramural people talk about impact factor like it means something. Pretty fucked considering that the publishers of glamour rags have vested interests that are very nearly the opposite of "careful scholarship".
[1]http://flybase.org/reports/FBgn0036414.html - short symbol is "nan"
https://baselinescenario.com/2013/02/09/the-importance-of-ex...
Then there was the infamous Reinhart-Rogoff paper in economics, used to justify harmful austerity policies worldwide post-2008, that came to false conclusions using a row formula that wasn't updated.
https://www.washingtonpost.com/news/wonk/wp/2013/04/16/is-th...
I've seen co-workers getting wrong results due to a wrong copy-and-paste, where no attention was paid. This is no bug: it's a misuse off the program.
Know your tools is always a good advice.
Ultimately though coming from the world of algorithms and nicely organized data, it was frustrating how disorganized the nomenclature seemed.
This is just one example. The entire gene naming area is a pile of bollocks.
HGNC seems to delight in nonsensical renaming. MLL (Mixed-lineage leukemia, a gene rearranged in many leukemias, particularly infants) is now KMT2A ("Lysine Methyl-Transferase 2A"), yet gobs of its fusion partners are canonically named, say, MLLT3 ("Myeloid/Lymphoid Or Mixed-Lineage Leukemia; Translocated To, 3").
But wait, what's its translocation partner? Oh, right, the Gene Formerly Known As MLL. Who thinks this is a good idea?
Journals typically insist on the latest HGNC (HUGO Gene Nomenclature Committee, a subset of the HUman Genome Organization) symbols, whatever they might be (see above). Not atypically, in review, someone will ask to use the old name because that's what they're used to. Best of all is when reviewers each suggest using a different name.
Science!
Yeah. Rigor is not really a concept to which even molecular biologists are at home, which is understandable and probably explains why their parties are much better than those in a lot of other disciplines, but maybe a little hard to get used to in some other ways.
Do you have some references to that project? Reports, comments, maybe software?
That's my single biggest suggestion for biologists: know that professional computer scientists and programmers are desperately hungry for interesting data and would love nothing more than to help you design the project up front so they don't get sucked into the vortex of technical debt that will swallow your project if you don't set it up right early.
This still doesn't actually tell us if the problem is getting worse. Or if it does it is badly worded. Even assuming this 3.8% is derived from their own data, you need the number of papers published that contain genelists (I would imagine this has probably risen faster than the number of papers overall).
In other words, the authors should have plotted the error rate over time rather than the number of errors over time.
The meaning changes. Depending on who you are talking to, What book you are reading, What part in a book you are reading.
Did using correct data and analysis pipelines actually matter to the conclusions these authors came to?
I'm going to make a sign and picket Microsoft when I'm next in the area.