Big Data Coming in Faster Than Biomedical Researchers Can Process It
npr.org
npr.org
Also, though there is a TON of 'data' coming in, most of it is not useful. For example, I have a 500Gb file of a stack of .tiff images per fish that I have imaged in a confocal microscope. I have a GFP filter on the scope and therefore only get the green part of the .tiff files exposed, the red and blue are just background noise. Also, most of the image is the dish I have the fishes in. I tickle the fish, they flick their tails, and I see this all in 120fps. Now, I measure how much of an angle the fish made their tails flick, all in 3-D, because that's what the scope records in. I have a half TB per fish to comb through, and I have ~20 fishes, say ~10TB. At the end, I get a single graph comparing the fish with some gene to those without it, and I have 10TB of 'data' left over. Yeah, someone could comb through it all and find something else to look at. But i forgot to record the precise temperatures, the orientation of the fish, the fish that I knew later died, etc. I had that all in my head. And, hey, what do you know?, the p-value is ~.45 and therefore there is no 'real' difference in the fish and we can't include this in a paper. Now all that 'data' is being kept on a drive on some computer somewhere and is counted towards the budget that the lab has on the shared spaces. It's not really 'data' anymore, in that it is useful to advancing knowledge for anyone (it counts as practice I guess), but it still clogs up space.
However, you still have to keep all the data around for possible re-analysis depending on the terms of the paper you submitted to, for up to 10 years or more in some cases, along with all the equipment and maybe some frozen fish samples in LN2 down in the basement. Now, not all labs are so well funded, and sometimes accidents happen and grad students are not always so informed about these rules, but you should be keeping it all ready for re-testing for some number of years. You simply cannot get rid of it all for the sake of science.
If you really want to help out bio-peepz, then helping them program is a great way to do so. Spaghetti does not even come close to describing their 'code'. If anyone really wanted to reproduce the experiment, then they would have to wade into the fetid swamp of 'code' that produced it. Most bio people cannot even begin to tell you how to code or what it even means. They can PCR and Western Blot better than Jesus himself, but code? No way. It is a real hindrance to the sciences, actually. The helping hand that code was meant to be has become a ball-n-chain that drags peoples minds away from them and 'lets the computer do it' for them.
I'd guess the model still wouldn't be considered "pure" enough by the scientific community unless it can be proven to be free from bias samples of raw information.
That's actually not correct. What if your validation method is broken? What if that unit test you wrote has a bug in the test itself?
In science you have to keep the raw data available.
This is not only for reproduction of your particular structured metric (i.e. tail flick angle), but for other groups to design novel metrics that may add to or trivialize the published result.
There is nothing I'm aware of saying that you can't losslessly convert tiff to png or even tiff to tiff.gz. Nobody needs to collect pcap files for their sensors (the raw data). They just collect the data and stick it in an appropriate file.
So, trying to tell them that you can losslessly convert them, though maybe true, is not a good idea. .tiff can handle a stack of images and you only have to call it once in a MatLab script, .jpg cannot and will result in a shitfit as they try to load in an entire folder of images. Then you get into bit-depth and God help you if the conversation ever has you trying to pronounce 'int', 'float', 'str', or 'double'. The class of a variable? Dude, this person could care less and will get out a protractor and measure tail flicks off a print-out before any of that will ever sink in.
No need to get into bit depth and int float str and double (though all of the biomedial researchers I've worked with knew these very well). Just: 'you can convert to png and back again without losing data. Just like zipping a file and unzipping it again. Let me show you how...'
So, inference to many biomedical folks is just 1-dimensional. Big data has a long way to go to penetrate fields where people cannot think in more than 1 dimension!
Unfortunately, there just doesn't seem to be much tenure-juice from innovating in statistical methods for most life-science fields. Not all, of course. Science moves slowly.
I was at an anesthesiology conference. Dr. Emery Brown is at Harvard and Mass. Gen. and has been triply elected to the Nat'l Academy in Eng., Medicine, and Biology (only 19 other people hold that record). Suffice to say, the guy is Smart with a capital S. He was talking about his new auto-anesthesia machine that records EEG on the head and them modulates the dosage of the anesthesia drugs to maintain or change the depth of anesthesia. It keeps you 'knocked out' better than any human can and will bring you 'back up' probably a lot better too (more testing is needed). Very basically, when knocked out for surgery, your entire brain rhythmically fires at ~11Hz. As you wake up, that rhythm deceases and goes away, the deeper you are, then your brain increases that rhythm. You measure that with the EEGs and filter out all the rest. So, to keep you knocked out, you increase the dosage when you see the rhythm slowing and decrease when you see the rhythm speeding up.
Anyone that has taken differential equations and the barest EE knows that the way to control that is with an Op-Amp circuit (https://upload.wikimedia.org/wikipedia/commons/f/fb/Voltage_...). It is incredibly straightforward if you were even barely awake in those classes. You don't need something as complicated as Machine Learning to do it, you just use feedback.
Now, when the Q&A session started with Dr. Brown, it was a mad-house. The anesthesiologists and neuroscientists were just dumbfounded that a simple circuit like that could control the machine with total clarity and reliability. No explaining by a guy with that level of gravitas and a credential list longer than your arm could convince them that it would work. The phrase they kept coming up with was 'I'm not a math person'.
Ok, got it? These folks are brilliant in bio and neuro and surgery. Just flabbergastingly good. They can feel how much drug is needed for a baby that just got their arm ripped off in a car accident. But they will never understand the math and they do not trust it at all as a consequence.
So trying to say that Machine Learning is the pancea to this flood of data is like saying to Eskimos that to heat their houses they need to simply invent a nuclear reactor. It's never going to happen.
This is extremely condescending.
I know many brilliant biologists who are also great mathematicians, statisticians, and programmers.
Anesthesiology isn't really a bio field at all. It's a medical field. A field that doesn't do a lot of cutting edge research. What you describe is engineering, and is likely headed by competent engineers, and funded by said Anesthesiologist.
>But they will never understand the math and they do not trust it at all as a consequence.
Again. Extremely condescending.
The people who researched, designed, and built the machine you describe are likely extremely competent in math. Yet, you are describing the end-user, and from that inferring the designer of the machine to be of the same expertise as the user.
Check your assumptions.
After a few years I understood stuff better and moved my pipelines to MapReduce. And I built a bigger simulator (Exacycle). It was easy to process 100+T datasets in an hour or so. It wasn't a lot of work, really. We converted external data forms to protobufs and stored them in various container files. Then we ported the code that computed various parameters from the simulation to MapReduce.
I took this knowledge, looked at the market, and heard "storing genomic data is hard". After some research, I found that storing genomic data isn't hard at all. People spend all their time complaining about storage and performance, but when you look, they're using tin can telephones and wind up toy cars. This is because most scientists specialize in their science, not in data processing. So, based on this I built a product called 'Google Cloud Genomics' which stores petabytes of data (some public, some private for customers). Our customers love it- they do all their processing in Google Cloud, with fast access to petabytes of data. We've turned something that required them to hire expensive sysadmins and data scientists into something their regular scientists can just use (for example, from BigQuery or Python).
One of the things that really irked me about genomic data is that for several years people were predicting exponential growth of sequencing and similar rates of storage needs. They made ludicrous projections and complained that not enough hard drives were made to store their forthcoming volumes. oh, and the storage cost too much, too. Well, the reality is that genomic data doesn't ahve enough value to archive it for long times (sorry, folks, for those that believe it: your BAM files don't have value enough for you to pay the incredibly low rates storage providers charge! Also, we can just order more harddrives, Seagate just produces drive to meet demand, so if there is a real demand signal and money behind it, the drives will be made. Actual genomic data is tiny compared to cat videos.
The real issue is that most researchers don't have the tools or incentives to properly collect, store, and use big data. Until that is fixed, the field will continue in a crisis.
I'm running a project that's 10gb in size and uploading the data to AWS S3 was absurdly slow.
Is there any way to speed up the upload that you found? 10 GB was painful though, I can't imagine uploading terabytes.
At the time, UDP based tools such as Aspera[1], Signiant[2] and FileCatalyst[3] were all the rage for punting large amount of data over the public Internet.
For UniProt a smaller dataset we just use it to clone servers and data from Switzerland to the UK and US at 1GB/s over wide area internet.
Very fast, and quite affordable.
that works well.
https://nsaunders.wordpress.com/about-2/about/
> You may be wondering about the title of this blog.
> Early in my bioinformatics career, I gave a talk to my department. It was fairly basic stuff – how to identify genes in a genome sequence, standalone BLAST, annotation, data munging with Perl and so on. Come question time, a member of the audience raised her hand and said:
> “It strikes me that what you’re doing is rather desperate. Wouldn’t you be better off doing some experiments?”
> It was one of the few times in my life when my jaw literally dropped and swung uselessly on its hinges. Ultimately though, her question did make a great blog title.
edit: To add an anecdote that I believe I read on HN; regarding the topic of the huge datamine of DNA and other health data provided by the U.S. government, a commenter said that the reason it was all on FTP was because professors couldn't download large datasets via their web browser, or some such technical hiccup.
I won't say that putting data on the Web makes it automatically more accessible, but data discovery through FTP requires a bit of scripting skill that I imagine the average biomedical scientist does not have.
We're a small startup that has partnered with UCSF Cardiology to detect abnormal heart rhythms, and other conditions, using deep learning on Apple Watch heart rate data:
https://wsj.com/articles/new-study-seeks-to-use-deep-learning-to-detect-heart-disease-1458240739
https://blog.cardiogr.am/three-challenges-for-artificial-intelligence-in-medicine-dfb9993ae750
https://a16z.com/2016/10/20/cardiogram/
We have about 10B sensor data points so far. If you're a machine learning engineer and interested in working on this type of problem, feel free to email me: brandon@cardiogr.am. * semi-supervised sequence learning (we have a paper in a NIPS workshop next week on applying sequence autoencoders to health data, for example)
* deep generative models
* variational RNNs
From a day-to-day perspective, we use tools like Tensorflow and Keras, similar to most AI research labs. In general, we try to act as a software startup that happens to work in healthcare, rather than as what you might think of as a traditional biotech or medical device startup.Does that help answer your question?
The logistics of dealing with big data is a solved problem in biology. We don't need data fast, and a laptop can run most analyses now, so it's just storage really.
The hard part is having familiarity with all the esoteric metrics used to compare genomes and all the unfamiliar ways your results can be biased.
An off-the-shelf corporate data scientist, while a critical thinker, lacks the domain specific knowledge to be able to ask good questions.
Yeah, I think that's the other side of the coin here (for biomedicine and most other fields). Even if better engineering is what is most greatly needed, it's not just about just being purely good at code and logistics. As a corollary, if you do have the domain knowledge and are in a position to act, you don't have to be an ace programmer to make a huge difference.
FTP is just bytes over TCP, it's just as prone to problems. HTTP recovery is far simpler because of ranged GETs.
FTP is definitely not better.
I've moved petabytes over the internet using FTP and HTTP. HTTP wins hands-down.
i.e. I do accept that http is a much better protocol than FTP but that social and organisational reasons lead to FTP being more stable and dependable in the field for large file download than http servers.
that's a server config issue, nothing to do with either of the systems. If you're setting up a CDN (which is what these genomic servers are) you just configure the servers to serve files, nothing else.
My hundreds of aborted recursive FTP fetchs compared to my almost-never aborted recursive HTTP show that anything you're seeing about FTP being more stable is just a PEBKAC issue.
All data on the sequence read archive (SRA), the main worldwide database for sequencing reads from NCBI, is accessible via FTP and via ascp. FTP is slow, ascp is blazing fast.
The post-docs and graduate students who do the heavy lifting on all of these projects don't make a living wage.
They can't raise a family, buy a house, or save for the future. The people in charge made them indentured servants and now those leaders are going to reap the whirlwind.
80% of postdocs I ran into came from either China or Russia. Almost all of them were "disposable" from our point of view. Long hours, low pay, little reward. Despite reaching out to them to build friendships, it was incredibly rare for them to do anything but socialize in their small postdoc circles consisting of people from the same province in China or whatever.
The best part is, from their point of view, they were here to take advantage of our "advanced" research infrastructure until they had enough experience to duplicate it back home. Fair trade, I think.
Huge shame, because I love biology. I'd love to work on these projects, but I'm not going to get a PhD just so I can make what I'm making now.
$500k for a new microscope, no problem! It's a fancy, four channel live microscope. So it takes four 1Mb images each frame, and you're running a 10min long experiment taking an image 5 times per second. So that's like 12Gb per imageset. And you take like 10/15 replicates per experiment under 3-5 different conditions. That data for that experiment which under-girds a 6-7 figure grant is now stored on a $100 3TB USB disk from Best Buy. Oh, and trying to process that 12Gb image over USB2.0 using MatLab on a student's personal macbook air is horribly inefficient - but there is no other option for the student really.
The students collecting the data and storing it on their local HD or laptop hard drive have no place to archive their data even if they wanted to. There are no repositories capable of generically storing that kind of huge data that needs to be frequently accessed at the price the students/labs are willing to pay (nothing).
And this speaks nothing about the code every student reinvents in MatLab to do basic scientific analysis. Or worse, does not reinvent and instead reuses 20-year old code written by some long-forgotten student who wanted to try their hand at a 'new' programming language like IDL.
The 'students' and postdocs are paid nearly minimum wage to do the high-tech biomedical research. There are no computer-scientists to be seen because they would be fools to give up making 5-8x money across the street at Twitter.
On the other hand, the scientists know their work really well, and it will take a truly integrated team to solve these issues. A computer scientist can't just come over for a day and write up an app to help out. The code will have to make scientific assumptions and must be custom for many/most projects. But it's very hard to build a capable team when the market salary for certain kinds of team members is multiples of another on the same team.
The HIPAA and other regulations aren't any more annoying than any other modern programming practices these days. As for liability in the case of a breach, that's what E&O insurance is for. It does raise the barrier of entry a little, but it's not by any means prohibitive.
I think the bigger issue is what format this data is in. Most medical and medical records data is in a variety of proprietary, non-open and difficult to integrate technologies. One look at HL7 is usually enough to send a programmer back into the loving embrace of anything else.
Working with this data is time intensive and expensive and I'm guessing most heathcare companies don't see it as worth the cost.
It's not worth the cost 99% of the time. The only reason we're still making real breakthroughs is because of the research institutes doing basic research on public funding. Even then, they're swinging for the low hanging fruit.
> I couldn't have gotten out of there fast enough for the way saner land of tech companies.
You know you have a problem when tech companies seem saner in comparison.
In many cases scientists are required to submit data to journals when they publish. However, data is a funny term! Using genomics data, there are many levels of data that you'll encounter - everything from raw image files (many gene sequencers are actually automated digital cameras, taking pictures of florescent markers attached to the DNA strands) to intermediate sequence files to formatted pieces of selected data. Then, there's the whole toolchain used to go from sample to formatted final analysis? What's required to be submitted? How long should it be kept? Who checks all this to make sure it's not NES roms or Shakespeare instead of the right data?
Finally, there's the big questio: how can we be sure that the data captured in the intermediate or final steps of analysis actually originates from the raw data? Should scientists store the raw data (TBs and TB), intermediates (GBs and GBs), or final analysis (MBs). For how long?
To answer your question about meaningful long term research - I've personally seen grad students' careers effectively ruined due to shitty data storage hygiene.
As someone in the field I beg to differ :). And regardless of who pays in the event of a breach, the effects may be sufficient enough to shut down the company for future projects.
You can basically self-certify, but most serious companies will bring in an outside contractor on an ongoing basis to certify compliance. Staff needs to be trained, computers need to be managed, software changes have to be very thoroughly reviewed, updates become slow. It makes it pretty unattractive to enter into for a lot of devs.
- Real, bespoke biomedical analysis is not trivial in effort, cost, or time. There are biomedical analysis systems-in-a-box (look at https://galaxyproject.org), but that's just canned analysis. To make real breakthroughs, you need rigorous analysis that requires years of experience to be able to perform.
- It's easier to get the money to collect the data than it is to effectively steward the data you collect. In a past life, I ran a biomedical research computing facility, and everyone got plenty of money for new sequencers, mass specs, and other fancy instruments. They got plenty of money for collecting all kinds of data. No one would ever add money to their grants to actually STORE the data. They would literally put the data on USB hard drives bought from Best Buy, and left them in file cabinets and on desks. There was absolutely nothing I could do about this, and so I quit.
- Research is balkanized to hell. Even though I ran the scientific computing for 20 research labs, each research lab was its own fiefdom. They could decide to obey or disobey my policies at will, since they controlled their own funding. You can imagine what happened when I proposed turning on quotas (~100TB per lab, to start!). Rather than work with my team to determine how to share resources, people would just jump off my high speed facility, buy a shitty cheap JBOD from Dell for their analysis, and store their archives on shitty cheap USB hard drives from Best Buy. The funniest part was that if the hard drive failed, and the data couldn't be restored, in theory the primary investigators could get into real legal trouble. No one seemed to worry.
There are a few biomedical research institutes that "get" scientific data stewardship - Broad, Scripps, but for the most part, biomedical research computing is a total clusterfuck and I couldn't have gotten out of there fast enough for the way saner land of tech companies.
The scariest part was that before I left, I built a highly scalable, long-term archive for scientific data built on LTO tapes that would allow ridiculously cheap (basically the cost of LTO tapes) on-line and near-line storage. When I left, no one wanted to bother with paying for the upkeep of the hardware, people got bored with swapping tapes, and it eventually died. Oh well. Your tax dollars at work.
This is such an interesting (and kind of worrying) problem. It's one of those problems where the obvious solution is only solving the obvious problem. As you have pointed out there is a more underlying problem that hides underneath and which you only know of if you know the well enough.
I need to think about this some more to understand why that is and why an incentive structure can't be created.
You gave me a new perspective on things today, thanks for that.
To solve this problem, I tried giving "monopoly money" to the professors, allowing them to trade data storage and cluster time for favors, analysis, and so on. For the researchers who didn't need as much storage, they could give their excess up. For those who gobbled up storage, they could "buy" the excess. It ended up failing because I didn't have backup from leadership to say "no" when I was asked to do things that were irresponsible:
Yes, you can buy a 2TB HDD for $100. No, that isn't the same as 2TB of storage on an enterprise-level storage array, clustered, with local mirroring, offsite tape backup, etc. No, I won't plug your 2TB USB HDD into my compute cluster.
check the job here, most of them are postdoc level. https://www.biostars.org/t/Jobs/
the postdoc level means you get about 50k~80k, even in bay area.
The situation is really bad.
1. Each organization has their own data silo and treats it as if it were gold. Therefore it is difficult and expensive to aggregate.
2. Because the data is sensitive, there are often many contractual restrictions beyond HIPAA on how the data can be used. Anything from research purposes only to you can build a commercial product but you can only sell it to the data owner.
3. Even if an organization wanted to share data with reasonable restrictions and pricing, it is often hard to do so because their focus isn't software. So it is difficult or impossible for them to share it.
Source: I made a serious attempt at a healthcare data analytics startup. It didn't work out.
In the modern world we promote a "one size fits all" health model that is encouraged by the many third parties involved in the doctor patient relationship (third parties such as insurance companies, employers and government regulators). It's going to be very difficult to adopt this model to the patient-tailored healthcare that is going to be required to fully leverage the recent advances in machine learning.
Except for when you want to cherry pick through it to help get your research paper published.
Try to make sure you aren't just helping people publish papers for the sake of it. As soon as you sense that, stop working with them. It is a very, very bad thing.
In Earth Sciences, Astronomy and presumably quite a few other fields there is a staggering amount of data coming in and the number of people who have the required domain knowledge, math knowledge and programming knowledge is not growing as fast as the data is. Teams can and do help, but well, there is lots of work out there.
The data processing pipelines need to be very efficient and questions clear to tackle these. Not to mention the processing algorithms for this data.
That said, this is a good reminder that data scientists are in high demand and can make a difference.
http://useast.ensembl.org/info/about/species.html
for example Humans: You end up here ftp://ftp.ensembl.org/pub/release-86/fasta/homo_sapiens/dna/
there is a read me which helps figure out what this stuff is, or ask your friendly neighborhood biologist.
ftp://ftp.ensembl.org/pub/release-86/fasta/homo_sapiens/dna/README
There are petabytes of open data there for you pick from every kind of organism. Go nuts.
For a more comprehensive list check out one of the many "Awesome Public Datasets" [5] (biology section).
[1] https://my.pgp-hms.org/public_genetic_data
[2] https://www.openhumans.org/members/
[3] https://www.openhumans.org/public-data-api/
[5] https://github.com/caesar0301/awesome-public-datasets#biolog...
What we need is to have people who can ask questions and think critically. Good, hypothesis driven science is how we discover entirely new concepts and mechanisms.
Maybe it only makes sense to store sections which deviate in a significant way from a range of error (lossy compression).
Maybe some of those inputs just don't make sense for the questions being asked.
A concrete reasoning for why the data should be kept needs to be presented, and THAT is what should call for the funding to back that need.