Nanopore MinION – $1k solid-state DNA sequencers
nanoporetech.com
nanoporetech.com
The error rate is stupidly high (somewhere between 10 and 20%) compared to Illumina or Ion Torrent who give error rates far less than 1%.
It can give very long reads, which are useful in some niche applications. But it’s been massively over-hyped (and over capitalized).
The neat thing is that it’s very small. But that isn’t really compelling given the very low accuracy.
This is not a "full" human genome, but a collection of 150bp fragments that can be realigned to an existing human genome. You cannot take this and infer the whole diploid genome of the individual. There is a huge amount that will be missed, and all of our current knowledge is based on this gappy picture of what's going on in single genomes and human populations.
> It can give very long reads, which are useful in some niche applications. But it’s been massively over-hyped (and over capitalized).
I think you're dismissing the technology out of hand because of biases derived from much more limited short-read technology that only allows us to reliably see small variants <50bp.
Without these long reads we can't see structural variation (SVs). There is an increasing amount of evidence that much of adaptive variation is driven by these kinds of variants. If you want recent evidence, see https://www.nature.com/articles/s41588-017-0010-y. There has long been evidence that there are huge copy number variations in humans, but these are still not evaluated reliably: http://science.sciencemag.org/content/330/6004/641.
We should be open to the possibility that our observational techniques are limiting our understanding how how genomes work. This has consistently occurred in the history of every observationally-driven science.
It's amusing to me that people assume that SVs are "niche" when even the limited surveys of genomes we've been able to do with short reads show that roughly an equal number of base pairs in the human population vary due to small variants like SNPs and indels and big ones like deletions, insertions, and large scale copy number variation: http://science.sciencemag.org/content/330/6004/641
Most structural variation I've seen based on whole genome assemblies is not even classifiable into neat categories like "deletion" or "insertion". If you think that "most" things are detected with short reads then you are deluded by the dominant technology.
There is, likely value in long reads, but what non-niche research applications are there for highly error’d reads that justify a valuation of several billion dollars?
The area where DNA sequencing will first be revolutionizing clinical practice is in sequencing pathogens for sake of identification. In these instances nanopore sequencing rules, because it can give answers in minutes.
The most compelling near term applications (NITP etc) use fragmented DNA, and long reads will have no benefit here.
So, yes. Long reads are useful, but you need to have at least reasonable performance in other respects. The same thing has been seen with PacBio, who have not played well in the market, despite having a read length advantage.
With sepsis, every hour counts.
For things approaching a read length the per-base error rate of a single read is simply irrelevant. In practice, with sufficient coverage (e.g. 20x) you simply don't care about the per base error rate of the reads.
The Insertion/deletion error rate is 20-30%.
The point mutation error rate is something 0.1-1% (higher than HiSeq but not crazy high).
This means with a semi-decent reference genome you should be able to do re-sequencing fairly accurately. It also means, that in conjunction with HiSeq reads you can do cheap genome assembly, using the HiSeq reads for coverage, and the minion reads for scaffolding.
The mismatch rate is much lower. But it's hard to calculate exactly the mismatch rate when the indel rate is so high.
I was able to get about 1-10% mutation rate, with a median of about 1.5%. Rate depending on quality of the run. In general it was on par with PacBio.
Are you using it for fun? professionally? academically?
Reagent kits run a few hundred dollars depending on what your doing and you get several uses out of them.
IIRC it's not certified for use as a medical diagnostics device so it's "research only".
It's a nice system with great (none) upfront costs that has a lot of potential. But read quality and quanity per dollar aren't as good as what we get with our NextSeq (which by contrast costs something crazy like 300-400k)
The problem here is that everything in that pipeline keeps improving at a fast base. Also the reference databases are updated all the time. Whereas sequencing will probably always give the right answer to the precise question it was given, I don't see any luck in certifying these methods any time soon.
That's what Illumina will tell you. But that really depends what you mean by "a genome" and what you intend to use it for. Getting anywhere near the quality of the Human Genome Project is not possible.
That is an unrealistic and inappropriate metric to use.
The HiSeq X isn't mean for making new human reference genomes. It's meant for re-sequencing at scale, and speaking from personal experience, it does an incredible job in quality, speed, and cost.
Even if we could reach reference assembly quality with a hiseq run, we wouldn't want to use it that way. Having reference genomes and annotations (hg19 and hg38) means we can compare and contrast individuals from a common foundation. It would increase the cost and time of genome analysis 100 fold if we had to do denovo assemble for each individual sequenced.
As for pacbio, if I remeber correctly, their cost was low (high equipment cost though) and sequencing was relatively fast, but at a reduction of accuracy. Never dealt with them first hand though.
https://www.youtube.com/watch?v=XaumUp4GpCw
https://www.youtube.com/watch?v=3K9whMm7vvc
It's almost like Moore's law.
Also, the initial Human genome wasn't that good either...
https://www.researchgate.net/profile/Orit_Shaer/publication/...
(You don't need a complete sequence for that, just enough unique markers to ID somebody.)
Goodbye "papers, please", hello "cheek swab, please".
This idea already here, but with technology based on older/simpler methods. One example is Intigenx: https://integenx.com/rapidhit-system/
The FBI has plans for implementing them: https://www.fbi.gov/news/testimony/fbis-plans-for-the-use-of.... The laws just need to catch up (or prevent it).
And with any other technology, including a miniaturized PCR implementation, it would be too slow.
[1] https://www.reddit.com/r/dataisbeautiful/comments/7rpdr0/wev...
The read lengths might be long, but the error rate is a couple of orders of magnitude higher. I can’t see a market large enough for highly error’d log reads to support their valuation.
https://www.jcapitalresearch.com/uploads/2/0/0/3/20032477/20...
To me the fact that they have a patent monopoly on "putting DNA through a protein pore with voltage sensing" is tragic. Who knows were we would be today if these patents had been granted to the public domain.
https://www.investopedia.com/terms/b/breakevenanalysis.asp
Are they anywhere near break even or are they going to require another large infusion of venture capital and can Nanopore find these new investors?
Their valuation is not about short term sales. Their technology really is somewhat of a holy grail. There is no "simpler" way to read DNA in terms of technology. Everything else involves complicated biochemistry. Illumina beats Nanopore Technologies right now because of the scale of their machines. But they can't scale them down much further without losing performance. The most efficient Illumina machines will always be cupboards or bigger.
No competitor can come up with something better than ONT. Nothing smaller, for certain. The only way to beat Oxford Nanopore in the future will be to implement a very similar nanopore sequencing scheme.
Illumina doesn’t beat nanopore because of the “scale” it beats it because the technology produces fundamentally higher accuracy data. The error rate on Nanopore reads is >10%. Even on first generation Illumina machines the error rate was 1%, and is now significantly lower.
Oxford Nanopore have been working on this for 15 years, this is the best they can do. It might be possible to create a Nanopore (maybe solid state, not protein) system with a lower error rate, but I think they have little hope of doing it.
I think they’d need to sell >10,000 units a month. That is if they are making any kind of profit. Unfortunately I don’t think there’s a market for anything like that number of flow cells.
Illumina (see recent IP battle with Oxford Nanopore) also have key IP in this space (for mspA).
If you're interested in Oxford Nanopore, you might also keep an eye out for Roswell Biotechnologies (https://www.genomeweb.com/sequencing/roswell-biotechnologies...). TL;DR: their sequencer involves immobilizing polymerases in circuits so you can measure the current changes that occur as the polymerase adds bases to a strand. This might be a good approach, since it doesn't involve optics (like IonTorrent and minION, keeping costs down) and you get a current event per base (as opposed to per 5-6 bases as with minION), making basecalling easier and potentially more accurate. And it will be fast, since you're reading as fast as the polymerase can work. They want to get up to 10kb reads, and I imagine they could increase the consensus accuracy per read by looping around a single molecule of DNA several times, like with PacBio. Seems feasible too (after all, PacBio has successfully been able to integrate polymerases into very small features rather well, so I see no showstoppers there).
Barring that, the Cas9-mediated targeting method looks somewhat promising (https://www.youtube.com/watch?v=DGDH-FdoARM). No need for amplification! Though for applications with human DNA you might run into some issues with the copy number of a sequence of interest if you’re starting out with a normal amount of genomic DNA.
https://www.nasa.gov/mission_pages/station/research/news/dna...
Since 2013, Illumina's brought the price of a genome to about $1k with the HiSeq X, which essentially scales up their existing tech. No idea what the next few years hold, but I'm excited to see insurance companies beginning to cover genome sequencing as an affordable diagnostic tool.
Interestingly the cost had not dropped in the past two years.
Also the graph references genome.gov but that site appears to be unavailable at the moment.
Source: Software Engineer for a well-known non-profit sequencing lab/center.
A lot of the costs are outside of the sequencing itself. You have to extract the DNA from the sample and prepare that DNA for the sequencing platform you're using. These costs are both the reagents needed to do this, and also the associated labour costs, even when done at scale.
Perhaps you meant $100 human genome though? However I think roughly the same principles apply.
Is this a similar product/result as buying whole genome sequencing from an established company for the $500 to $800 they charge these days?
We use ONT's minION for strange niche stuff, where either the DNA prep method matters or the latency (the time from loading to the first data out of the machine) matters. Read quality is typically not as good as illumina, so for applications where we need quality we stick with the NextSeq. There's no reason to think nanopore technology won't get better though, but for now, if you've got the money illumina is still the way to go most of the time.
That is, without handing the DNA over to somebody else.
But you also need to consider how you prepare your DNA to go into the Nanopore. It's a lot more investment than you might think. Extracting your DNA in a clean enough way for it to work with the Nanopore will require equipment/facilities that most people will not have access to (phenol/chloroform is the optimal way to extract DNA, which will require a fume hood and toxic chemical disposal). Depending on how you want to prep your DNA, you might need more specialized equipment/reagents (Ampure beads, ligation kits, end-prep/TA-tailing, etc).
> You're buying a service for the 500-800$. Here you're buying a flowcell for 1000$ (the "sequencer" is more or less free). The technology is pretty different and a lot more bleeding edge. If what you want is just a genome aligned to a known reference and you can wait a few days then this is not what you want. We use ONT's minION for strange niche stuff, where either the DNA prep method matters or the latency (the time from loading to the first data out of the machine) matters. Read quality is typically not as good as illumina, so for applications where we need quality we stick with the NextSeq. There's no reason to think nanopore technology won't get better though, but for now, if you've got the money illumina is still the way to go most of the time. >
Here's some more detail on the tech works: https://www.mongodb.com/blog/post/oxford-nanopore-technologi...
It's remarkable the perverse economic rules that make disposables with vendor lockin a feature. Opex vs capex.
(note: I work at MSR, but on other projects)
Pshhh, easy: Gattaca
kidding
Many thanks
I've personally used the oxford nanopore to diagnose malaria subtypes from ~1ml blood draws. Though it took awhile to do so, well beyond clinically relevant time periods. Our best turn around was about 2 days for sequence, but analysis takes much longer.
The enormous reads offer a different window into genomes, by spanning repetitive regions where short reads can't be accurately placed. For lots of applications, the sweet spot is to use long reads for "scaffolding", then build up coverage with short reads.
Did you get it back as a file?
https://store.nanoporetech.com/catalog/product/view/id/122/s...
And in my experience, the washing buffer doesn't do a great job. I read somewhere there is ~10% carryover.
Apparently the cost is competitive for what it does.