When your first run fails, are you willing to pay again as much for the kits to run it again (these machines are very much following the cheap printer/expensive ink model IMO)?
Once you have some data, do you know how you will BLAST it against annotated genomes to figure out if you have mutation X? How do you interpret e-scores, etc?
For the OP company, when they say "download the data" do they mean the raw reads, or assemblies? Make sure this is spelled out (it likely is, I haven't looked lately). Do the downloads/service provide adequate metadata on the data generation process so that you can tease out errors in reads from reality (real single-nucleotide mutations), etc.?
All fun stuff, but only for a very serious hobbyist, so to speak.
On the products page it says they provide FASTQs, aligned BAM and VCF. Which is exactly what one would expect. They're almost certainly just running the DRAGEN pipeline (or something similar) and giving you the output it creates.
You aren't trying to sequence the human genome from nothing, but instead have a reference genome to work with. You take each of the reads you get from the sequencer and find where it best matches up with the reference genome.
I think the wet portion is a far bigger challenge for the typical computer programmer hobbyist.
(Also I don't think $1,000 will get you, in addition to the sequencer, a flow cell sufficient to gather enough reads to tell the difference between sequencing errors and mutations?)
Alignment. You could easily automate alignment of raw sequencing reads to their position in the reference genome. From there, even variant calling (finding your individual changes from the reference) is pretty easy as well.
Assembly is a bigger challenge that takes the raw reads and puts them together into a larger (graph) structure. This is usually de-novo, but could be able to be done with a reference as well. But this is not really something that is routine or possible with 30X sequence coverage to the degree one would expect.
The wet-lab portion isn't terrible and anyone with skills making anything (soldering, for example), should be able to do it -- once you figured out some of the jargon.
The bigger challenge someone not trained in this would have is interpreting the variations that you'd find. This can be quite difficult to figure out if an A>C at a particular position is meaningful or not. Again, much of these can be filtered using existing tools, but even then, it can be quite difficult to interpret results.
Whoops; fixed!
> This can be quite difficult to figure out if an A>C at a particular position is meaningful or not.
What about https://promethease.com ?
Note: I do this work for cancer genomes (WGS, tumor and germline). Promethease is not something that I use for annotations, but my workflows are significantly more complex.
The biggest thing I’d look out for is mindset. Biology is very different from comp sci. CS is very deterministic. You change something here, you get a result there. Biology, on the other hand, is very messy. There are compensation mechanisms on top of compensation mechanisms, and everything favors life. If you see a deleterious variant, that still doesn’t mean you’ll see a phenotype (effect on you). It’s quite a different world and can be very scary if you are looking at your own data.
And to complicate things further, one variant in a gene could do nothing if the other copy is still good. Even if you have two bad variants in a single gene, they might be on the same allele (from the same parent). In which case, you still have one good copy! It’s like how Iceman got hit twice in the same engine at the end of the original Top Gun. He still had one good engine to get home. For many genes, that’s all you need. For others, you might not even need one good copy as other genes could take over the same job.
Yes, there are some variants that you don’t want to see. Things that might indicate high likelihood of disease down the road. But even that is bound by statistical probabilities. Good news — some of those can’t be seen by this type of sequencing.
You’re absolutely right though — it can be scary. After looking at many genomes, it can be amazing that we all work as well as we do.
Additionally, we're talking about nanopore data, which is much longer reads, and so is going to be computationally a bit easier.
Still, ONT users would need to understand how to convert raw reads into useable basecalls at loci of clinical interest
Last I looked into it, it was a minimal set of extra equipment -- votexer, centrifuge, etc... nothing unusual for a bio lab. And even most of that could be handled with a good snap of the wrist, if you weren't trying to be super precise.