Reads mapping is a massively embarassingly parallel computation, again not something you would need or want a supercomputer for. You mainly need disk IO to/from the source reads and the mapping table you produce.
Reads mapping is a massively embarassingly parallel computation, again not something you would need or want a supercomputer for. You mainly need disk IO to/from the source reads and the mapping table you produce.
For these guys, it's likely their best and only option. They probably weren't given the money to build a cluster optimized for their needs or to maintain a cloud instance. Why? Its oakridge, their main priority is HPC physics. It's hard to argue when you have access to such a HPC center. And HPC sites need all the customers they can get, lest their clusters get shoved into the cloud. It's a real fear. They'd end up with hidden costs, data lockin, and poor interconnects. To help pay for those peak massive simulations, traditional HPC need to fill up that last 10-15% and bioinformatics needs most of what they offer. Perhaps all that hard won knowledge will rub off on the burgeoning field too. :)
A vision of bio-oriented HPC is IU's clusters. Though they have a shiny new cray shasta with ampere and slingshot, several of their other clusters are 10 gigs with high mem nodes. All connected to the same storage too. The hospital is the main customer and dictates their designs.
Ditto on the paper though. It's what I disliked about Bioinformatics. All the glory to the researchers designing the experiments and they can't even bother to mention what software they used.
> RNA-Seq analysis was performed using the latest version of the human transcriptome (GRCh38_latest_rna.fna, 160,062 transcripts to which we appended the SARS-CoV-2 reference genome, MN908947). Mapping parameters were set with a mismatch cost of two, insertion and deletion cost of three, and both length and similarity fraction were set to 0.985. TPMs were generated for all 160,063 transcripts for the nine COVID-19 samples and the 40 controls (Supplementary file 2). The resulting transcript mappings for genes of interest were manually inspected to account for any expression artifacts, such as reads mapping solely to repetitive elements such as the Alu transposable element or all reads mapping to a UTR or pseudogene therein. Transcripts whose counts came solely from (or were dominated by) reads at repetitive elements were removed from the analysis. For the controls cases we ran an outlier analysis using the prcomp function in the R package factoextra. Input data were TPM for transcripts that averaged greater than one across all samples (30,102, Supplementary file 2).