"the genome is a sequence that can be language modeled/auto-regressed for depth of understanding by the network"
The genome is not a sequence so much as a discrete set of genes which are themselves sequences which specify construction plans for proteins. That distinction is important.
Language modeling in the context of machine learning typically means NLP methods. Genetics is nothing like natural language.
Auto-regression is using (typically time series) information to predict the next codon. This makes very little sense in the context of genetics since, again, the genetic code is not an information carrying medium in the same sense as human language. Being able to predict the next codon tells you zilch in terms of useable information.
"Depth of understanding by the network" ... what does that even mean???
The above sentence is a bunch of popular technical jargon from an unrelated field thrown together in a nonsensical way. AKA word salad.
aka a sequence. "a book is not a sequence so much as a discrete set of chapters which are themselves sequences of paragraphs which are themselves sequences of sentences" -> still a sequence
these techniques are already being used, such as in the paper I just linked.
> Being able to predict the next codon tells you zilch in terms of useable information.
You have absolutely no way of knowing that apriori. And autogressive tasks can be more sophisticated than just next codon.
> bunch of popular technical jargon from an unrelated field thrown together in a nonsensical way
Okay, feel free to think that.
There's always this assumption of it "will never work on my field." I've done work on NLP and on proteins and read others' work on genetics. I think you will end up being surprised, although it might take a few years.
Let's say you have discrete sequences that are a product of a particular distribution.
Unsupervised methods are able, by just reading these sequences, to construct a compact representation of that distribution. The model has managed to untangle the sequences into a compact representation (weights in a neural network) that allows you to use it for other, higher level supervised tasks.
For example, the transformer model in NLP allowed us to not have to do part-of-speech tagging, dependency parsing, named entity recognition or entity relationship extraction for a successful language-pair translation system. The compact transformer model managed to remap the sequences into a representation that allows direct translation (people have inspected these models and figured out the internal workings of it and realized it does have latent information about a parse tree of a sentence or part-of-speech of a word).
Another interesting note is that designers of the transformer architecture did not incorporate any prior linguistic knowledge when they were designing it (meaning that the model is not designed to model language but just a discrete sequence).