How is it known that there truly is junk dna?
How is it known that there truly is junk dna?
This isn’t just “we don’t know what it does, so it must be junk.” It’s more like, “We can’t find any sign that it matters, and everything we know about evolution says if it mattered, we’d see fewer random changes there.” Down the road we might uncover small roles for some of these regions, but at this point, calling them junk is just an honest read of the evidence we have.
Over long evolutionary or environmental timeframes, these sequences may take on important functions, potentially becoming critical under conditions we can't currently foresee.
Anyways there definitely are non-coding regions that just don’t do much and evolve neutrally. I’m hesitant to call them junk but only because that designation has burned biologists so many times.
To borrow an example: an onion likely doesn’t need 5x more DNA than a human, and a lungfish probably doesn’t need 30 times more than we do (and 350x more than a pufferfish). And yet, these enormous genomes exist. It’s very likely that portions of these sequences are what we’d call “junk,” i.e., DNA that doesn’t confer a meaningful functional advantage and can accumulate due to the relatively low cost of carrying it along.
If we want to avoid the term “junk,” we could say something like “areas of the genome for which we assign a very low prior probability of functional importance.” But “junk” is a concise shorthand to acknowledge that, while some non-coding sequences matter, there are also huge swaths of DNA in many eukaryotes that show no signs of being anything other than evolutionary baggage.
Bacteria have high population sizes. Selection can be quick and brutal. Low levels of “code of unknown function” in bacteria is perhaps related to replicative efficiency. Fast DNA replication is highly advantageous in nutrient-rich environments. No space (or time) for junk DNA.
But the rate of mutation is itself subject to selection. There isn't a base rate, just a setting that's different for different parts of the genome. Some parts have more copy errors than other parts. Some are hung out in the sun more often.
So you can conclude from the mutation rate of a particular stretch that it would probably be bad if it started mutating more, and that it would probably be bad if it started mutating less, but not that nothing's influencing the mutation rate.
Rates of mutation in these regions, and lack of conservation are hallmark clues which show that there isn't function in these regions. That doesn't mean totally useless, these non-functional regions provide the raw material for the creation of genes and functional elements. Its just that, right now, those regions aren't doing anything.
No biologist calls it "junk DNA". That is just a simplified layman's term for media press releases.
Furthermore, in rats and mice, large swaths of junk DNA have been experimentally removed, without any detectable effect on the phenotype.
If the junk DNA has any positive effect, it may be to protect against viruses or transposons inserting themselves into random areas of the genome. Keeping the majority of DNA "useless" may decrease the risk that these insert themselves into vital parts of the DNA.
> A gene is defined as any portion of chromosomal material that potentially lasts for enough generations to serve as a unit of natural selection. In the words of the previous chapter, a gene is a replicator with high copying-fidelity. Copying-fidelity is another way of saying longevity-in-the-form-of-copies and I shall abbreviate this simply to longevity. The definition will take some justifying.
> On any definition, a gene has to be a portion of a chromosome. The question is, how big a portion—how much of the ticker tape? Imagine any sequence of adjacent code-letters on the tape. Call the sequence a genetic unit. It might be a sequence of only ten letters within one cistron; it might be a sequence of eight cistrons; it might start and end in mid-cistron. It will overlap with other genetic units. It will include smaller units, and it will form part of larger units. No matter how long or short it is, for the purposes of the present argument, this is what we are calling a genetic unit. It is just a length of chromosome, not physically differentiated from the rest of the chromosome in any way.
> Now comes the important point. The shorter a genetic unit is, the longer—in generations—it is likely to live. In particular, the less likely it is to be split by any one crossing-over. Suppose a whole chromosome is, on average, likely to undergo one cross-over every time a sperm or egg is made by meiotic division, and this cross-over can happen anywhere along its length. If we consider a very large genetic unit, say half the length of the chromosome, there is a 50 per cent chance that the unit will be split at each meiosis. If the genetic unit we are considering is only 1 per cent of the length of the chromosome, we can assume that it has only a 1 per cent chance of being split in any one meiotic division. This means that the unit can expect to survive for a large number of generations in the individual’s descendants. A single cistron is likely to be much less than 1 per cent of the length of a chromosome. Even a group of several neighbouring cistrons can expect to live many generations before being broken up by crossing over.
[end quote]
I think a reasonable extrapolation from this is that “genes” (things that are subject to natural selection in our genome) can survive more generations if they are surrounded by genetic material that are not “genes” (ie. Their copying fidelity is not subject to any selection pressure.)
If a gene in this sense is recombined in a way that makes it no longer the same gene, it most likely isn’t going to be beneficial to the organism (and thus the gene’s longevity.) Most genetic mutations aren’t.
Highly recommend David Reich’s book for are good overview of the math of recombination in humans.
I got into Dawkins’ books initially from The God Delusion (as I suspect many laypeople), and heard about The Selfish Gene from there, so evolutionary biology is not my area.
It makes sense that TSG is considered the dark ages as it’s such an old book. I was always curious to read more about the topic from other — and hopefully more recent — biologists, since Dawkins sometimes feels like he’s more of a communicator than a practicing biologist (and one with a particularly anti-religious chip on his shoulder, not that he’s wrong.)
The idea that you can have DNA of some critter and there aren't some errors, unused bits, and so on, after what must be trillions of copies, well, I would find it statistically unlikely. Like saying you have a program with millions of lines of code and it is completely error-free.