The Lurker: How a Virus Hid in Our Genome for Six Million Years (2013)
phenomena.nationalgeographic.com
phenomena.nationalgeographic.com
For a hypothetical we could imagine the scenario where viruses existing in say a remote rain-forest, generally innocuous and self-contained. When deforestation occurs the virus is now open and can potentially infect whatever mechanism (human) that has disrupted it, binding to the same virus DNA that has existed for millions of years.
It's not that far-fetched to think unintelligent systems have defenses that we are largely unaware of - or cannot understand.
There's an entire endocannabinoid system, and it's not because pot is natural. Nicotine broadly matches up with acetylcholine and nicotonic acetylcholine receptors.
It's far more likely that plants evolved and selected for compounds which have beneficial/detrimental effects on possible consumers, like nicotine as a pesticide, than it is that we have receptors for plants found in the wild on the off chance that we may have consumed them or that they would have been so prevalent in the past that receptors for them would have been selected for.
Any more modern books of note?
we would need a lot of experimentation to figure out what's safe to remove.
you can think about it as a physical "dependency hell"
Besides that... we need to understand how a cell recognizes its purpose in its local environment - some kind of local communication probably. If we knew what the various environments that control cell responses are, we would have the basis for something like unit tests... and then we could try randomly removing parts of the DNA and see if its failing to perform as expected or not.
The animations I've seen that purport to show the general public how DNA replication works indicate that it is a sequential process. The DNA is split into two strands, kind of like a zipper unzipping, and then bases are added to the two strands to form the two new complete DNA molecules. One strand (the leading strand) has the new bases added one after the other in the same direction. The other strand (the lagging strand) has them added in short called Okazaki fragments.
Errors on the leading strand should be independent and occur at a constant rate, and so having useless zones should have no effect on the number of errors a given useful zone gets.
The lagging strand is more complicated, because you have at least 3 distinct things going on: finding where to start on Okazaki fragment, filling it in, and recognizing the end. I suppose that allows for a different kind of error on the lagging strand (messing up recognizing the start or end of an Okazaki fragment) that would affect multiple consecutive base pairs. Useless segments would increase the average spacing between useful segments, and so would decrease the chance that a given multi-base error in a useful fragment affects multiple segments.
However, the problem here is that the enzymes aren't perfect and coding errors are quite common (they get fixed, sometimes). A big problem is that some chemicals can look like a nucleotide (A or T or C or G) and after insertion it can "decay" into a different one, hence causing an error. During replication it is possible for a DNA fragment to be cleaved (once again, it is often repaired, but thats how the Y chromosome came to be AFAIK) however sometimes enzymes "mess up" and reattach them at wrong positions. At other times base-pairs are deleted or inserted shifting the whole strand. There is a lot of things that could go wrong.
The artillery that lands on the plain may strike an advancing unit, or it may fall (possibly harmlessly) between a set of advancing units. The artillery that lands on the narrow beachhead is more likely to hit a unit.
This analogy is far from perfect: sometimes mutations are good, which is one primary driver of evolution. Non-coding regions and/or "baggage to be refactored" (paraphrased great-great-gp comment) in DNA (the regions of the plain/beach not occupied by an advancing unit) can absorb "errors". Also, there are other types of mutations (insertions, deletions, ...), aside from the single point mutations that this analogy was attempting to help convey.
The point is: it's like bunching up a lot of important things over a few points of failure. If you increase "the genetic surface area", you lower the chance of the important thing getting hit.
On evolutionary scales, viable DNA has been selected with a lot of non-coding (and sometimes useful) regions, we know that if we reduce that down, we are more likely to be susceptible to fatal mutations on coding regions (e.g. a region that codes for a vital protein).
In fact, copying DNA is more like downloading a large file over an unreliable network. There's a certain chance that each individual bit is flipped and the file becomes useless. You can reduce that chance by sending it multiple times, or introducing checksums, both of which add redundant data. But simply adding an extra TB of junk bytes to your download won't help preserve the integrity of the original file.
However, if genetic mutation count is time-dependent but totally independent of the size of the genome, then having a larger genome actually does protect you from individual mutations, and it would do so exactly using the mechanisms described previously.
Think of two genomes, one large and one small, both existing throughout time. Both will accumulate a similar quantity of mutations from mutagenic processes which are time-dependent like radiation exposure.
Basically, trying to give an example of the grandparent's point. (i.e. fewer nucleotides to be flipped -> more likely that an important one will be). I agree that it was a poorly executed analogy. The metaphor I was trying to make is that on the 'vast field' a random single point mutation is probably going to land on an individually unimportant nucleotide, and in the 'narrow beach' (the smaller strand/higher geninfo density) an individually important nucleotide is more likely to be hit. I'm still probably not articulating my point well, sorry.
But I think your analogy is better for a subtly different point; describing how DNA replication works in a system, where stands can be selected out, errors corrected, and genetic information can be preserved at a systemic level.