As I mentioned in another post, the triangle inequality roughly holds, with some counter examples in proportion to what one would expect from a stochastic process.
If you’re asking specifically about instances of non-treelike gene transfer, the answer is twofold:
First, and I may be mistaken here (I’ve been out of the field for a few years) I think that any metric fulfilling the triangle inequality has a corresponding, consistent tree. So you’ll get a tree, totally. It may just not be the correct one.
Second, if you are asking for our ability to quantify divergence from that model, I’ll give a short explanation of the codon usage method mentioned before as an example:
The genetic code has 4 letters (ACGT) that translates to a different sequence consisting of 21 Amino Acids (a chain of amino acids is a protein, such as the enzymes doing catalysing all the fancy chemical reactions in your body)
Each three letters of DNA code one “letter” of the protein sequence, meaning one amino acid. Since 4^3 > 21, the code isn’t quite optimal.
There are some three-letter sequences that contain meta instructions like start/stop, and some that don’t have any meaning. But there are also instances of two or more three-letter sequences translating to the same amino acid. They are functionally identical. You can replace them in the lab, without any change in phenotype.
For some reason, some species (or even branches in the teee of life) still show preferences for using one or the other three-letter code for a given amino acid.
If you plot which of the (functionally identical) codes are used within a bacterial genome, you can find sudden changes, where the preference switches dramatically for some lenght, then returns to the previous preference.
That’s indicative of a piece of DMA that jumped across the tree. It’s really not very subtle when you know to look for it.
Another, even more obvious, example is viral DNA: this tends to end up as mangled, non-coding (“junk”) DNA, containing (fragments of) genes that often have nothing in common with the rest of the DNA, but have long stretches nearly identical to, say, a known gene for some viral coat.
In terms of quantification I’m really at the limits of my memory (and, unfortunately, in-flight internet), but I’ll take a stab and the former mechanism can be found in 5-10% of bacterial genomes, and amount to usually less than 8% of he genome, with maybe one or two exceptional cases with 20% or so (some bacterial species may have developed a tendency to exploit this sort of buffer overflow to cheat at evolution)
Phylogenists (the biologists working in trees, but not those “trees”) do consider all this (and much more). Where they encounter “known unknowns”, you will often see trees with nodes branching into more than two branches. That’s essentially what cartographers would lable “here be dragons”. Or, prosaically: it isn’t quite clear which split happened first in evolutionary history.