I worked in phylogenetics for a while and it was a pretty confusing area. Originally, phylogenetic trees (not the biological trees that are the subject of the OP) were created by finding physical features (yes, just like ML) and using those to build a semi-supervised tree-structure of classifications. However, eventually we began to use DNA sequences to compare organisms, which restructured the tree in many ways, even close to the root. It was a controversial time as the the historical physical-feature classifier group was certain their way was right, and same for the DNA folks. I sort of assumed that the DNA would be a much higher quality source for clustering but it hasn't really always worked out that way.