But it's still no different to any other language. An XML element has two collection members, one representing the attributes, and another representing the child nodes.
How is this different from a Java object which contains two list members called "attributes" and "children"?
How is this different from an S-Expression containing two lists?
How is this different from a json object like {"attributes":[...], "children":[...]}? Bear in mind that there is no requirement for JSON lists to be homgeneous. {"things":[1,true,"hello",3,{"addressee":"world"},[{"greeting":"Hola"},7],false]} is a perfectly valid JSON object. You don't have to define it as {"numbers":[1,3], "strings":["hello"] ...}.
It is a pretty common behaviour, that if you want a homogeneous list of things that differ, then you make abstractions until the differences disappear, e.g in an OO situation, you go up the inheritance tree until you are at the lowest common base class. In an XML situation, that common base class is "Node".
Even in a strongly typed language that requires homogeneity in lists, the only thing you know about the list members is that they can be cast to the same type, not that the members are of that type and no other, and certainly not that they all have the same name. Consider a C++ array of CFruit objects. It may have a member that is of class CBanana, one that of class CApple, and another of class COrange. If you want to do something COrange specific with the oranges, then you have to perform dynamic_cast<COrange> on any member of that list you suspect of being an orange. The same is true in a duck-typing situation.
The reason for my male & female example is that you would normally simply have a list of people. Of course, if your model contains no base that is common to both men and women, then you can't expect the machine to work that out, but if Man and Woman both inherit from Person, or if there is no Man or Woman, just Person with a member that specifies a gender, then it's trivial.