DOM is completely optional, and you only need it for convenience or to observe the modifications, e.g. the way browsers make use of it.
Just think how you would compile any XPath expression — it’s very similar to REGEX, only for structured data.
Worst case scenario XPath expression will not yield any results before traversing the entire tree. However any XPath 1.0 (and I imagine 2.0 is no different in that regard) can be compiled into a deterministic state machine (DFA), which only needs to keep tabs on how many elements it has seen and what conditions has been met.
What XPath 1.0 specifically doesn’t allow are arbitrary sub-expressions in predicates. Those would be problematic in certain conditions.
The only potential issue with performance is when running a large number of XPEs against a single stream. So there are various techniques to remedy that, including merging state machines for branching expression etc.