Processing documents with Clojure transducers
blog.juxt.pro
blog.juxt.pro
"xml/parse parses the file into a tree structure like this"
So aren't we fitting the whole document into memory anyway?
For one thing, transducers can be used in alternative "reducing contexts", for example, a core.async channel. If you define map/mapcat/filter/etc in terms of a concrete data structure (such as lists), you can't reuse them as readily.
Another perf-ish reason for transducers is separate compilation. It's dramatically easier to fuse loops for a high-level symbolic representation, but can get much trickier once you only have byte-code left. All Clojure functions are compiled immediately upon creation and the source code is discarded. By being built out of function calls, you can have package A define a transducer and package B compose it with another tranducer, without having to perform inter-module optimizations. And the JIT will perform inline across modules at runtime!
Having said that, there's an alternative approach that can be made to work too: yield. Not shallow yield; delimited-continuation / monadic yield. Scala's collections approximate this idea with effectful Traversable and such, but really it's not quite right in both performance and flexibility.
So yes, Transducers are a bit of a hack to accommodate the host, but no, they are not totally without novelty or intrinsic value.
I will admit that the multiple arities is a bit awkward when an interface (or two) could have done the trick. I'm not 100% sure I understand why Rich choose to do it the way he did. I suspect it was so that `comp` would work.
I probably would have just defined ITransducer or something and supplied a custom composition function, but then again, the Haskell lens package takes the same approach, preferring composition via `.` on functions instead of extending a hypothetical "Composable" or "Pipelinable" type-class upon which `comp`/`.` could be built.