The new programs, called Context-Free Grep and Hierarchical Diff, will provide the ability to parse blocks of data rather than single lines. For each new type of data structure, a vendor would provide a pattern library identifying the basic structure of the data, which the software would then use to "extract the constructs of interest from the document," Weaver said.
So how does the above fit into the Unix Philosophy's dictum that "Everything is just a stream of bytes (of text)?" Or does it?