Dartmouth computer scientists expanding diff, grep Unix tools
itworld.com
itworld.com
http://eelco.lempsink.nl/thesis-eelco-lempsink.pdf
It allows you to have meaningful diffs for any kind of datastructures, not just lists of strings. The code is also available as a Haskell package: http://hackage.haskell.org/package/gdiff
The new programs, called Context-Free Grep and Hierarchical Diff, will provide the ability to parse blocks of data rather than single lines. For each new type of data structure, a vendor would provide a pattern library identifying the basic structure of the data, which the software would then use to "extract the constructs of interest from the document," Weaver said.
So how does the above fit into the Unix Philosophy's dictum that "Everything is just a stream of bytes (of text)?" Or does it?
Neither grep nor diff operate on a stream of bytes in their most common form. They operate on data structures called "lines" extracted from the stream of bytes according to a pattern hardcoded into the tools themselves.
Now that the camel's nose is under the tent…
But that exactly was my thinking as well. Complicated data-types sure need different tools, but if you want to use somethinng like grep and diff, it would be nice to have simple data. Which leads to the thinking that not grep and diff are the issue, but the data. Though the article even states it's not about replacement, but about a different area of usage:
>Such utilities could provide administrators the ability to work with more complex forms of data now being generated by network equipment and infrastructure software.
The natural inverse to hierarchical diff should be hierarchical patch, which would be (more-or-less) a generalization of what coccinelle does.
Is there a link to the paper somewhere?
Having read the paper, it seems its more specialized to cisco IOS config files - they're building a library of patterns you can match against, the paper doesn't explain how you can add your own. And it does seem to be CFG rather than PEG - there's a separate token library.
Ward's work looks more usably generic than this, it's not cited so I presume they weren't aware of it.
Relying on vendors to supply machine readable data explaining their file syntax seems... optimistic. Maybe "It's JSON" will be enough?
Also, I'm curious how the patterns will be specified. I can imagine it quickly becoming complicated.
I'm also not sure the problem is enough of a problem for people to learn the new pattern matching syntax. The occasions where I would have needed this have been inconvenient, but not enough that I bothered to look if tools existed. I usually just crank out a quick Python script to do it.
Specifically for conffiles, there's already Augeas (http://augeas.net/), a good, practical tool to map many Unix/Linux formats to trees of text.
(I'm referring to the paper: http://www.cs.dartmouth.edu/reports/TR2011-705.pdf )
From the article:
"...Gabriel Weaver, a Dartmouth graduate student who, along with Dartmouth computer science professor Sean Smith, is creating the variants of grep and diff. Weaver presented the new utilities at a poster session at the Usenix Large Installation System Administration (LISA) conference, being held this week in Boston."