There's projects on this - one of the best known ones (to my knowledge) is the Stanford IP Litigation Clearinghouse Project:
The good news, though, is that legal documents tend to follow a fairly narrow channel of variations, when isolated to particular practice areas (e.g., leases, sales of goods, service agreements, motions, etc.)
I've always wanted to run a huge number of documents through Beyesian filters or something similar to develop some interesting classification rules, but it's damn hard to get a pool of representative documents that isn't strictly confidential.
http://topsy.com/pycon.blip.tv/file/4879824/?allow_lang=en
which then links to a video of the presentation (watching it now)
http://blip.tv/pycon-us-videos-2009-2010-2011/pycon-2011-how...