And, one other very important feature I'd need, would be to be able to run the dictionary on itself, to find words that are used in the descriptions for which there isn't yet a definition.
This would be amazing, for example, to run on a large corpus, generate the dictionary, and then run it again to find words that are used but not defined - not just in the original corpus but in the definitions too.
I've yet to find anything like this and have managed, over the years, to do this with some cobbled-together sed and awk hacks .. but I still think this is something that would be quite a viable commercial product - especially useful for international translations and creating properly-defined glossaries for documentation, etc.
Anyone know of such tools? I'd love to have my own Dictionary builder, proper, and stop fantasizing about turning sed and awk scripts into a proper app ..