Learning how dictionaries work
arrantpedantry.com
arrantpedantry.com
Adding “deplatform” or whatever to the dictionary isn’t a statement about how proper it is as a word. It’s a statement that it’s a word in use, and so if you come across it, here’s what it means.
Without removing things you end up with the Oxford dictionary which is currently 20 volumes long. And I'm pretty sure they still remove things once it's considered obsolete.
And, one other very important feature I'd need, would be to be able to run the dictionary on itself, to find words that are used in the descriptions for which there isn't yet a definition.
This would be amazing, for example, to run on a large corpus, generate the dictionary, and then run it again to find words that are used but not defined - not just in the original corpus but in the definitions too.
I've yet to find anything like this and have managed, over the years, to do this with some cobbled-together sed and awk hacks .. but I still think this is something that would be quite a viable commercial product - especially useful for international translations and creating properly-defined glossaries for documentation, etc.
Anyone know of such tools? I'd love to have my own Dictionary builder, proper, and stop fantasizing about turning sed and awk scripts into a proper app ..
There are more complete versions of this kind of thing publicly available: https://github.com/turtlesoupy/this-word-does-not-exist
> This would be amazing, for example, to run on a large corpus, generate the dictionary, and then run it again to find words that are used but not defined - not just in the original corpus but in the definitions too.
I think this would be how you would gauge success of the model. That is to say, you would evaluate model accuracy on a set of held-out words with definitions that never appeared in your dictionary training set but appeared in context in your corpus. You would have to manually annotate whether or not the generated definition of these held out words was acceptable.
>I think this would be how you would gauge success of the model.
Yes, exactly. I think there would definitely be edge-cases, but the general rule is that there should not be any undefined terms/words in the final dictionary. The degree to which this can be achieved is of course related to the cyclomatic complexity of the original materials. But this is why I want this tool - to see how effective it is for creating training materials that prepare students for obtuse subjects.
It addresses some of the things you’ve mentioned.
EDIT: after visiting your site, somehow my browser language has been switched to Kannada and I'm now finding myself highly amused at seeing Twitter threads in this script. Would love to know how you did that!
They build definitions by the words, directly and indirectly, associated with them. You could then use those clusters to create individual definitions in some way.
Have you looked at them?
I think the biggest problem is that definitions are semantic which computers are terrible at. We’re at the infancy of being able to with transformers and large language models nowadays. So is start looking around for the proto-tools that will lead to what you're thinking about.
- words are created, or existing words bent, mainly by non-users of dictionaries.
- dictionaries cautiously track this process the benefit of dictionary users.
The nice thing about dictionaries is that nobody is forcing you to throw out your old ones.