Are you on postgres? One thing you could do-- and I only suggest this cumbersome idea because you might just be crazy enough to try it-- would be to use the pg_trgm (trigram) extension with the following in mind: a) Theory being, when someone greps /[a-z]{4,8}/, they're either interested in {anthem, aardvark, ambition, ...} or {nltk, xkcd, json, zzxx, xxzz, ...}, likely not both. b) Neither (nor any third set you might come up with) is so inherently superior that it deserves default status over the other. c) Even with limiting results, half are bound to be totally uninteresting to the user. So what does that even accomplish?
So my pg_trgm suggestion is to take that same /[a-z]{4,8}/ result set and offer the user a relative sliding-scale by which they can push their visible 1,000 closer to/further away from a predefined set of dictionary words.
http://www.postgresql.org/docs/9.3/static/pgtrgm.html
You may also consider tech acronyms - maybe steal those from StackOverflow tags. Human names would be too big a hassle, IMO.
Again, I love the ambition of the damn thing. Kicks ass.