The longest word you can type on the first row
rubenerd.com
rubenerd.com
I happen to have a corpus which includes pretty much every word ever written in a book, including many misspelled, mistranscribed, or otherwise non-dictionary words.
After eliminating nonsense, non-English, or other mistakes, I think the real winner, coming it at 12 characters, is:
teetertotter
That's a relatively common word. Even though it's usually seen hyphenated, the unhyphenated form is recognized by all the online dictionaries I found.----
And some other candidates, just for fun, in the 13 or 12 character range:
proproprietor
priorityqueue
reporterette
preprototype
"proproprietor" seems more like a misspelling. Should have a hyphen, or be two words."priorityqueue" is of course familiar to hackers here, but is more of a jargon term, and is only concatenated due to appearing in source code. Invariably it's two words when actually written out.
"reporterette" is antique, but appeared in a NYTimes headline as late as 2018 - the author reflected on her career, including sexist epithets. https://www.nytimes.com/2018/12/02/opinion/george-hw-bush-ma...
"preprototype" is used exactly as is, in lots of scientific papers, up to the current day. That's a pretty good one too, and could be a tie for "teetertotter", but it's verging on jargon.
Sorry for the questions, but it seems like an interesting, yet probably common data set and as someone who is venturing down this path, I’d like to learn more about building my own dataset similar to this from scratch.
lol, it's a 42MB text file from Google Books Ngrams.
The format looks like this:
$ head words-all.txt
a 14219615690
a! 196012
a" 84
a' 47713
a'0 3036
a'1 4070
a'10 99
a'11 56
I queried it with perl and sort. $ time perl -wlane 'if ($F[0] =~ /^[qwertyuiop]+$/) { print length($F[0]), "\t", $F[0] }' words-all.txt | sort -rn > qwertywords
real 0m1.915s
user 0m1.896s
sys 0m0.025s
I can't remember exactly which file I downloaded, but according to my notes I got it from here back in 2012 or so.https://storage.googleapis.com/books/ngrams/books/datasetsv2...
There seems to be a newer corpus published in 2020:
https://storage.googleapis.com/books/ngrams/books/datasetsv3...
grep '^[qwertyuiop]*$' /usr/share/dict/words | \
awk '{ print length(), $0 }' | \
sort -nI note that hardheartedness and hotheadedness threaten the darnedest nonstandard assassinations. Such sordidness!
Middle/Second row result is
8 flagfall "Flagfall, or flag fall, is a common Australian expression for a fixed start fee, especially in the taxi, haulage, railway, and toll road industries."
8 galagala "A name in the Philippine Islands of Dammara Philippinensis, a coniferous tree yielding dammar-resin."
Lower/Third Row: - None
There are no vowels on the bottom row. So no words. I've been typing at ~ 50wpm for 30 years, and I don't think I'd ever actually consciously recognized this fact about the bottom row.
(standard US keyboard layout)
And yeah, nothing in the bottom row other than acronyms and similar pseudo-words.
',.PY FGCRL pry or Lyly
AOEUI DHTNS tendentiousness
;QJKX BMWVZ xxxv, www, bbq or mm
After 'apt install wbritish-insane' pyrryl (a chemical group)
unostentatiousnesses (and anaesthetisations is good too)
mmmm $ grep '^[qwertyuiop]*$' /usr/share/dict/words | awk '{print length, $0}' | sort -rn | head
11 rupturewort
11 proterotype
11 proprietory
10 typewriter
10 tetterwort
10 repetitory
10 repertoire
10 proprietor
10 pretorture
10 prerequire
On Debian GNU/Linux 11 (bullseye): $ grep '^[qwertyuiop]*$' /usr/share/dict/words | awk '{print length, $0}' | sort -rn | head
10 typewriter
10 repertoire
10 proprietor
10 perpetuity
9 typewrote
9 typewrite
9 territory
9 repertory
9 puppeteer
9 prototype ↑7↑{⍵[⍒≢¨⍵]}words/⍨{''≡⍵~'qwertyuiop'}¨words
┌→─────────┐
↓peppertree│
│perpetuity│
│prerequire│
│proprietor│
│repertoire│
│typewriter│
│etiquette │
└──────────┘
Reading from the right, "test each word by removing 'qwertyuiop' and see if it leaves an empty string, use the test results to filter the input word list, descending-sort the length of each word and use that to arrange(index) the remaining words, flatten the array and take the top 7".(Longest from the middle row is 'haggadahs' then 'alfalfas', third row is 'mm')
% grep '^[qwertyuiop]*$' /usr/share/dict/words | awk '{print length, $0}' | sort -rn | head
11 rupturewort
11 proterotype
11 proprietory
10 typewriter
10 tetterwort
10 repetitory
10 repertoire
10 proprietor
10 pretorture
10 prerequire
So it seems that in addition to having parts of its kernel based on FreeBSD, there is also a lot of similarities in the wordlist at /usr/share/dict/words of macOS to that of FreeBSD :) perhaps even the same? C:\>grep '^[qwertyuiop]*$' /usr/share/dict/words | awk '{print length, $0}' | sort -rn | head
Bad command or file name
Bad command or file name
SORT: Too many parameters
Bad command or file name @echo off
mkdir c:\t
echo prompt echo %1 $g C:\t\%1 >> c:\temp.bat
command /c temp.bat %1 > c:\h.bat
del c:\temp.bat
type words.txt | find /V "a" | find /V "s" | find /V "d" | find /V "f" | find /V "g" | find /V "h" | find /V "j" | find /V "k" | find /V "l" | find /V "z" | find /V "x" | find /V "c" | find /V "v" | find /V "b" | find /V "n" | find /V "m" > c:\words2.txt
echo > h.bas DIM word as STRING
echo >> h.bas OPEN "C:\words2.txt" for INPUT as #1
echo >> h.bas DO WHILE NOT EOF(1)
echo >> h.bas INPUT #1, word
echo >> h.bas CMD$ = "C:\h.bat " + word
echo >> h.bas SHELL CMD$
echo >> h.bas LOOP
echo >> h.bas CLOSE #1
echo >> h.bas SYSTEM
qbasic /RUN c:\h.bas
dir /OS /B c:\t
@del c:\t\*
@rmdir c:\t
@del c:\words2.txt
@del c:\h.bat
@del c:\h.bas
You can play with MS-DOS 6.22 in a virtual-machine-in-browser here[1]. That VM comes with Vim (non-standard) so use Vim or Edit to create a word list and save as c:\words.txt. Then yype all this code into a batch file using `edit run.bat` and then run it with `run %1`. MS-DOS 6.22 came with QBASIC so I think that's allowed; I tried to avoid it but wasn't able to. NB. DOS is way less capable than Windows cmd prompt so there's no `for /f` or anything. "Dir /OS /B" sorts files by size and that view will leave the largest files on screen as the answer. The files will be one per word, containing the word so the size in bytes is the word length and the filename is the word to see it in the file listing. The words will be echoed into the files by a helper batch file containing `echo %1 > %1`. Building the helper batch file is hard because echo cannot echo > into a file. The qwerty filtering is a chain of `find /V "a"` for excluding each of the other rows cough. I then couldn't loop over the file lines without QBASIC.[1] https://copy.sh/v86/?profile=msdos
If you never used MS-DOS classic, try "edit test.txt" and see how it has a nice TUI, where Alt+F brings up the File menu, the brightly coloured letters are the hotkeys, so Alt+F, X will quit. Shift+Down will select a line, Shift+Delete to cut and Shift+Insert to paste. Ctrl+left/right arrows to jump forward/back a word, Ctrl+Shift+Left/Right to select a word. 29 years later those keyboard patterns still work in this FireFox editor, in current notepad, WordPad and Word, and in my muscle memory. Escape tends to exit back out of popups and menus. Quit and try "help date" and see the TUI help, where the green angle brackets are hyperlinks and can be TAB'ed between, Enter to activate and Escape back. F1 is still the help key, only it actually showed offline help back then instead of doing a Bing search for 'get help in notepad'. Quit and run QBasic, see how F5 runs the code.
[2] by Geoff Cutter: https://groups.google.com/g/alt.msdos.batch/c/Ozg2C-ANCqI
[3] Phil Robyn's QBASIC loop sample https://groups.google.com/g/alt.msdos.batch/c/44NbdZJ2-p4/m/...
awk '/^[qwertyuiop]+$/ {print length, $0}'I found a few earlier this year and I was going to file a bug so I did some research to find out this is a known and expected behavior.
If the computer say reads stubborn as stubbum, the smaller dictionaries are the ones that have cross checked and filtered those out. The insane ones do not. It's a good name. "Lack of sanity checks"
Here's an example word I found, "suabilities". You'll find it only on wordlist sites that used this wordlist and I guess, now here.
I wonder how often that happens - surely there's tons of people dealing with japanese text who can't read it and just use diligence to make sure the "letters are the same"
In the start The One has risen the stars and the earth.
The earth had no order, and nothin' resided there; and shade resided on the nonendin' 'neath. And The One rided on the seas.
Then The One said: "I desire it to shine"; and it shone.
And The One had seen the shine, that it's neat; and The One sorted the shine on one side, and the shade on the other.
The One then denoted the shine and the shade. So the nite and the shine that are date no. one had ended.
[1]: https://colemak.com/Funhttps://godexperiment.org/beginnings-an-alliterative-rewrite...
It was inspired by these versions:
https://llamasandmystegosaurus.blogspot.com/2017/05/alpha.ht...
First row
$ awk '/^[,.pyfgcrl]$/ { print length(), $0 }' /usr/share/dict/words | sort -nr | head
3 pry / ply / fry / cry
Second row
$ awk '/^[aoeuidhtns]$/ { print length(), $0 }' /usr/share/dict/words | sort -nr | head
15 tendentiousness
14 assassinations
13 instantaneous
13 insidiousness
Third row
$ awk '/^[;qjkxbmwvz]*$/ { print length(), $0 }' /usr/share/dict/words | sort -nr | head
4 xxxv
3 xxx
3 xxv
2 xx
[...] 15 sententiousness 15 sinuatodentated 15 soundheadedness 15 tendentiousness 15 uninitiatedness 16 antisensuousness 16 ostentatiousness 17 dissentaneousness 17 instantaneousness 18 unostentatiousness
$ awk '/^[;qjkxbmwvz]*$/ { print length(), $0 }' /usr/share/dict/words | sort -nr | head 4 xxxv 3 xxx 3 xxv 2 xx 2 xv
If I start at the top row and go down, I can make TAXES but couldn’t think of a longer word. The third row having no vowels makes it so hard.
Starting at the bottom row and going up, I came up with CHICKEN which is delicious and neat that it ends where it started. Chickens is longer but ends on the middle row which is not as neat I feel like :(
A dictionary search turned up "paxwaxes" as the longest word I could find that starts in the top row and goes down, wrapping around to the top every three letters.
> Starting at the bottom row and going up, I came up with CHICKEN which is delicious and neat that it ends where it started. Chickens is longer but ends on the middle row which is not as neat I feel like :(
Chickens is indeed the longest.
If you start at the bottom row and go up-and-down: cataclysms, or catamarans.
If you start at the top row and go down-and-up: escapable.
If you start in the middle and go down-and-up: scarabaean
If you start in the middle and go up-and-down, I didn't find anything longer than 7 letters, and there were 39 seven-letter words, including "discard", "grandpa", and "stacked".
FYI, rupturewort is the sole 11-letter word answer to TFA in YAWL; found using:
grep '^[qwertyuiop]*$' word.list |while read -r line; do echo "${#line} ${line}"; done |sort -n | tail
1: https://github.com/elasticdog/yawlIt's 170k English words, no placenames or people's names or anything like that, but does have some that I question how valid they are.
/usr/share/dict/words is always destination zero for words.
There's four: a hamlet named Treopert, a mountain range in Snowdonia named Eryri, a house in Bangor University named after that mountain range, and TUI outdoor shop.
Curiously the longest top-row places I can find anywhere on OpenStreetMap are almost all roads in France:
Rue Pierre Ropert
Rue Pierre Riquet (x2)
Rue Pierre Poutot
Rue Pierre Potier (x4)
Rue Pierre Perret (x3)
Rue Pierre Pietri
Route Petit Peyre (x2)
Poirier Pierrotte
Petite Rue Pierre
Rue Pierre Perrier (x3)
Rue Pierre Pottier (x3)
Rue Pierre Routier
Rue Poirier Piquet
Rue Pouyer Quertier (x2)
Route Pourpre et Or
Tyrepower Port Pirie <-- a shop in Australia
Petite Route Petite RueMy understanding of German grammar is that words can be of infinite length since they allow unlimited compounding. But they also have a language authority that makes words official, and the current longest is 68 letters.
So it would indeed be an interesting exercise in German.
Duden (the standard German dictionary) is the closest I know. And they list "Aufmerksamkeitsdefizit-Hyperaktivitätsstörung" (the German word for attention deficit hyperactivity disorder) with 44 letters as the longest in the dictionary [1].
*Update:* There are also [2] and [3] but they are both not anymore part of a law (the "-gesetz" suffix in the word) or regulation (the "-verordnung" suffix in the word), respectively.
[1] https://www.duden.de/sprachwissen/sprachratgeber/Die-langste... [2] https://en.wiktionary.org/wiki/Rindfleischetikettierungs%C3%... [3] https://en.wiktionary.org/wiki/Grundst%C3%BCcksverkehrsgeneh...
$ axel -n4 https://dumps.wikimedia.org/enwiktionary/latest/enwiktionary-latest-pages-articles.xml.bz2 && bunzip2 enwiktionary-latest-pages-articles.xml.bz2 # 1.1G -> 8.1G
$ grep '<title>' enwiktionary-latest-pages-articles.xml >titles # 282M
$ sed 's/ *<title>//' titles |sed 's/<\/title>//' >clean-titles # 126M
$ grep '^[qwertyuiop]*$' clean-titles | awk '{ print length(), $0 }' | sort -nr|grep ^1[23456789]
15 retroproiettori # italian <-- joint italian winner
15 retroproiettore # italian <-- joint italian winner
13 topprioriteit # dutch <-- dutch winner
13 ripitturerete # italian
13 retorqueretur # latin <-- joint latin winner
13 repropitietur # latin <-- joint latin winner
13 purpurroterer # german <-- german winner
13 proterreretur # latin <-- joint latin winner
13 perterreretur # latin <-- joint latin winner
13 perquireretur # latin <-- joint latin winner
12 teetertotter # english <-- english winner
12 riproporrete # italian
12 ripitturerei # italian
12 riotturerete # italian
12 retorquetote # latin
12 retorquerere # latin
12 requiriertet # german
12 requireretur # latin
12 repropitiere # latin
12 repperiretur # latin
12 purpurrotere # german
12 prototypique # french <-- french winner
12 proterruerit # latin
12 proterrituro # latin
12 proterrituri # latin
12 proterriture # latin
12 proterretote # latin
12 proterrerere # latin
12 protereretur # latin
12 proriperetur # latin
12 proietterete # italian
12 priorytetowy # polish <-- joint polish winner
12 priorytetowo # polish <-- joint polish winner
12 prioriteetti # finnish <-- finnish winner
12 prepotettero # italian
12 portretterte # bokmål <-- joint bokmal winner
12 portretterer # bokmål <-- joint bokmal winner
12 portretteert # dutch
12 piroetterete # italian
12 pipettiertet # german
12 perterruerit # latin
12 perterrituro # latin
12 perterrituri # latin
12 perterriture # latin
12 perterretote # latin
12 perterrerere # latin
12 perquiritote # latin
12 perquirerere # latin
12 perpetuerete # italian
12 perpetrerete # italian
12 perpeteretur # latin
12 perequitetur # latin
12 iperprotetto # italian
12 iperprotetti # italian
12 iperprotette # italian
12 eqqoqqortooq # greenlandic <-- greenlandic winner canvasbacks
counterconvention
photofluorographies canvasbacks omits letters from the top row
counterconvention omits letters from the middle row
photofluorographies omits letters from the bottom row
Presumably the longest word available in each case?this allows you to, among other things, tune the comprehensiveness/accuracy tradeoff to your liking for a particular task by cutting the list off at a given point
probably i should download the google 1-grams now that i have a bigger disk
anyway
$ grep ' [qwertyuiop]*$' ~/wordlist | perl -lane 'print length $F[1], " $F[1]"' | sort -n
indeed suggests only 10 perpetuity
10 proprietor
10 repertoire
10 typewriter
possibly a more practical problem is, what are the most common words you can type entirely with the left hand while your other hand is on the mouse†; 'redraw' was a significant one with early versions of autocad $ grep ' [qwertasdfgzxcvb]*$' ~/wordlist | head
2150885 a
923975 was
664780 be
478178 at
470949 are
uh that's somewhat boring $ grep ' [qwertasdfgzxcvb]*$' ~/wordlist | perl -ane 'print "$F[1] " if length $F[1] > 6' | fmt | head
greater started effects address average regarded created treated affected
greatest streets readers afterwards referred database degrees dressed
arrested attract decades attracted addressed stressed stewart awarded
targets estates abstract assessed terrace reserve reverse barbara reserves
careers defeated dragged grabbed creates greeted secrets bastard extract
edwards detected deserve affects traders adverse retreat addresses rewards
aggregate decrease scattered breasts referee databases deserted actress
debates deserves reversed deserved defects reserved fastest rewarded
steward battered erected deceased exaggerated cassette exceeded servers
warfare extracts drawers stresses asserted reacted sweater regards stabbed
that's a bit better. how about words you can type alternating the two hands, so you can type faster $ egrep ' [^qwertasdfgzxcvb]?([qwertasdfgzxcvb][^qwertasdfgzxcvb])*[qwertasdfgzxcvb]?$' ~/wordlist | perl -ane 'print "$F[1] " if length $F[1] > 6' | fmt | head -4
problem problems england chairman alright element ancient visible penalty
quantity visitor signals amendment claudia bicycle authentic antibody
malaysia naughty dickens entitlement antique paisley rituals auditor
endowment blanche chairmen siemens chaotic suspend uruguay mcleish
it's a curious experience to touch-type a sequence of these words because after a while you notice that something is unusual. the one-handed words are a bit more conspicuous. try typing 'a better career award as edward created database facts after we agreed we stared at dear steve' into the comment box, it's super weirdof course the bnc has some built-in biases which dramatically understate the frequency of certain words
$ grep fuck ~/wordlist
2568 fucking
1236 fuck
158 fucked
53 fucker
23 fucks
20 fuckers
13 motherfucker
5 motherfuckers
5 fuckin
'ngl' appears but only 5 times. 'yeet' doesn't occur at all______
† mouse?
$ grep ' [qwertyuiop]*$' ~/wordlist | perl -lane 'print length $F[1], " $F[1]"' | sort -n
in common lisp (defconstant *wordlist*
(with-open-file (in #P"~/wordlist")
(loop for line = (read-line in nil)
while line
for p = (position #\space line :from-end t)
collect (list (parse-integer (subseq line 0 p))
(subseq line (1+ p))))))
(defconstant *topwords*
(let ((w (loop for (freq word) in *wordlist*
if (loop for c across word always (find c "qwertyuiop"))
collect word)))
(sort w #'< :key #'length)))
those last four lines of code seem very acceptable to me but not really competitive with the unix approach for interactive experimentation(parse-integer (subseq line 0 p)) is shorter (parse-integer "12345678" :end 3)
(It's still amazing to watch him explain that they don't look at the mouse while moving it, they look at the pointer).
I don't know how it compares to left-handed (or right-handed) Dvorak.
https://en.wikipedia.org/wiki/One_hand_typing
https://en.wikipedia.org/wiki/Dvorak_keyboard_layout#One-han...
http://web.archive.org/web/20090727071756/http://lists.canon...
my hack was kind of half-assed but worked well enough that i was convinced that it was a bad idea, so i decided not to do the second buttock
the critical path for fast typing is the precision with which you can synchronize the motions of different fingers, and in particular different hands; if keydown events happen in the wrong order, you start to get tranpsosition erorrs
chording keyboards don't care in which order the keys in the chord start; they only care what the set of keys in the chord is and when it ends (so they can stop looking for new keys to add to the set). fast typists on chording keyboards can do 300 words per minute, which is about 7 chords per second, which is about as many "strokes" as a normal typist on a non-chording keyboard, just with chords (producing a syllable each) instead of individual keys (producing a letter each)
adding a required modifier key that has to happen before a keystroke, and end before the next keystroke, is the opposite; it adds more things you have to sequence correctly. in the half-keyboard patch case, in particular, it adds 50% more things. this slows down your typing by about a third
exaggerated
stewardesse