A Spellchecker Used To Be A Major Feat of Software Engineering
prog21.dadgum.com
prog21.dadgum.com
While it seems trivial for Google or a database backed server to provide real-time intelligent suggestions (e.g. suggest 5 words that follow: Harry), implementing such a feature on iPad took me over a month even though I knew exactly what I wanted to make and had all the necessary data beforehand. I had a list of 1 million words and phrases with frequency of usage (16mb of text data) and wanted to suggest 3-7 words in real-time as the user typed each letter. And implementing this on iPad required quite a bit of engineering and optimization.
Doing a half-decent production spell checker is STILL a major feat. Same as "just crawling the web" (further down the discussion). Both require problem understanding and engineering you can't see and appreciate at a glance.
And no, looking up individual words in some predefined dictionary doesn't qualify as half-decent spell checking, especially for non-English languages. Spelling correction is another step.
"Their coming too sea if its reel."> And no, looking up individual words in some predefined dictionary doesn't qualify as half-decent spell checking,
Well, but the author is talking about that problem! Even if you don't consider that real spell-checking, his point still stands. Let's define crappy-spell-checking as "looking up individual words in some predefined dictionary"; that problem used to be hard and now it's very easy, as in, you could write one in 15 minutes using Python.
I read the point of the article to compare "spell-checking then (80s) and now", whereas others read it more along the lines of "looking up static English words then and now". Your nickname sounds japanese, but I assume you're talking about English as well, with those 15 minutes.
That corrector has no context though, so it will not correct misspelled words that happen to come out as other words (incl. the "Their coming too sea if its reel.").
This is especially important on the web, where pretty much every conceivable word is "correct" (a name of a company or otherwise). False positives are costly, and context disambiguation critical. Let the fun begin.
This is GRAMMAR checking, or at least grammar-assisted spell checking.
Very few, if any, shipping mainstream spelling correctors do that.
The example sentence has misspelled words, hence is within the domain of spell checking. This type of misspelled words are called "homonyms", which are one very common spelling problem. The academic terminology is uninteresting to most users, however.
Or if you mean to posit that "spell checking really means looking up if a word exists in a static dictionary of English", then yes, that's easy and solved, no argument there.
No, it doesn't . It is grammatically incorrect, but all the words are spelled correctly. You're definitely talking about a grammar checker, not a spelling checker.
In your view, "spelling check" is applied to individual words, to see if they appear in a fixed dictionary (see my other comment about difficulties with choosing this "correct" dictionary in reality, though). I can imagine this view is inviting for programmers, because it's easy to implement, but I doubt anyone else finds useful a definition that says there are no misspellings in "Their coming too sea if its reel."
In my view, "spell check" applies to utterances and roughly means "all words are spelled as per the norm of the language; I can send this document to my boss/customer and they won't laugh at my spelling." It is a more user-centric view, and more complex too, because it covers intent and norms, as opposed to the comforting lookup table for a few hand-picked strings. Modern spell checkers make heavy use of statistical analysis of large text corpora, to reasonably approximate context needed to model such intent.
Once we agree on the terminology, I believe we are in agreement, so let's not split hairs. The "correctly spelled" sentence under question comes from the Wikipedia article on spell checking, by the way.
I realize this is just a debate over a definition, so it's not very meaningful, but the fact is, you're on the wrong side of the common definition here. I just tested, and every spell checker that I just checked (Chrome, Firefox, MS Word, TextMate, TextEdit - perhaps some of these rely on the same underlying engine, I'm not sure?) accepts that sentence as not having a spelling error, so clearly there's some use for such a definition.
The grammar checkers, on the other hand, don't like it, but by changing "their" to "they're", they all accept it, despite the fact that it's still garbage. So don't overestimate how good "modern" spell checkers are...though there may be techniques to do a better job, they're not in common use, at least in the most common spell-checking contexts (which, lets be honest, pretty much means MS Word).
There is also no non-words. It happens that today's spelling checkers are generally just non-word detectors. This doesn't mean that anyone defines "misspelled" to mean "misspelled in such a way that the result is not a word at all".
And yes, a non-word detector is still very useful, and it's much much easier to make than something that also determines reliably when words are misspelled in ways that produce other words, so there's lots of software out there (perhaps essentially all of it) that contains only a non-word detector.
I think it's perfectly reasonable to define "spelling checker" to include mere non-word detectors. Or for that matter non-common-word detectors. (I expect most spelling checkers will reject "hight", and very sensibly because if someone types that they probably meant "high" or "height" or something -- but it's a perfectly good word, albeit a rare and archaic one.) But there's no way it's correct to say that "sea" isn't misspelled in that sentence merely because the mistake happens to have produced something that's an English word.
Sure it is (reasonable).
You just cannot go around acting as if such a definition is already widespread, established in tech discussions, and followed by spell checkers.
No, it contains correctly spelled words used in place of other desired words.
"I sea your eyes", "I she your eyes"
If you want to correct these kind of errors, you must now a lot about natural language. Also, the above are trivial cases. There are tons of edge cases and far more difficult distinctions. Here's an amusing one, that can lead to "Microsoft paperclip" like interactions:
"I gave him the new pink dress as a present" => "I gave her the new pink dress as a present"
No, idiotic spell-checker, I do mean him. My friend is a cross-dresser, shut up and let me type.
Such a spellchecker would also be useless for poetry. And if you find poetry obscure, so it doesn't really matter, then such a spellchecker would also be useless for irony. Suddenly, you lose all the hipsters from your potential users (except if they start using it ironically).
Anyway, no spell checker in widespread use attempts this --and it's probably a very hard nut to crack, and probably uncrackable in the general case.
If you intend to write a word meaning "actually existing as a thing" and spell it "reel", it is monstrously unlikely that you genuinely thought a word meaning "a cylinder on which flexible materials can be wound" was a suitable substitute. If you did, that would be the use of a correctly spelled (but incorrectly selected) word in place of a desired word: a grammatical error, in other words. However, if you just pick the wrong letters to construct a phoneme, as is almost certainly the case here... yep, that's a spelling error.
To put it another way, imagine that the word "reel" didn't actually exist, and I make precisely the same error, substituting an 'e' for an 'a'. All of a sudden, by your argument, what was a grammatical error is now a spelling error. But the mistake I made hasn't changed, so that makes no sense.
> If you want to correct these kind of errors, you must now a lot about natural language. >... > Anyway, no spell checker in widespread use attempts this
So? The category of error doesn't change with how difficult it is to fix.
I would say that Norvig's corrector is a first-order spelling corrector, since it works within the context of a single word.
A second-order corrector would take into account the word before or after it to choose the spelling that is more likely to make sense. ("Their coming" would suggest a correction to "They're coming")
Third-, fourth-, (and so on) order expands the distance of words considered.
How about: "Her relatives would visit us for Christmas. Their coming filled us with dread!"
That no spell checker might be able to catch this specific class of error doesn't change the type of error it is.
What makes working in a finite, very limited set of resources so rewarding is that those limitations turn mere programming into art.
You can't have art unless you constrain yourself somehow. Some people paint with dots only and some people express themselves in line art. If they allowed themselves any imaginable method that is applicable, they wouldn't be doing art. They could just take a photograph of a setting, and that photograph wouldn't say a thing to anyone.
Endless bit-twiddling and struct packing may turn trivial methods into huge optimization-ridden hacks and not get you too far vertically but given only few resources, those hacks are required to turn the theoretic approach into a real application that you can actually do useful work with. Those hacks often exhibit ingenious thinking and many examples of that approach art—and the best definitely are. And any field where ingenious thinking is required will push the whole field forward.
Similarly, for example, using a Python set as a rudimentary spell-checker is fast, easy, and convenient but it's no hack because it requires no hacking. It's like taking that photograph, or using a Ferrari to reverse from your garage out to the street and then driving it back. Which ingenious tricks are you required to accomplish that? None.
The bleeding edge has simply moved and it must lie somewhere these days, of course. Maybe computer graphics—although it has always demonstrated the capability to lead the bleeding edge so there's actually no news there. The fact is that the bleeding edge is more scattered. Early on, every computer user could feel and sense the bleeding edge because it revolved around tasks so basic that you could actually measure with your own eyes. Similarly, even a newbie programmer would face the same limitations early on and learn about how others had done it. Now you can stash gigabytes of data into your heap without realizing that you did, and wonder why the computer felt sluggish for a few seconds. Or how would you compare two state of the art photorealistic, interactive real-time 3D graphics demos? There's so much going on beyond the curtains that it's difficult to evaluate the works without having extensive domain knowledge in most the technologies used.
Findind the bleeding edge has become a field in itself.
> You can't have art unless you constrain yourself somehow.
> They could just take a photograph of a setting, and that photograph wouldn't say a thing to anyone.
While I realize that art is subjective, I'm very surprised that you would put these conditions around what you consider to be art. Especially since it seems to be a condemnation of a large subset of photography.
A large subset of photography isn't art. In fact, most everything people create isn't art per se—if it were, there wouldn't be good art nor bad art, just art and lots and lots of it. Spend a few hours on some photo-sharing site and see what people shoot. They're photographs, but rarely art.
But there are grades of art. Look at this search: http://goo.gl/2mLVI — a thousand sunset pictures, while maybe pretty, aren't generally art and not because it's the same sun in each picture. Most of these pictures have nothing to say. Now, some object lit by the sunset or silhouetted against it gives a lot more potential to be art. A carefully crafted study of a sunset in the form of a photograph can be art, but it requires finding certain constraints first, finding a certain angle that makes the photograph a message, and eventually conveying through the lens something that makes the viewer stop for a moment, to give an idea, to give a feeling, to give a confusion.
> A large subset of photography isn't art
> a thousand sunset pictures, while maybe pretty, aren't generally art
Seriously? Because you get to decide? Your response just seems incredibly egocentric. You can define what art means to you all day long, but you can not define what art means to everyone.
I guess we'll just have to disagree.
Some may say something is lost and resources wasted - taken for granted as we now brute force our way through such problems. Surely going backwards?
but now we are, a million times a second; free to disambiguate a word's meaning, check its part of speech, figure out whether its a named entity or an address, figure out if it is a key topic in the document we are looking at and write basic sentences. I agree. That is progress.
I think it is progress: Yes, we can brute force today through the problems that were a feat a decade or two ago. But: Not having to solve those problems frees a lot of time. Time that can be spent on problems that are a major feat today.
Getting better at a problem in one domain can spin off benefits in others.
Spelling is part of language, and language is something that computers are really bad at. Brute force helps a bit with that (auto-correct; siri;) but better understanding would be cool.
[1] (http://www.spellingsociety.org/journals/j20/spellchecking.ph...)
Morris, Robert & Cherry, Lorinda L, 'Computer detection of typographical errors', IEEE Trans Professional Communication, vol. PC-18, no.1, pp54-64, March 1975.
Spell checking is not about comparing a list of words to what the user wrote and tell him what didn't match. It's much more about understanding the users intention and helping him shaping that intention into an officially recognised grammar/spelling. Example: "then" is a correct word. But in the context of "Google's spell check is better then Word's" "then" is actually wrong. (Google also doesn't tell you about that mistake but the first search result actually contains a "than", which is recognised as what you actually meant)
I hope I could make it clear, why I think it still is a major feature to have the best spell checker and I think, the title really should be revised.
Actually it seems that everybody reading the article got the same WRONG impression.
The author full well knows what it takes to do a i18n full-featured spell checker.
That is BESIDE the point.
What he says is that doing a basic (lame ass) spell-checker in the 80s used to be a MAJOR undertaking, and now, doing EXACTLY THE SAME is trivial.
His point is not about spell-checking.
It's about modern OS, language, CPU, HD and memory conveniences, vs what one had to deal with in the olden days.
See also the Digital Antiquarian's fascinating (and reasonably technical) blog on the origins of Infocom: http://www.filfre.net/tag/infocom/
They used a data structure they dubbed a DAWG (directed acyclic word graph) that was a trie for common start patterns with words and also collapsed together common endings (every "tion" ending word ended in same portion f the graph. As I recall they used 5 bits to store ASCII uppercase characters and either 11bits or 19bits to address nodes in the graph leading to a data structure that compressed words as well as any lzw type compression with the advantage that it was directly readable with respect to word extraction.
While the paper was published in 1988, the program was written and being used much earlier in '82 or '83.
http://wiki.cs.pdx.edu/cs542-spring2011/papers/appel-scrabbl...
That site also has a paper on the development of spell (http://unix-spell.googlecode.com/svn/trunk/McIlroy_spell_198...).
An Apple II spellchecker that I used a few times seemed to spell-check using brute force means with the dictionary filling an separate 143K floppy disk. It was a somewhat slow batch-like one document-at-time operation, but it was good enough to be useful.
As a programmer in the 80's, I saw a spellchecker that used a small hash table to tell whether a word was misspelled or not. The hash table could be made even smaller if we limited the dictionary to about 16000 of the more common words. If I remember correctly, the smallest hash table was generated as an array of 16 bit integers in the source code before compiling and it's small footprint kept in memory of the built program. Once a word was identified as a misspelling, using a very fast hash look-up, suggestions of words could be pulled from the actual dictionary on disk. Using the hash, there was the possibility of false spellings. Maybe this was a Bloom Filter.
They do away with a dictionary.
As an example, a few months ago a young developer I worked with implemented a set of attributes on a table as a bit field. He had just read an article about then and did it because it sounded cool. It made running reports against the data and doing imports via plain SQL impossible without writing a bunch of extra functions. His clever solution that saved about $0.0001 worth of disk ended up costing several hours of developer time because he didn't just a join table and a foreign key.
Eh, that needs qualification to make sense. Just crawling and trivially tokenizing the web counts as "indexing all the individual web pages" but that's never been such a monstrous task.
Having results fresh to-the-day/hour/minute is a better example of something that will probably be looked at is child's play even though that's a very new, very computationally expensive development in search.
Yes, you don't have the memory complications that are really hard, but you still need to think. Get a proper data-structure (Trie, for ex), and fill it with all forms of the words.