What's your best attempt? :-)
What's your best attempt? :-)
For example, "words" in agglutinative languages[1] (e.g., Turkish) act very differently from "words" in English. It's hard (impossible?) to capture all that variety in a pithy way. "A string of morphemes" might work, but that's hardly a satisfactory definition!
Maybe a good analogy for the HN crowd would be like asking, "How many characters are in a string?"
Each character is a syllable, with a particular pronunciation and a constellation of meanings, usually closely related to one another. There are very common combinations of characters that appear together, which one could define as words. However, often, you could just as easily view the individual characters as words, or the combination as a word.
In some cases, the combination of two characters means something totally different from what the characters alone would mean (e.g., 东西, where the characters literally mean "East-West," but the combination means "thing"), so the combination is clearly a word. But sometimes, the meaning of the combination is basically a combination of the words' meanings (e.g., 吃饭, where the characters literally mean "eat-food," and the combination means "to eat, have a meal").
Because written Chinese doesn't use spaces, I guess it doesn't really matter what one defines as a word. The issue just doesn't come up, practically speaking.
Also, the problems you raise here are mostly just as applicable to English, though perhaps for somewhat fewer words. Is a "walkie-talkie" one word, or two? How about "unmarried"? The un- prefix has a distinct meaning on its own, even if it never appears alone, after all. Or how about "can't"? Technically it's a contraction of "can not", and those words do sometimes appear separately as well, even in this same meaning.
In Chinese, just about every syllable has its own set of meanings. In English, there are compound words, and some words have prefixes or suffixes, but you can't just arbitrarily break a word into its syllables and assign a meaning to every syllable. Imagine if the word "syllable" could be broken down as syl-la-ble, and every person who spoke English could tell you what "syl," "la" and "ble" individually meant. That's the situation in Chinese, for almost every polysyllabic word you can utter. They can almost all be decomposed into syllables that have individual meanings. It's a very different paradigm from English.
I think it's obvious that there might not be one unifying definition that spans all human languages. "Word" is an English word refering to a specific of the English language. A class in JS is not the same as class in C++, big deal? ;)
As for English, I'm happy defining it through written language and spacing. "Can't", "unmarried", and "walkie-talkie" all one word each.
We might as well think of "word" in foreign enough languages as separate concepts. Doesn't seem meaningful to try to fit fundamentally different structures into the same conceptual molds. Which I guess ties back to your original point regarding if it's a useful excercise taxonimizing English as creole.
I’m not a linguist, but I would define a word as a part of a sentence composed out of one or more syllables, with word boundaries either implicitly or explicitly specified by different methods in different languages, e.g. by using pauses, longer or shorter phonemes, by using accents, rhythm, or intonation, or simply by remembering words as part of learning a lexical vocabulary.
A word is something that can be categorized as to which part of speech it belongs (noun, verb, adjective, adverb, etc.)
Depending on the languages it’s not always clear whether prefixes and/or suffixes are part of the word a separate words.
Similarly with compound words - do they count as a single or multiple words?
A short sentence in one language may enter another language as a single opaque word.
;)
BTW, should be "\b\w+\b". The \b is a zero-width match for the start or end of a word. Your pattern requires a space before and after:
>>> import re
>>> re.compile(r"\b\w+\b").findall("What's the problem?")
['What', 's', 'the', 'problem']
>>> re.compile(r"\s\w+\s").findall("What's the problem?")
[' the ']I was being facetious, obviously, but you've already identified a serious problem with that definition (is "what's" two words or one?)
cljs.user> (re-seq #"\s\w+\s" "やり直して")
nil
Joking aside, is it common to use regular expressions? Seems like the method only works for languages with spaces. I think a more sophisticated lexer may be necessary, but are there are non-regex, "fast approximations" that work across most languages? This is a problem that I have not tried solving before.