\s\w+\s
;)
;)
BTW, should be "\b\w+\b". The \b is a zero-width match for the start or end of a word. Your pattern requires a space before and after:
>>> import re
>>> re.compile(r"\b\w+\b").findall("What's the problem?")
['What', 's', 'the', 'problem']
>>> re.compile(r"\s\w+\s").findall("What's the problem?")
[' the ']I was being facetious, obviously, but you've already identified a serious problem with that definition (is "what's" two words or one?)
cljs.user> (re-seq #"\s\w+\s" "やり直して")
nil
Joking aside, is it common to use regular expressions? Seems like the method only works for languages with spaces. I think a more sophisticated lexer may be necessary, but are there are non-regex, "fast approximations" that work across most languages? This is a problem that I have not tried solving before.