NLP in Python vs other Programming Languages
nltk.googlecode.com
nltk.googlecode.com
NLTK is supposed to be an educational toolkit. It's used by linguists taking their first steps in programming, and by CS students taking their first steps in complexity and mess of human language. They're not looking for the shortest code, the fastest code, or the most <quality attribute X> code, just the most readable, insofar as readability can be supported and encouraged by a language.
Clearly the author has made a strong assumption that the peculiarities of Python syntax and semantics (`import`, `sys.stdin`, `for` ... `in`) are somehow clear to Python novices.
for line in ARGF
for word in line.split
if word.match(/ing$/) then
puts word
end
end
end
I'd write as for line in ARGF
puts line.split.grep(/ing$/)
end
Or puts ARGF.map{|line| line.split.grep(/ing$/)} ARGF.each {|line|
line.scan(/\b\w+ing\b/).each(&:puts)
}
You know, why are we breaking this by line? :P ARGF.read.scan(/\b\w+ing\b/).each(&:puts)
Probably requires ActiveSupport for some versions of ruby :) The &:symbol trick isn't in all versions. while line = gets
puts line.split.select { |word| word =~ /ing$/ }
endThe equivalent Python is something like
import sys
for line in sys.stdin:
print '\n'.join(filter(line.split(), lambda word: word.endswith('ing'))
which I would argue is more readable to a layperson because function calls look like function calls. (By the way, does Ruby's puts automatically do the join with newlines, or what?)The post seems to be about language independent readability. Rather than "functions should be called with () - that is clearly not a function - can not compute!!", the question is "can this code be understood even ignoring unfamiliar syntax?"
Most developers familiar with languages like Java, JavaScript, Python, Perl, C# or Ruby could correctly guess what "split" refers to and that it is unlikely to be used as a variable name. Further, if a developer is familiar with the idea of method calls, it will be inferred that "whatever.split.whatever" is a sequence of method calls.. much as if it were, in some fictional languages, whatever->split->whatever or whatever:split:whatever. Similarly, a(b(c())) and a[b[c[]]] could both be easily inferred to be a set of nested function calls.
"import sys" is not exactly clear either, but again, the meaning of this could be accurately inferred by most developers.
print '\n'.join(filter(line.split(), lambda word: word.endswith('ing'))
I want to join the array's elements together, not join the string. Joining together an array's elements by calling a method on the delimiter is no less raving than any Ruby I may cook up. Your example is both longer and has more syntax.Yes, puts joins newlines.
#include <iostream>
#include <string>
int main()
{
std::string s;
while (std::cin >> s) {
if (s.size() >= 3 && s.match(s.size()-3, 3, "ing") == 0) {
std::cout << s << '\n';
}
}
} import Data.List
main = putStr . unlines . filter ("ing" `isSuffixOf`) . words =<< getContents
(To be read from right to left.)A Forth example would be interesting.
: print-ing ( filepath encoding -- )
[ [
" " split
[ dup "ing" tail? [ print ] [ drop ] if ] each ]
each-line ]
with-file-reader ;
It's kind of ugly because it's designed to print out the values without leaving anything on the stack. It can be prettier if it produces an array of matching words : collect-ing ( filepath encoding -- seq ) { } -rot
[ [
" " split
[ "ing" tail? ] filter append ]
each-line ]
with-file-reader ;
and then just prints that out : print-ing ( filepath encoding -- )
collect-ing [ print ] each ; foreach my $line (<STDIN>) {
foreach my $word (split($line)) {
print "$word\n" if $word =~ /ing$/;
}
}
The programmer made a conscious decision to read STDIN implicitly in a while loop and then call split with implicit usage of $_, then went on to complain that too much was hidden away. I could make the python example equally obtuse: import sys, re
for a in sys.fopen(sys.stdin.fileno(),'r',1):
for b in re.split(a,r''):
if re.search(b,r'.*ing$')
print b
Compressed forms: # PERL
map { print "$_\n" } grep { /ing$/ } map { split($_) } <STDIN>;
# PYTHON
import sys
for a in [ a for a in [ b for a in sys.stdin for b in a.split() ] if a.endswith('ing') ]: print a
# PYTHON
import sys
print "\n".join(filter(lambda x: x.endswith('ing'), [ var2 for var1 in sys.stdin for var2 in var1.split() ]))An entirely different angle on all of this is that NLTK is aimed at novice programmers (linguists) as well as expert programmers (CS students). Perhaps readability is a different concept for these two user groups...
They are also inconsistent with their coding style. In the while() loop, they make use of $_ implicitly (even in the split statement in the foreach loop), but in the foreach loop, they don't use $_, but instead define $word. Then they go off on split being 'difficult to guess what it represents.'
> Having used Perl ourselves in research and teaching
> since the 1980s, we have found that Perl programs of
> any size are inordinately difficult to maintain and
> re-use.
Says more about the programmer than the language. You can write maintainable code in any language. "Too many choices" doesn't cause code to be unmaintainable. Lack of discipline in your programming practices does (as well as lack of documentation). Saying that a language is better in this respect is just to say that "X language took the ability to choose away from me so that I'm forced to do things a certain way, whether I like it or not."In the following example I'm using the portable http://www.cliki.net/SPLIT-SEQUENCE to split the words:
(defun endswith (string suffix)
(let ((size (- (length string) (length suffix))))
(unless (minusp size)
(equalp (subseq string size) suffix))))
(loop for line = (read-line *standard-input* nil) while line do
(loop for word in (split-sequence #\Space line) do
(if (endswith word "ing")
(write-line word))))
It's still wordier than python, though.Reading this hurts.
isalnum(c)
looks much better than (c >= '0' && c <= '9') || (c >= 'a' && c <= 'z') || (c >= 'A' && c <= 'Z') if word.endswith('ing')
looks much better than if word[-3] == 'i' and word[-2] == 'n' and word[-1] == 'g'
clearly the authors of that article weren't proficient in the C standard library :) File standardInput readToEnd split select(endsWithSeq("ing")) foreach(println)
(Given that I don't actually use Io a lot, there might be shorter ways to do this.)Anyway, as usual, code samples prove very little. =)
if word.endswith('ing')
of course, it's great standard library design to have a string method called endswith() rather than making people use a regexp ending in '$', since finding suffixes is a common operation. but such a simple operation is hardly indicative of hardcore NLP (which would mostly be hidden in special-purpose library code anyways) for line in io.lines() do
for word in line:gfind("[^%s]+") do
if word:find("ing$") then
print( word )
end
end
endThe book is available at the same site: http://nltk.googlecode.com/svn/trunk/doc/book/book.html
really nice work on both the book and the NLTK library.
Console.ReadLine().Split().Where( word => word.EndsWith("ing"))
.ForEach( word => Console.WriteLine(word));