NLUlite: Natural language parser and database
nlulite.com
nlulite.com
I'm the developer of the Redshift NLP library ( http://github.com/syllog1sm/redshift ). Currently the software lacks documentation, but it offers a good speed/accuracy trade-off. Documentation and a good tokenizer are coming. You can read a tutorial for a simplified version of the algorithm on my blog: http://honnibal.wordpress.com
I suspect it might be expecting too much, but I'd love to integrate my browser history, people I've contacted, where I've been etc. in order to produce an easier way to search and find some webpage (e.g. when I remember some contextual information about the day or the place when I visited a page but am unable to find it by re-googling)
@unsane: This is of course not supposed to happen (we tested it on many different machines). Apologies. Please write at contact@nlulite.com and we'll look into the problem.
@garblegarble: The system is still under development. Please write at contact@nlulite.com to suggest features you would like to be present.
@Syllogism: 93.6% of accuracy is impressive. At this stage, however, we prefer to use proprietary algorithms. We feel we can reach similar accuracy for version 0.2.0 (out in January)
@CGamesPlay: The server is supposed to be installed in the $HOME directory. If you wish to use a different path, you can use the option -d <YOUR_NEW_PATH> when starting the server.
@Rhapso: You are right, the non-commercial download is somewhat byzantine. The problem with wget is that you don't get to sign a non-commercial agreement. Let us think about it for a few nights.
@toblender: We are working on that ;-)
Not to belittle the tremendous effort, but most projects I have seen are "English Language Parser"s.
Are there any actual generic language parsing projects out there?
That don't try to overfit to English but actually attempt to do a job of whatever quality in whatever language?
Like I'm a native English speaker, I can understand English say 100%, Japanese 80-90%, I can understand a bit of a few European languages and I can identify a bunch of other languages.
It would be wonderful if there were software with this design in mind.
Chalmers University has impressive results on this - http://www.grammaticalframework.org/
http://en.wikipedia.org/wiki/Horse (I don't like snakes)
File "/home/drace/dev/NLUlite/client_python/NLUlite.py", line 375, in add_url
parser.feed(page)
File "/usr/lib/python2.7/HTMLParser.py", line 114, in feed
self.goahead(0)
File "/usr/lib/python2.7/HTMLParser.py", line 158, in goahead
k = self.parse_starttag(i)
File "/usr/lib/python2.7/HTMLParser.py", line 305, in parse_starttag
attrvalue = self.unescape(attrvalue)
File "/usr/lib/python2.7/HTMLParser.py", line 472, in unescape
return re.sub(r"&(#?[xX]?(?:[0-9a-fA-F]+|\w{1,8}));", replaceEntities, s)
File "/usr/lib/python2.7/re.py", line 151, in sub
return _compile(pattern, flags).sub(repl, string, count)
UnicodeDecodeError: 'ascii' codec can't decode byte 0xe2 in position 8: ordinal not in range(128)It's also really slow at learning. I have a ton of everything, cores, memory etc and it takes minutes to process web pages. I guess you do say that on the website that the free version is slow.
[append] Turns out the server will silently do nothing if you do not extract the archive to $HOME.