Show HN: Neural network that impersonates writers
github.com
github.com
The output seems to me around the level a Markov chain might produce. Karpathy's RNN code produces much, much better results[1].
I wonder if manually extracting features and training the RNN on that is a mistake? RNN's tend to work well on text because they encode understanding of the parse tree themselves.
from sklearn.neighbors import KNeighborsClassifier
# Create a sperate neural network for each identifier
for index in range(0, len(NaturalLanguageObject._Identifiers)):
nn = KNeighborsClassifier()
self._Networks.append(nn)Wow.
There's some good reasons to think this approach won't work at all. If I understand it correctly I think it is attempting to predict part of speech using previously observed values.
That's an interesting idea, and might be somewhat valuable as a feature to use in a text generator, but on its own won't be enough to ever generate sentences that make sense (because some specific sequences just don't make sense).
He mentions that he used sklearn's Neural Network libraries in his blog post, but sklearn doesn't have any aside from RBM.
it must be a robust tool.
What the hell.
For comparison with Andrej Karpathy's RNN code (http://karpathy.github.io/2015/05/21/rnn-effectiveness/) training on the "HarryPotter(xxlarge).txt" (76K) file using the default hyperparameters and a batch size of 25 gets me:
> But Atfa the loom proset! No contarin — mibll,’s just pucking to live
> note left them hard and fitther, clooked of course little happered to
> trige on the fistpened. Their knew Harry mear from the shind-beas
> eveided, at Uncle Vernon’s thepped to spept were pelled and beadn
> Harry, distine dy use. Harry had in a amalout, into the fish sfary door.
The difference here is tokenizing on words vs letters: the RNN code is trying to learn the structure of English from completely zero whereas the code here gets to work with well-formed words from the beginning. But otherwise, the results in the linked post are about as silly semantically: > Input: "Harry don't look"
> Output: "Harry don't look , incredibly that a year for been parents in .
> followers , Harry , and Potter was been curse . Harry was up a year ,
> Harry was been curse "
EDIT: Updated the RNN output text. Was sampling from a checkpoint file for a different input corpus. Got confused by the long similar-looking filenames. Doesn't change the overall point though.It is learning phrases like so that you need the and that's the one from scratch. This version looks better at first glance because it is using correctly spelled words, but it repeatedly makes syntactic errors like for been/was been.
As I've posted here before, people have been training character n-gram models and getting language modeling performances comparable to those from word-based models---without using neural networks---for at least a decade. That it works with RNNs is no surprise because it worked just fine with the much more constrained predecessor technology.
Deep RNNs are simple and produce good results for a huge, diverse range of problems with no new domain information. As Andrej Karpathy wrote:
> Sometimes the ratio of how simple your model is to the quality of the results you get out of it blows past your expectations, and this was one of those times.
N-grams don't have nearly the power (eg longer-than-N-range structure like grammar) and don't generalize nearly as well, making them a lot less surprising.
Anyway, there are countries that don't have a fair use policy. So in this countries your repository could not legally be used.
IMHO, this is an unnecessary use of copyrighted material when there are thousands of equally well suited texts that have fallen out of copyright.
This person has no idea what they're talking about. sklearn has no neural network code whatsoever.
EDIT: this feels like a testament to sklearn's greatness, honestly.
I'd be even more interested if it could be turned into a sublime text plugin that highlights words / phrases that deviate most strongly from the house style.