HNHacker News
TopNewBestAskShowJobs

youngprogrammer

249 karma · joined January 17, 2015

submissionscomments
youngprogrammer··on Determining Gender of a Name with 80% Accuracy Using Only Three Features
> The SSN name data seems almost certainly flawed to a small degree...I.e. It's just hard to believe that there are dozens of boys named Jennifer (and yet, strangely, no boys named Sue!)...but we're talking about infinitesimal rounding errors. The vast, vast majority of names are 99% one way or the other...with a few exceptions such as Leslie...though you can mitigate that by using older years of the SSN database.

I agree that the SSN data has flaws, but I only took names with at least 20 people. But the classification is probably iffy, as some names are classified as male and female.

> So I have to strongly disagree with OP that 80% accuracy is something to be astounded by when it comes to gender classification...

I originally hypothesized, I could reach 90% accuracy, but I could only get up to 82% max. As stated in the blog, 80% is the accuracy of a mammogram detecting cancer in a 40-45 year old woman which is pretty good for 3 features!

> I wonder how much more using Soundex would add to the accuracy? Creating a trained name classifier would be a fun project in service of a tool that could gender classify how masculine or feminine a made-up name sounds like...which would be a slightly useful tool if you were a fantasy fiction writer, though I suppose if you were to be a successful writer, your ear would be trained well enough for he purpose to not delegate it to a computational tool.

This would be very interesting to see!

youngprogrammer··on Determining Gender of a Name with 80% Accuracy Using Only Three Features
For me, this was an exercise in practicing machine learning and I found it very interesting that you could get 80% accuracy with such few features. A look up table works very well, but if a name does not exist in the table, you could possibly use some kind of ML to guess the gender.
youngprogrammer··on Determining Gender of a Name with 80% Accuracy Using Only Three Features
80% is not very good for practical uses (the Gender Classification as a Service on the web most likely use a map of names to probability of gender), but I think it is very good for 3 features.
youngprogrammer··on A Simple AI Capable of Basic Reading Comprehension
> So my basic objection is that you're really calling this reading comprehension, but I don't think anything is actually being "understood"; just parsed. A better title would be as I suggested: simple AI correctly answers reading comprehension questions.

I would argue that my program can understand the relationship between different objects but I agree that it does not understand the meaning of the relationships.

> I couldn't answer a REAL reading-comprehension test about this: I just have no idea what it's REALLY talking about, I don't actually understand it. (Obviously on a syntactic level, it's not hard to parse.) I don't know what over-expression is, I don't know what a kinase is, I don't know what phosphorylation is. I don't understand the text. But if the questions are simple, perhaps I could answer some reading comprehension questions about this by parroting back quotations from it. Syntactically, there's nothing difficult here. I just don't understand it.

I would also argue that you are doing very basic reading comprehension here. You might not know the meaning of individual objects, but you understand that "what is wrong" is "the overlapping substrate specificities of many AGC kinases, it is likely that the over-expression of one member of this kinase subfamily will result in the phosphorylation of substrates that are normally phosphorylated by another AGC kinase". You might not know what that whole phrase means, but you understand that its related to "whats wrong" with "studying the physiological role of AGC kinases by overexpressing the active forms in cells". I agree that my program is unable to do full comprehension in not understanding the meaning of objects and relationships, but it can do very basic comprehension in understanding what the relationships are.

> So I think it's unfair to call sentence parsing real reading comprehension, even if sometimes reading comprehension tests fail to differentiate between the two. You can parse sentences perfectly while understanding nothing.

I did not really call it "real" reading comprehension, but "basic" reading comprehension. But I suppose "basic reading comprehension" is still a little of a stretch. I think the real question here is: how can you really determine if a program can understand something? What does understanding something really mean? It is difficult to define something like this and it seems we need some kind of "Turing test" for understanding.

> I did find your work very interesting, thank you.

Thanks!

youngprogrammer··on A Simple AI Capable of Basic Reading Comprehension
> The essential question is: Can it go from basic reading comprehension to advanced just by adding more rules. Is intelligence simply 10 million rules? If so, how do we go about creating new rules as language evolves? By hard-coding them, as in the example code?

I believe that it can go from basic reading comprehension to more advanced by adding many rules but of course manually adding them is not very feasible or scaleable.

> In my opinion, the fuzzy, statistical methods @davesullivan mentions have a better chance at generalizing (although they may well be augmented by rules-based AI).

I agree that a statistical model would be better since it will be able to handle more complexity. It would be much easier to train the rules from a dataset instead of hard coding all of them and it would be able to adapt to new rules as well. However, I could not find a good data set for the task I wanted.

youngprogrammer··on A Simple AI Capable of Basic Reading Comprehension
> It's a great demo project, and I love seeing these on HN.

Thanks!

> OP - Instead of linking directly to en/Stanford Parser etc, you should get together a list of dependencies people need to run your application. Usually as easy as 'pip install pattern' (for the 'ImportError: No module named en') which is `import pattern.en` :-)

I didn't actually pip install anything for my project, I just downloaded and extracted the Stanford Parser, and Nodebox Linguistics libraries. The setup should be in the readme. I'll try to see if I can find the pip dependencies and update the readme.

youngprogrammer··on A Simple AI Capable of Basic Reading Comprehension
> Bobby picked up the toy. Then he put down the toy.

When we read this sentence, our brains automatically augment additional information based on the verb. However, in this example, my program will fail to answer because my program does not augment any additional information but it can be extended to.

This can be implemented in our program by created a new property for each object called "location". If a verb is location based, we can set the location of the object based on what the verb describes. For example, "the toy"'s location could be "Bobby's hands" after the first sentence based on the verb phrase "pick up". So the program will understand where the toy is and be able to understand queries related to "where".

As you can imagine, implementing this would be very tedious since there are too many cases for all the verbs. My program may not be able to do advanced reading comprehension (reading between lines and augmenting information) but I argue that it can do simple reading comprehension, in that it can understand the relationship between objects. There is still a long way to go before my program is capable of more sophisticated reading comprehension, but in theory, I think my approach seems possible.

youngprogrammer··on A Simple AI Capable of Basic Reading Comprehension
> I was intrigued by the initial transcript, but disappointed that it was edited (lightly). For the question "Why did Mary cheer?" the screenshot showed that the simple AI literally answered "because IT be HER FIRST TIME WINNING". This is correct, but the author edited it into a correct sentence for us.

For some reason, the library I used to get the present tense of "was" is "be". I had a manually fix for this but accidentally removed it when I was cleaning up the code. Sorry if I disappointed you

> I think it is too much to call this "capable of basic reading comprehension". Surely, "simple sentence parser can answer reading comprehension questions" would be more correct?

I would say it is capable of basic reading comprehension because it is attempts to build relationships between different objects although the relationships are weak. When humans do reading comprehension, we do what my program does which is try to parse the sentence and understand the relationships. But brains are also able to augment a lot more information and thus be more flexible with understanding.

> Would you say that 14-line Perl program is capable of basic reading comprehension? I wouldn't!

I would say it is not because it does not understanding the relation between objects; it only understands that sentences starting with "why" should be answered with everything after the string "because". Also, in your program, if there are two instances of "because" in the source, your program will only choose the first one.

← PreviousPage 2 of 2