HNHacker News
TopNewBestAskShowJobs

ddod

460 karma · joined July 16, 2012

submissionscomments
ddod··on Show HN: Showing HN
Thanks for the feedback! It's actually creating its own database so as time goes on, it will show everything since it's been running. I didn't want to scrape through the backlog as I've already gotten banned a couple times while making this (prompting pg's recent unban post) and I didn't want to push my luck.

The upvotes thing is tricky because it's very time sensitive. I thought about it for a bit, but besides the difficulty in judging when to update the points, it also goes against the goal that all Show HNs are given the spotlight even if they didn't do so well.

ddod··on How to get your IP unbanned on HN
Thanks Paul! I'm reluctant to try this in conjunction with developing any HN scrapers since I'm not sure what set it off in the first place and your language suggests it will only unban the IP once (I will, however, make sure the CMU IP I was using gets unbanned). It would be helpful to know what, precisely, that hair trigger is so we can make sure to avoid it.
ddod··on Ask HN: personal hosting
I can't know what you mean by minimal load, but I'm guessing you could get by with 128mb RAM to run a few personal sorts of sites, but I've seen $7/m for 1-2gb RAM. Be sure to read comments to make sure the company is decent, as most are run by just a single dude who could be flaky.
ddod··on Ask HN: personal hosting
I'd assume most HN users would spring for a VPS, and if you're looking for a good price, check out lowendbox. You can easily get something for $12 a year.
ddod··on Introducing the 4chan API
Could someone explain to me how this could be leveraged (or if it could be) to gather a sort of stream of messages, a la the Twitter streaming API or reddit.com/r/all/comments.json?

I'd be interested in doing some language statistics and comparing them to the aforementioned networks.

ddod··on Anonymous allegedly hacked Romney tax records
There's no actual strict definition of terrorism; just a bunch of disparate attempts at one by various government agencies. The reason this is, is because it's so painfully, obviously subjective. If we use the definition provided by drivebyacct2, the use of the term itself is the best example of its definition. Most accounts of terrorism are merely just criminal acts that a certain party wants to paint as extranormative. It's interesting to bring up the question of whether this hack could be considered terrorism, as I'm sure if it's real, some Republicans will have interest in calling it that. That's the nature of the illusive concept of "terrorism". Sorry for the off topic discussion!
ddod··on Show HN: Automated writing analysis of Twitter and Reddit feeds
I'm working on the assumption of statistical significance of people using "yea" pronounced "yeah" vs. the very few people who might have reason to take to Twitter to use "yea" in a voting or biblical context.
ddod··on Show HN: Automated writing analysis of Twitter and Reddit feeds
Great points. I went into this knowing full-well that to many, many people, these misspellings were not considered as such. As the graph suggests, though, the majority of people still use the "older" spellings of "yeah" and "anyway". There are two ways of spelling each one (with "yea" and "anyways" considered to be the more nascent), and I'm approaching this from the standpoint of why the spellings are changing. Is the language evolving (suggesting improvement) or devolving (suggesting perpetuated errors). We could discuss linguistics, but what indicators are there from the data showing Reddit has a lower percentage of yeas and anywayss compared to Twitter? Is it simply the medium? How does that data compare to your anecdotal classification of the respective aggregate intelligence level of both sites' users? What do you think the data would say about HN vs. Youtube on the frequency of these (as you would consider) both valid spellings? This all boils down to me labeling the graph as common misspellings, but ignoring that, you should be able to make of the graphs what you will, as there's no associated hypothesis or conclusion.
ddod··on Show HN: Automated writing analysis of Twitter and Reddit feeds
There's no link between spoken and written English in the context of yea:yeah, as they're pronounced the same but either written correctly or incorrectly.

As for anyways, I think your classification there will depend on who you speak to (which is the point of the analysis).

ddod··on Show HN: Automated writing analysis of Twitter and Reddit feeds
Those might work for Twitter, but there may be some lurking variables on Reddit. While they never bother to catch yea and anyways, Redditers do have a tendency to pile on about the few grammatical/spelling errors they can spot. For example, if a post title says "could of", there's likely going to be multiple comments discussing the error and thus tripping my hooks.

Edit: As for the axes, I made some stylistic decisions that I wouldn't dare to make were this being submitted to the Annual New Media Socio-Linguistics Journal.

ddod··on Show HN: Automated writing analysis of Twitter and Reddit feeds
That's why I'm not relying on any single metric, as well as framing it in the context of time. If Justin Beiber asks his fans to vote yea or nay on something, it might mess up the yea:yeah metric, but it should eventually normalize, and while it isn't normalized, the anyways:anyway metric should still be fine (Unless JBeibs is asking his fans to vote yea or nay on naming his new album "Anyways").
ddod··on Show HN: Automated writing analysis of Twitter and Reddit feeds
Sample size is the anyway|anyways|yeah|yea instances. I get those as they get posted, so that's why it fluctuates with time. Since they're so commonly used, it should incidentally give you an idea of all of Twitter's load. I'm also grabbing some other words that I haven't implemented on the clientside yet, but I don't include them in the sample size.

HN comments tend to be more varied and sparse, so I think you're right that measuring yea:yeah, etc. wouldn't be too enlightening. That said, I've noticed a defined qualitative shift in HN comments over the past few months, and I'd like to develop ways of measuring that before they reach Reddit/Twitter levels.

As for your last point, I could track individuals but a relational comparison based off of /all/ data would be pretty difficult due to the number of comments vs. the few number of any individual's comments. Also, ARI isn't a great metric (hence me putting it in a tiny graph) because it measures chars instead of syllables. For example, "FFFFFUUUUUUUUU" has the same score as "constructivism".

ddod··on Show HN: Automated writing analysis of Twitter and Reddit feeds
I chose those two to start things off for two reasons: The first is that they stand out to me as pretty basic misspellings that don't appear in any sort of popular literature, so it speaks to someone's reading experience. You can take that for whatever importance it is to you, personally. The second reason is their frequency. If the unit of measurement was larger than a half-hour or hour, I could measure a bunch of other things (which I might still do). As it stands, anyway|anyways|yea|yeah occur 25k-70k times per 30 minutes.
ddod··on Show HN: Automated writing analysis of Twitter and Reddit feeds
Thanks. It compares misspellings to correct spelling counterparts (anyways:anyway), so that, in itself, should account for sample size changes in Twitter. The Reddit sample size (should) stay constant, as I'm grabbing /all/comments every 15 seconds, which from everything I've seen is about 3 times longer than it takes for it to turn over.

I have been thinking about subreddits and hashtags, but I'm not sure what would be interesting to people and still have enough statistical significance to update every half-hour/hour. If I were going to release this on Reddit, I'd have probably arranged it as a comparison between all the major subreddits. As for hashtags, their impermanence makes them a difficult measure.

I don't have the source up yet, but it's pretty simple. I just get the json output of /all/comments and get the Twitter streams for the words I'm checking with ntwitter and then do the basic math to analyze it every half hour. I render the page myself with the backlog of values and then send it updated info as I process it, so if you leave the page open, it should keep fresh.

ddod··on Show HN: Automated writing analysis of Twitter and Reddit feeds
I made this over the last two days in Node.js. The analysis is still pretty simple, and I'd like to expand it over time. It updates every 30 minutes, and you can already see that there's some significant shifts in literacy between the daytime and very early morning.

I had originally started the project to grab Flesch-Kincaid grade levels, but it was taking me 45 minutes to analyze about 20 pages of comments, so I switched it out for Automated Readability Index.

I really wanted to include HN, as I feel it's in a period of social shift, but it wouldn't have the same sort of statistical significance due to its size (it would most likely just be meaningless jarring jagged lines across the graphs). If someone can think of a good way to include HN, I'm all ears. I'd also be interested in other methods of writing analysis that I could include.

ddod··on Show HN: Tiny weekend project from scraped Wikipedia data: whodiedhere.com
Instead of learning SPARQL, I wrote my own script to parse from their format into my own. I should probably devote some time to learning SPARQL, but dbpedia's documentation was really confusing for me and I was basically learning it all just looking through the actual db files.
ddod··on Show HN: Tiny weekend project from scraped Wikipedia data: whodiedhere.com
Their stuff really isn't very user friendly (or at least wasn't for me) so there's a lot of room for improvement. I'm sure if you made a simple (and useful) Wikipedia API, you would be loved by all.
ddod··on Show HN: Tiny weekend project from scraped Wikipedia data: whodiedhere.com
Thanks! https://github.com/benwasser/whodiedhere
ddod··on Show HN: Tiny weekend project from scraped Wikipedia data: whodiedhere.com
I just put it up on Github now: https://github.com/benwasser/whodiedhere

After I built this I played around with some other information in dbpedia to see if I could get anything juicy out of it. No luck so far, but I'm sure there's a lot to uncover.

ddod··on Show HN: Tiny weekend project from scraped Wikipedia data: whodiedhere.com
I had to use a couple databases from dbpedia to get all the information (the death location syntax wasn't uniform) to make my own more manageable database.

As for the server, it's just Node.js and Socket.IO

ddod··on Show HN: Tiny weekend project from scraped Wikipedia data: whodiedhere.com
Yeah, I wanted to keep things simple and let people search by name, too, even if it's a little outside the scope of the site. I'll run the issue by the board of directors.
ddod··on Show HN: Co-optimal - A Steam matchmaking app made in Node.js and Socket.io
Those recent articles about finding friends in adulthood gave me the idea for this. Let me know if you run into any bugs or have any other feedback. I'm still a bit of a novice at Node, so I'm sure I've done a few things wrong.
← PreviousPage 4 of 4