HNHacker News
TopNewBestAskShowJobs

slashcom

367 karma · joined October 18, 2007

submissionscomments
slashcom··on Meet the algorithm that can learn “everything about anything”
FWIW, the paper is much more modest and honest than the gigaom article. It's certainly a hot topic and getting wide coverage, but most scientists are being fairly conservative about their claims.*

* Grant applications and future work sections excluded :)

slashcom··on The Reformation: Can Social Scientists Save Themselves?
You mean to say, granting more wide recognition for the collection itself, rather than the analysis stemming from collection?

Possibly, I don't know. Mostly this is just my understanding from talking to peers in the social sciences and asking why they don't have open access like my field.

I think the problem there though, is knowing what questions to ask (what data to collect) is the hardest part, but fits much more in line with the analysis aspect (ie forming a hypothesis) than collection (conducting a designed experiment). So an ideal researcher would show good work in both areas, not just one.

slashcom··on The Reformation: Can Social Scientists Save Themselves?
My understanding of why this isn't often the case in the Social Sciences is that data is usually extremely expensive to collect. The researchers usually spend many, many hours writing proposals, revising it, resubmitting it, until they finally have the funding to collect the data their interested in. Then they actually have to get people into a laboratory, or send someone out to do field work (sometimes halfway across the world), or any number of expensive procedures.

And then, after data is collected, they begin analyzing it and seeing what are the interesting patterns in the data. And this takes a lot of time, and often you just want to show one thing at a time. In the current environment of publish-or-perish, and grant providers asking what you've done with their money, you might want to take that data you worked so hard to obtain, and milk quite a few publications out of it. If you openly release the data, then others can easily find interesting patterns in your data before you can, beating you to publication, and diminishing your track record for future grants, tenure, etc.

So your point is extremely valid. But the counter point of why open data isn't feasible is, if a researcher does all the hard legwork necessary to get to the interesting analysis stage, why shouldn't they reap the rewards of their hard work (namely, publications and recognition).

A few things have been proposed: verification projects (one particular other research group gains access to your data and verifies your results, but is not allowed to use the data otherwise), and grace periods (you have to release your data 3 years after a study or something, giving you time to milk out publications, but still allowing for general verification later).

Generally speaking, true scientific verification should also involve the complete recollection of the data. But when data collection is extremely expensive, this is typically infeasible.

[I work in a field that's almost all open access, and research data is almost always available upon request. So I agree with you. But data is typically incredibly cheap in my field, so the problem above doesn't really apply.]

slashcom··on A Face Recognition Algorithm That Outperforms Humans?
If we're going to go that route, linking to the actual paper might be even more appropriate. Even the blog article is only a second hand source.

http://arxiv.org/pdf/1404.3840v1.pdf

If you read it carefully, there's a caveat about how this particular dataset (recognition of /cropped/ pictures of /unfamiliar/ persons) has relatively low human accuracy (97.5% as opposed to 99.2%), because humans also use other features. That is, there's more work to be done, and facial recognition isn't a "solved" problem yet.

All the same, very impressive work. Congratulations to the authors on achieving such an important milestone.

slashcom··on Valve open sources Mesa fork from SteamOS
What is Mesa here? There's no readme on the github.
slashcom··on Computing 10,000X more efficiently (2010) [pdf]
Shit HN Says: the adjunct professor at CMU who's designing floating point hardware for processors is "totally stupid."

I'm guessing he intentionally underspecified for his audience, and if asked, could tell you in excruciating detail what 1% error means.

But skimming the patent [1], it seems that there's sometimes an error in a mathematical operation, measured as the computed value minus the expected value. The result doesn't seem surprising if you read how he defined addition in his hardware.

[1]: http://www.google.com/patents/US8150902

slashcom··on Execute SQL against structured text like CSV or TSV
Even then, it should scale reasonable well if one added a repl, and kept the database in memory. Probably would be fine then up until at least a couple gigabytes, depending on how well sqlite handles in-memory tables (probably very well).
slashcom··on Artificial intelligence: Our final invention?
With respect to 4: probably the funniest way I've ever seen vacation homes ever painted.

Presumably it goes Top AI -> Successful and wealthy -> has vacation home in quiet, rural area.

The author seems to paint it as Top AI -> has a rogue AI bunker

slashcom··on News organizations respond to Fed lockup questions
I'll throw this out there: what if the clock in Chicago was just 5ms slower than the one in NYC. The logs would show the transactions happened at the same time, but they didn't.

The NYSE mandates that business clocks never drift more than 1s from the atomic clock [1]. What is the resolution guaranteed between the two clocks being compared? After all, it's impossible to guarantee perfect synchronization of two clocks at any distance (bounded by the speed of light and the drift rate of the clocks) [2].

The only retort I can think of is why do the other players react at the proper time.

[1] http://www.nyse.com/nysenotices/nyse/rule-changes/detail;jse...

[2] Cristian, F. (1989), "Probabilistic clock synchronization", Distributed Computing (Springer) 3 (3): 146–158

slashcom··on How one man turns annoying cold calls into cash
Well he only gets 7p according to the article, so more like $6.50/hr. AFAIK, receiving calls in the UK is generally free (as in, it doesn't count as minutes as it does in the US), and the calling party foots the bill. He's essentially made his line a toll number.
slashcom··on I'm Beating The NSA To The Punch By Spying On Myself
Another problem is he doesn't include any sort of meaningful baseline. Maybe 67% of the people he talks to are male. In which case, he's not predicting anything. Just guessing randomly biased by his own calling habits.

[EDIT: I had some other criticisms that were incorrect and overly harsh anyway. I've removed them.]

slashcom··on Show HN: Global HTTP Latency Test
I like the idea greatly, but it appears to be broken at the moment. Perhaps the HN load is too high?
slashcom··on TL;DR — Faster News
As an NLP researcher, this is interesting as a sort of summarization data set.

The thing is though, summarizing news articles is best done by just reading the first paragraph of the article. News articles are intentionally written this way, and it's a very difficult baseline to beat in automatic summarization.

Still nice site though.

slashcom··on BlaBlaMeter detects how much bullshit is in your text
It should also be noted that on about 400 short texts (~300 words each), it did not correlate with the Flesh-Kinaid readability measure at all. So it's not measuring something like average word length or syllable counts.

But it's not QUITE a true lexicon, as it handles Out-Of-Vocabulary words quite strangely. If you use as input text:

"PR-Experts, politicians, ad writers or scientists need to be strong here! BlaBlaMeter unmasks without mercy how much bullshit hides in any text. A useful tool for everyone involved in writing! Simply copy your text into the white field and check your writing style. It works with english text up to 15.000 characters (overhead will be cut off). For a meaningful result we recommend a minimum length of 5 sentences."

Then you get 0.16. If you replace the last word 'sentences' with 'strategy' you go up to 0.44. However, if you change the last word to 'sentstrategyences' you get 0.47. Try it: you can basically insert 'strategy' inside ANY word and really raise your score. Actually, if you just insert "strateg" anywhere inside the text, it goes up massively.

So I actually think it's just doing string search counts over a lexicon.

slashcom··on BlaBlaMeter detects how much bullshit is in your text
If I were to make the software, the corpus of PR, licenses, etc. would be the way I go. But "they did it statistically" doesn't answer the question "what is the model?" There are many different statistical models one could use. My other post has a few things we've figured out.

But I'm starting to think a rule-based lexicon isn't out of the question, given these >1 scores on some texts.

slashcom··on BlaBlaMeter detects how much bullshit is in your text
Here's a few things I've gleamed from experimenting with it:

- It uses a unigram language model. You can take the same text, randomly permute the words, and you get the same score. This means it also can't be using things like POS tagging, phrases, etc.

- It normalizes words by making all letters lowercase. The exact same text in all upper case has the same score.

- The score is eventually normalized by the length of the text. The same text copied multiple times gets the same score.

- It does not form a valid probability distribution, as someone's managed to get some 1.16's. This makes me believe it's not a Naive Bayes classifier giving you the P(Bullshit|Text). Though this is what I originally thought it would be.

slashcom··on Part 2: I analyzed the chords of 1300 popular songs for patterns.
Please train a Conditional Random Field on this data. (A hidden markov model would also be interesting, but risks going out of key more easily).
slashcom··on What makes one appear smarter and more sociable?
Indeed. While the results are very interesting, and some of them are clearly very strong differences, I'd like to see some t-tests.
slashcom··on Silly Python riddle
Not a dumb question at all.

Short answer: it throws a syntax error otherwise.

Long answer: I don't know.

slashcom··on WhiteyPaint Turns Walls Into Whiteboards Without Cramping Your Wallpaper’s Style
Sweet idea, sweet product it seems. But that ad did not feel tasteful.
slashcom··on Facebook To Build New Server Farm Near Arctic Circle
Facebook's smart for taking advantage of the cold and such, but there are a lot of factors that go into choosing a location. Tax incentives, proximity to users, and a million other things I'm not aware of.

And while they will probably save massively on the electricity bill and carbon emissions for this, it's not like energy is prohibitively expensive here in Texas...

slashcom··on Siri says some weird things
I would think so, but if you put a number of these into Wolfram Alpha, you usually just get simple definitions rather than Siri's clever responses.
slashcom··on The Programmer Salary Taboo
$x/hour * 40 hr/wk * 50 wk/yr = 2n1000

The general rule is always "times 2, add 3 zeros" to do it in your head.

slashcom··on M.C. Escher: More Mathematics Than Meets the Eye
Obligatory reference to Gödel, Escher and Bach.

(For the record, I do understand this article explores a different mathematical aspect of Escher than Hofstadter)

slashcom··on I took "abusing the HTML5 History" to the next level.
Took me a while to come back here to comment. :)

Very fun, very creative. Abuse is most definitely the correct word.

slashcom··on The polynomial algorithm for 3-SAT problem (or P=NP)
Zero other (english) publications by the author in the Cornell archive and only four references within. Not an indicator as to whether the paper is correct (I haven't read it), but that's "smelly".
slashcom··on Ask HN: Who/What will do for Prolog what Clojure has done for Lisp?
Prolog has its limitations, but it still stands out as one of the oldest, influential and important declarative programming languages.

IMHO, it's hard to say what is the natural successor to Prolog. For all that it shares with Lisp, it definitely has never shared even a fraction of the popularity. They have two entirely different approaches and purposes though, so this isn't terribly surprising. Even if someone does create a new Prolog ala Clojure, it would probably still remain in academia.

I think a better question to ask is what direction is declarative programming headed in? Prolog is just one language in this field, as are some of the Prolog-alternatives listed by the OP. Remembering that SQL is also a declarative programming language, I believe that declarative languages are far from dead; it's a common paradigm, just not one we hear a lot of buzz about.

So all this said, it may be that Prolog doesn't necessarily need a successor. It does it's job well, but logic-based declarative languages are inherently too specialized to expect anything causing a surge of popularity.

One interesting variant of declarative programming is called Answer Set Programming. (http://en.wikipedia.org/wiki/Answer_set_programming). It's particularly good at modeling and solving NP-hard search problems, usually has Prolog-esque syntax, but does well on programs where Prolog would infinite loop (e.g. p :- not q. q :- not p.). As a disclaimer, I'm just beginning research in ASP. :)

slashcom··on I’ve created an online ebook creator called eBookBurn.com
You should put screenshots/demos on the home/sale page.
slashcom··on Experiment HN: Crowdsourcing Tech Stock Predictions
You should post aggregate values in a few days as well.

I hope you're also storing timestamps with the submissions. Just guessing, but I'd bet any predictions made on 04/30 will correlate much better than those made today.

slashcom··on Two Kinects at once is now possible
That was the worry and it is indeed the case. However, as demonstrated in the video, the interference wasn't as bad as expected.
← PreviousPage 3 of 4Next →