HNHacker News
TopNewBestAskShowJobs

ltjohnson

61 karma · joined December 14, 2009

submissionscomments
ltjohnson··on Markov Chains Explained
This is correct. Sadly, the first paragraph of the article contains some glaring errors.

"For Markov chains to be effective the current state has to be dependent on the previous state in some way;"

This is trivially untrue. A sequence of independently and identically distributed (iid) random variables is a Markov chain. An iid sequence is clearly effective at many things (e.g. Monte Carlo integration).

"Not every process has the Markov Property, such as the Lottery, this weeks winning numbers have no dependence to the previous weeks winning numbers." As lambdaphage pointed out, the Lottery does have the Markov property.

ltjohnson··on Normal vs. Fat-tailed Distributions
Most statisticians have their own definition of fat-tailed, I like the definition you linked to (polynomial decay), but it is not universally accepted/known. These fat-tailed distributions will have a variance as long as they decay faster than x^(-3).
ltjohnson··on Starting a Bike Shop
I also used to work for Erik's (dinkytown location).

That was the gross margin on bikes, which I'm pretty sure means (sale price - purchase price)/(purchase price). Correct me if I'm wrong about the formula.

Not included in this is 36% number is the shops cost in selling the bike, which includes,

* Assembling the bike (unpacking, assembling, getting rid of packing materials, which are substantial).

* Tuning and Selling the bike.

* Post sale service.

At least for Erik's, a big part of the post sale service was to get people back into the shop and try to sell them accessories. I don't mean un-necessary parts or work, but things like clothes, bike computers, etc..

For many shops I suspect there is not much profit per bike left after these other expenses.

Erik's stream-lined much of this process, for example most bikes were assembled at the central warehouse by a dedicated crew.

ltjohnson··on Paul Graham spills: Why some companies get his cash and others don't
I sort of addressed this in my above comment. Success in academic research, like success in startups, is more dependent on determination than intelligence.
ltjohnson··on Paul Graham spills: Why some companies get his cash and others don't
I'm finishing up my PhD this year (haven't set a defense day yet, but I'm aiming for June at this point), and finishing a dissertation requires piles of determination.

What I came to realize is that you can't take what they say too seriously especially when you're aspiring to create something out of the ordinary.

I can't comment on your personal experience with grad students (I use grad to mean both Masters and PhD students), but I can say that just like people everywhere else, there is a wide variety of interests and abilities. Just like with job applicants, it's easy to see a sample biased by the "bad" ones. The sampling problem is even worse for grad students, the ones lacking in drive, creativity, or intelligence tend to stick around their programs longer [1]. The only other comment that I'll add is, depending on the program, there probably isn't much incentive to work with undergrads in the capacity you are talking about. So the determined/driven/focused grad students will not do it very much. You may have been steered towards projects because it was closer to what they work on, or would explore a question that they had personal interest in.

do 4 years as an undergrad, go straight to your Masters in the same school and then to your Doctorate in the same school

This is a warning sign. Staying at the same school for undergrad and grad school is called academic inbreeding. If it's a rarity, it's probably a special case that you can ignore (maybe that student was just particularly awesome). If it's common, the program may not be trying hard enough in admissions or perhaps something worse.

In my program, there are many students (including me) that did not go straight through undergrad to grad school.

tl;dr Just like everywhere else, the awesome people move through grad school quickly and the less awesome people tend to stick around. It's more likely that you were encountering the less awesome people.

[1] That's not to say that sticking around a PhD program for more than 5 years is an indication that a person lacks drive, creativity or intelligence. Life happens. Babies, sickness, bad breaks etc..

ltjohnson··on No, shut up. What statistical programming languages can learn from Dropbox.
Possibly. R is really great at doing "fancy" statistical analyses. It's very lousy at doing things like text manipulation. When I have a project that needs some text manipulation on the front end, I frequently use other tools (Python, vi, sed, ...) on the front end to beat text data into a nicer form for R. I couldn't say without knowing more about your project.
ltjohnson··on No, shut up. What statistical programming languages can learn from Dropbox.
No, but I can give some suggestions. It would help to know what you want to do.

First of all, you need to decide if you want a language reference, or an application guide, as R books fall into those two categories.

If you have a specific type of work in mind (bio-informatics, data mining, data visualization, ...) I'd say to find a book that focuses on that topic. I haven't looked in a while, but I haven't seen a general R book that I like, anything I suggest there would be guessing on my part.

There are plenty of good references on the web. I'd start by looking at the material available from the R web site:

R's core manuals [1] are typically correct and reasonable to use. The "Introduction to R" guide will get you up to speed fairly well if you already know another programming language. There is also the contributed documentation [2]. I haven't gone through these, so I can't say much about them, or promise that they are up-to-date. I suspect not, as R develops rapidly. The one reference I can recommend highly is "The R Inferno" by Patrick Burns [3]. This is not a starter guide, but something you read after one. It gives excellent advice on avoiding common pitfalls in R.

[1] http://cran.r-project.org/manuals.html

[2] http://cran.r-project.org/other-docs.html

[3] http://www.burns-stat.com/pages/Tutor/R_inferno.pdf

ltjohnson··on No, shut up. What statistical programming languages can learn from Dropbox.
I am a statistician that does both research and applied work.

I use R for three reasons: (1) It's Free Software; (2) It's a programming language; (3) Other statisticians use it so it's easier for me to collaborate.

There are the usual supporting arguments for (1). (2), I've only used SAS a little bit, and it was extremely unpleasant to use it for non-built-in stuff, which makes research harder for no good reason. For (3), I have nothing against Python but most other statisticians don't use it. If I want to share my work in R, it's easy (statisticians know how to install R packages). If I want to share my work in Python, I first have to teach [most] other statisticians how to use Python. There's nothing wrong with that, but why raise the start-up cost for them?

tl;dr I conjecture that most statisticians don't want what the author is suggesting. Also, there are plenty of companies that are trying to do what the author is asking for, but most of them seem to miss the desired sweet spot, or charge lots of money, or both. I haven't taken a survey of the available software in quite some time.

ltjohnson··on Appsumo reveals its A/B testing secret: only 1 out of 8 tests produce results
You're right, reading the paper that way, they are trying to detect a change in convergence rate from 5% to 5.25%, I was confused by the use of % in two separate contexts in the same sentence. That being said, I think this is not a good argument against A/B testing.

Fair enough, it would take a very large sample (122K is close enough) to detect a change from 5% to 5.25%. Being concerned about a change that small seems really silly unless 0.0025 * N visitors * revenue per user is a big enough number to be concerned with. I contend it won't be unless either

(1) N visitors is very large or

(2) revenue per user is very large.

If (1) is true, then testing on 122K users is not a big deal. If (2) is true you probably want to have a much more targeted approach, like someone doing sales.

ltjohnson··on Appsumo reveals its A/B testing secret: only 1 out of 8 tests produce results
122,000 seemed absurdly large to me so I just looked up the paper you reference [1]. It looks like there is a calculation error in the paper.

The formula they use is n = 16 σ^2 / Δ^2. σ is the standard deviation, and Δ is the size of the difference. Thus in this problem, Δ = 0.05 and their formula gives n = 16 * (0.05 * (1-0.05)) / 0.05^2 = 304. This is much more in line with what you get from using a 2 sample proportion test (with H_a: p_1 =/= p_2), ~440 in each group [2].

But maybe I misunderstand their formula.

[1] http://exp-platform.com/hippo_long.aspx

[2] http://statpages.org/proppowr.html

edit: fixed Greek letters and added final comment.

ltjohnson··on HackerBuddy
Just signed up.

But, my main areas of expertise (statistics, e.g. data analysis or mining, modeling or forecasting, designing and analyzing experiments, ...) don't mesh well with your categories. Have you thought about allowing users to search profiles, or some means of adding categories (sugestion box?)?

ltjohnson··on Snippets of code for things you do all the time in R
You gotta do what you gotta do, but your example probably doesn't do what you want in R.

  > system.time(vec1 <- rnorm(1000000))
     user  system elapsed 
    0.200   0.000   0.199 
My computer takes .2 seconds to generate the million random normals and assign them to vec1.

  > vec1 = rnorm(1000000)
  > system.time(vec1)
     user  system elapsed 
        0       0       0 
 
My computer only takes 0 seconds to evaluate the existing vec1 of length one million.
ltjohnson··on Snippets of code for things you do all the time in R
I've added some examples to my above comment. My example is perhaps a little closer to 'median(x <- 10)' than you were asking for, but I think it illustrates the point that using '=' will not always have the desired result in places where '<-' would.
ltjohnson··on Snippets of code for things you do all the time in R
Using '<-' is the general style of the R community. The accepted (but not universal) practice is to use '<-' for assignment and '=' in argument passing. This is a good idea because '<-' and '=' are not equivalent and '<-' is more likely the operator you want. From the help page [1] The operator ‘<-’ can be used anywhere, whereas the operator ‘=’ is only allowed at the top level (e.g., in the complete expression typed at the command prompt) or as one of the subexpressions in a braced list of expressions. That being said, some R community members still use '=' for assignment, but it's frowned upon. It certainly annoys me when I see it in R code.

Simple example to show when '=' won't work but '<-' will. Pilfered from "R Inferno" [2], which is an excellent reference for R idioms and best practices.

  system.time(vec1 <- rnorm(1000))
  mean(vec1)
  system.time(vec2 = rnorm(1000))
  mean(vec2)
And to demonstrate that you need to use '=' for argument passing:

  x <- c(3,4,NA)
  mean(x)
  mean(x, na.rm=TRUE)
  mean(x, na.rm<-TRUE)
[1] Type help('=') at an R prompt, or view it online: http://stat.ethz.ch/R-manual/R-patched/library/base/html/ass...

[2] http://www.burns-stat.com/pages/Tutor/R_inferno.pdf

edit: clarified statement.

edit: added examples.

ltjohnson··on Proofs By Contradiction And Other Dangers
Speaking as an almost not-amateur (finishing my PhD this year, and working on getting a faculty job), you're both right.

Logically, there's nothing wrong with proof by contradiction. As long the steps between "Assume P is true" and "therefore FALSE" are valid, you have a valid proof that "P is false".

The issue with proofs by contradiction is psychological, but it's important not to ignore such things when you are doing math. Proofs are written for people to read, and people make mistakes in reading, writing and constructing proofs. The longer and more complex a proof is, the more likely it is that there are mistakes. I can't speak for all fields, but in my field (statistics) proofs by contradiction are acceptable but not in style and constructive proofs are generally considered to be more informative.

ltjohnson··on Prof Gives Lecture to Prove He Knows Students Cheated; Over 200 Students Confess
I'm not interpreting your comment as flippant. :)

In terms of shortcuts, that would depend on what they were. Some are fine, others are not. E.g. 'borrowing' answers from other students on homework is a shortcut, but it's okay if the student understands the material in the end; cheating on an exam is not an okay shortcut.

"Surprisingly hard" is the name of the game. I mentioned that writing exams is hard to motivate why an instructor would use a test bank. At a research school (I'm at a research focused school), being an instructor is only a small part of a professors (or grad students) job. Using a test bank (from books/publishers or private ones) is considered to be fine, as it lets them focus more time on the things they "should" be doing. Many things that are considered okay in this context are abhorred in others, e.g. having grad students be primary instructors is okay here but would be taboo at a teaching focused school. Of course, most teaching focused schools don't have big grad programs.

Incidentally, I'm on track to graduate this year. I'm sending out academic job applications and I find myself more drawn to teaching schools than I thought I'd be.

ltjohnson··on Prof Gives Lecture to Prove He Knows Students Cheated; Over 200 Students Confess
I'm a 5th year PhD student who is teaching a large (80 student) section of a course, this is the 4th course I've taught. I've also taken plenty of exams as a student, and they are still fresh in mind.

I would want to know more information before I decided the students were cheating or not. The instructor refereed to an "exam room", and gave an hour range that the new exam could be taken. So the students are not all taking the exam at the same time, this makes it seem possible that the exam is online. If the exam is online, and the students can take it at home vs take it in a proctored room, that would change what would be cheating. If it were online at home (I don't think so from the video) then reviewing the test bank while taking the exam would be cheating. If not, then having seen a question before the exam may or may not be cheating, depending on HOW you saw the question.

If you did not acquire questions in an unethical way, then it's not cheating, it's just studying. As an instructor, I will sometimes put problems from the book onto my exam. If the students worked the problems before because they were studying hard, then good for them! I want my students to study, because it will help them learn. I also provide a sample exam with previous exam questions on it; I write most of my own questions and it's important for students to get used to my style. As a student, I had to take a written exam for my PhD. When I was studying for the exam I asked Professors for help, one of my Professors gave me some of his questions. I worked out every single question. He also submitted one of his existing questions to the exam and I recognized it when I was taking the exam. Cheating? No. I just got lucky (and worked my ass off).

If test questions are acquired by malicious means, or knowing that they are going to be on the exam, or are the test bank that is going to be used to make the exam. Then it is cheating. So if students knew that the questions came from a test bank, and downloaded the test bank (I'm sure it's on the web somewhere) to gain an advantage they cheated.

Finally, as an instructor. Writing a decent exam is surprisingly hard. My goal with an exam is two-fold, figure out how well the class as a whole is doing, and separate the students into their grade groups. The ideal exam has some problems that even the D students can answer (to separate them from the F's) and some problems (usually just 1 problem) that are a stretch for the A students. And a mix of medium problems for everyone. If you have too many easy problems, the grades will creep up and you won't separate students. If you have too many hard problems, the grades will creep down and you won't separate students. Writing an exam from scratch is very time consuming. I use my private test bank, and try to add 1 or 2 new questions to the bank when I'm writing each exam. I can understand (but don't agree with) an instructor pulling entirely from an existing bank to write an exam.

ltjohnson··on Rationing, errors and mammogram math
I'll start by saying that I'm not trolling.

Using conditional probability, or bayes theorem to calculate a conditional probabilitiy, does not make something bayesian analysis.

I haven't read the original report (it's on my list) but what is presented in this article isn't really anything except probability. It would be bayesian if they put a distribution on the proportion of people that have a "certain cancer."