HNHacker News
TopNewBestAskShowJobs

EvanMiller

1,676 karma · joined April 29, 2010

submissionscomments
EvanMiller··on A Hacker News for grad students?
I would also like to see such a community.

I've found it somewhat in the Julia community, which tends to lean heavily towards scientific computing and has a lot of Ph.D.s, grad students, and Ph.D. dropouts. We just had JuliaCon last week, and the talks tended to be first about a problem domain (statistical modeling, optimization, natural language processing, finite element methods...) and secondarily about Julia. I loved it.

If you're in the Chicago area, I'd encourage you to drop by our Julia Meetup group, which has a similar format. Previous talks have been about solving problems in climate modeling, machine learning, and molecular dynamics, and I'm trying to line up talks later this summer about machine learning in psychology and MLE modeling of longitudinal econometric data. (All using Julia, of course.) The group is small, but the quality of discussion is very high. Join the group and you'll get an email when the next meeting is added to the calendar:

http://www.meetup.com/JuliaChicago/

EvanMiller··on Startup advice: cold recruiting
I occasionally receive emails like these, and I agree with most of this advice. I will add a couple of personal turn-offs:

* Overtesting. I have received multiple emails from different people at the same company with the exact same subject line ("Hey Evan, let's chat"). I am sure this subject works better than the other ones that they tried, but a little variation would make me feel like I am more than a conversion goal.

* Being vague about the purpose of "chatting". If an engineer emails me and saying he or she "would love to hear more about X", where X is something I've done, it's not immediately apparent that they actually don't give a shit about X unless they can hire me. Don't be bashful, just say your company is hiring. At least when a recruiter (or founder) emails, I know what page we're on.

EvanMiller··on Founder Depression
Dig deeper, Sam.

Achievement-oriented people are given to depression both when they fail and when they succeed. If your identity is tied up in your work, then you feel bad about yourself when work isn't going well. That's obvious, and that's the message of this blog post. The implicit message is that you're depressed because you're not succeeding, so get your shit together and succeed and be happy like everyone else.

But then if you do succeed, you start to wonder, why did I just spend my youth in this masochistic, narcissistic path, and why the fuck am I not as happy as I was expecting, and is this really all there is in life. This is a classic "achiever in crisis." The problem is that you realize all along you've been doing things that OTHER people wanted -- that is, you've been doing things that make you valuable in society -- perfect summed up in the raison d'etre du jour, "making the world a better place." And nobody stopped you, because who can argue with making the world a better place? (Or being a doctor, or whatever.) But upon reflection, you quickly realize that this was in many ways easier than asking yourself what YOU wanted out of life. I.e. you've pushed aside your innate feelings and desires, whatever they may have been, and replaced them with the external motivation of achievement, under the rationale that you'd be able to "figure it out" after you had "made it".

Unfortunately achievers aren't really sure what they want "deep down" because achievement is inherently defined by society, and then after they've "made it" they freak out because they start to wonder if there even is a "deep down" or if they're just a highly educated donkey chasing a carrot.

If you talk to e.g. people who've gone through rigorous Ph.D. programs, you'll find a number of them were severely depressed after their defense. It was just kind of a let-down after such a long buildup, and then they started to wonder why they invested the entirety of their twenties into it and question whether that's really what they wanted their life to be. At least before the defense they could have something look forward to, and the various requirements provided a source of manic energy to propel the achiever forward.

Anyway I don't think the problem here is "not enough success," and I don't think the solution is having more coffee meetings. Founders need to take a hard look in the mirror and ask themselves why they're doing what they're doing and whether their depression is truly a function of their free cash flow or if there's a deeper dissonance between the founder's feelings and the expectations of society, i.e. the heroic mythology of the founder that Silicon Valley has been inculcating in susceptible teenagers for the last 20 years.

Just my 2c. I am not a founder just an observer and aspiring societal psychiatrist. If you want to learn more I highly recommend "The Wisdom of the Enneagram":

http://www.amazon.com/Wisdom-Enneagram-Psychological-Spiritu...

It looks a lot like astrological pseudoscientific trash but read it and see if things in it resonate with you.

Ok back to work.

EvanMiller··on A Formula for Bayesian A/B Testing
I'm pretty sure this formula is correct, but I haven't seen it published anywhere. John Cook has some veiled references to a closed-form solution when one of the parameters is an integer:

http://www.mdanderson.org/education-and-research/departments... [pdf]

But he doesn't really say what that closed form is, so I think his version must have been pretty hairy. (My version requires all four parameters to be integers, so I doubt we were talking about the same thing.)

Sadly I couldn't get the math to work out for producing a confidence interval on |p_B - p_A| so for now you're stuck with Monte Carlo for confidence bounds.

Thanks to prodding from Steven Noble over at Stripe, I'll have another formula up soon for asking the same question using count data instead of binary data. Stay tuned!

EvanMiller··on Bayesian A/B Testing
This is a good discussion, but the author is confused about what the "Low Base Rate Problem" is. It doesn't have anything to do with the null hypothesis being true most of the time -- the example the author gives is actually a second form of repeated significance testing, which could be addressed with a Bonferroni or Šidák correction.

The Low Base Rate Problem is when you have a binary outcome and one of the outcomes is rare (say, less than 1%). There is so little entropy in the information source that you have to acquire a heck of a lot of samples in order for the statistical test to have any power. The problem is not unique to frequentist statistics; it's a consequence of information theory and so it affects Bayesian statistics as well.

Nonetheless, I highly recommend examining Bayesian test techniques to avoid repeated significance testing (both within a single trial and across multiple trials). A side benefit is that when someone says "What's the probability that the new purple dragon logo outperforms the old one?", you can give them an answer without backpedaling and explaining null hypotheses, p-values, significance levels, and all that jazz.

The major drawback to Bayesian techniques is that it tends to be computationally expensive. For example, to evaluate the A/B test and answer the purple-dragon question with normal priors, you have to integrate a normal distribution in two directions, and there's not a clean analytic formula for that. That's why there's a jagged histogram in the blog post; it changes every time you hit "Calculate" because it's being integrated with Monte Carlo techniques, which take a lot of juice compared to (frequentist) analytic methods.

EvanMiller··on JuliaCon, first-ever conference for the Julia language
We'll definitely record the talks and post them online -- but whether we'll have live-streaming is TBD.
EvanMiller··on JuliaCon, first-ever conference for the Julia language
JuliaCon's been in the works for a while now. I'm helping organize it, feel free to leave any questions here and I'll try to give you a prompt response.
EvanMiller··on Show HN: Visualise the structure of a spreadsheet
Thanks for sharing your work. I love your approach to showing relationships inside a spreadsheet.

This is definitely a feature that Excel ought to have, which is both good news and bad news for your business. The good news is that you'll probably find willing buyers. The bad news is if you find too many buyers, Microsoft will probably implement something similar in their products, which will put your business at risk.

There are a few ways to survive in these situations. One is to keep moving with more products and add-ins. That's basically what Xobni had to do, but it will be exhausting, and I think at some point you'll run out of steam. Another way is to take out patents and hope for a licensing deal or buy-out from Microsoft.

A third survival strategy is to have a better product than what Microsoft will inevitably come out with. It's become cliche at this point, but high product quality requires, above all, focus. Microsoft might have more money than you, but if you're willing to focus, you'll actually have more time to think about the product than they'll have to think about their feature.

To be honest, I were you guys, I'd pick either the Cloud version or the Excel Add-in and go all in (i.e. dump the other one). Then make that version the absolute best it can be.

If you're going to focus on one version or the other, there are a lot of factors to consider, and you guys know the market and technology better than I do. One factor in favor of the Add-in is that you'll be able to charge a single purchase price up-front, which will give you more short-term revenues that will enable you to grow your business without taking (as much) outside investment. Subscription models are great for big companies and VC-funded startups, but they're actually a raw deal for cash-starved, self-funded companies.

Anyway, thanks again for sharing. I look forward to seeing where you take Slate!

EvanMiller··on The Mt. Gox alleged insolvency document has been posted
And this, children, is why markets are regulated.
EvanMiller··on Volkswagen's US workers vote against joining union
This has been a knock-down drag-out fight and a lot of fun to watch. VW hates the UAW, but they can't say as much because being seen as anti-union would hurt their image back in Germany. Thanks to NRLB rules, VW can't make any statements about rewarding the workers for not unionizing, but that hasn't stopped Senator Corker from saying he heard it on good authority that VW would bring production of a second car to the Chattanooga plant if the workers rejected the UAW. That was a dirty/brilliant move on Corker's part, which both advanced his anti-union agenda and set the expectation that VW will add a second production line to its factory, which VW has been considering since the Passat hasn't been selling as well as they initially planned.

Framing the unionization effort as a "works council" was ingenious marketing on UAW's part, but apparently wasn't enough to persuade the Tennessee good old boys to take the Rust Belt gambit. If you've ever heard stories about how GM & Co. run their factories you'll be amazed anything gets built at all.

EvanMiller··on An Interview with John Foreman, Chief Data Scientist at MailChimp
If you want to learn the nitty-gritty of machine learning and optimization, I highly recommend his book, Data Smart:

http://www.amazon.com/Data-Smart-Science-Transform-Informati...

It's one of the few books on the subject that doesn't get bogged down with mathematical notation, nor does it cheat with "And Then A Miracle Happens" library calls. The book is primarily aimed at business analysts, but programmers can get a lot out of it too.

EvanMiller··on Startups that have bootstrapped their way to profitability
Braintree is a Chicago company.

Adrian Holovaty (of Django fame) recently caused a stir here when he argued in front of a group of investors that Chicago is an ideal place for bootstrappers:

http://www.holovaty.com/writing/chicago-bootstrapping/

EvanMiller··on Setting OmniGraphSketcher Free
I've enjoyed watching the "arc" of GraphSketcher. I went to college with the original developer, who, upon being sufficiently annoyed with having to draw supply and demand curves by hand in Econ 101, wrote the code that eventually became OmniGraphSketcher. Classic itch-scratching if I've ever seen it.

Robin later managed to turn the program into his master's thesis, a conference paper that won "Best Paper" (http://dl.acm.org/citation.cfm?id=1518870), and an acquisition by the Omni Group, where he went on to develop the iPad version.

He later left Omni, which is part of the reason that they decided to discontinue and (later) open-source the product.

I'm look forward to seeing what people do with this code -- as well as what Robin might have up his sleeve next.

EvanMiller··on Why Wikipedia's A/B testing is wrong
I won't comment on the virtues of bandit algorithms, but the author grossly mischaracterizes the requirements of a multivariate A/B test. Just do a regular 50/50 split and run a regression on the outcome using the observable characteristics as explanatory variables. Then take that model and feed the observables back in to get an individualized prediction as to whether A or B will perform better. You can bin continuous variables or you can including higher-order terms to achieve a polynomial fit. The data requirements really aren't that large.

I believe the author's confusion stems from attempting to model every interaction term, e.g. if you think there will be some extra effect to being Male and using Windows and being Swedish that is more than the "Male effect" and "Windows effect" and the "Sweden effect" considered individually.

I wrote a little article on implementing a simple multivariate A/B test from scratch:

http://www.evanmiller.org/linear-regression-for-fun-and-prof...

All you really need is a function that inverts a matrix and you're good to go. The cool part is that you can actually throw away most of the data that comes in from the firehose, just keep updating the Hessian until you're ready to invert it and bam, you have a model.

EvanMiller··on The Death Of Expertise
I think what the author misses is the large institutional structure that restricts the supply of experts and (as a result) enhances their prestige. These days, with the growing availability of technical information, existing systems of licensure, credentialing, and professionalization are breeding resentment on both sides of the lectern.

The author's focus in the article is on the layman's resentment for the expert. While most people would agree that experts are better qualified to talk about something than Joe Blogger, I think the urge to disagree with the expert is engendered by the larger power structure of expertise that has grown up over the past 100 years.

To take a trivial example, I wear contact lenses. While I agree that an optometrist knows a lot more about vision correction than I do, I resent the fact that I have to pay a lot of money to sit in a chair in a dark room and wait for The Expert to deliver his Opinion about whether I need the -4.5's or -5's.

The author of the article feels like he is the victim of resentment at the hands of his students (references to "intellectual valet" and so forth). I don't think students are necessarily in the wrong -- most are probably in college to get a job, and they're right to resent the stranglehold that universities have on social prestige and career respectability. They end up taking it out on the professors, because, well, they're stuck in class listening to The Expert for hours a day because of this or that degree requirement.

The experts then start to resent the laymen for failing to pay what they feel is a proper amount of respect for the superior state of their knowledge. They conclude that the solution is to come up with ways to enhance their prestige even more, for example by writing articles like this one that talks about how great experts are. I'm afraid this will only poison relations between laymen and experts even more.

What I love about computer programming is that basically all you need to be an expert on something is time, persistence, and an Internet connection. Contrary to the author's claims, the alternative to institutionalized expertise is not that "Everyone is an expert". It's that "You don't have to go to a prestigious university to become an expert". Sure, there are a lot of charlatans running around in the programming world, but the free exchange of knowledge and the absence of licensure has led to both a flowering of human creativity as well as (outside of San Francisco) non-resentful relationships between experts and laypeople.

EvanMiller··on Homogenization of scientific computing – Python is eating other languages’ lunch
I'd just like to add that what I love about Julia is that it actually lets you go deeper than C code. For high-performance computing it's easy to hit a wall with C (i.e. with SIMD vector instructions), and it's fairly difficult to jump the barrier to programming assembly. Julia makes it easy to muck around with the generated LLVM IR code as well as native assembly code. You can go as deep as you want without leaving the Julia REPL.
EvanMiller··on High-end CNC machines can't be moved without manufacturers' permission
Another possibility that hasn't been mentioned is that the purchase agreement might include some kind of first refusal for the manufacturer to repurchase the equipment if the original owner wants to sell. This kind of provision prevents the emergence of a used-equipment market, the existence of which would cut into the manufacturer's pricing power on new equipment. Requiring the manufacturer's consent to relocate the equipment would be one way for the manufacturer to enforce such an agreement.

tl;dr if you sell expensive machinery then do everything in your power to prevent buyers from reselling.

EvanMiller··on Ask HN: How to increase self-discipline as a self-employed person?
I have to challenge the premise that most of the responders have taken, which is that working on a daily schedule is "better".

On the contrary, I find that my work habits are a lot like yours -- and I think that's a good thing! Sometimes I will be possessed by a coding demon and crank out work for days (weeks?) on end. Other times I will putter around watching TV or brainstorming ideas.

For me the whole point of being self-employed was to NOT have to show up to an office (or home office) and work 9-5 every single day. A creative human brain is a rare and marvelous creature, and we understand very little about how it works. I think the best thing to do is to let it run around and work when it feels like working, or read a book when it feels like reading a book. I personally find my creativity withers away under a strict work regimen.

If your work is not creative and you're just grinding it out for money every day, then by all means, follow the advice in the other posts. But if your work requires imagination and making unexpected mental connections, then don't worry too much about "efficiency". As long as you're thinking about something related to work most of the time, over the long run your real productivity will exceed that of all those poor saps who measure output as a function of mindless hours in front of a computer.

Embracing your "lazy" side requires a certain amount of courage, but if you can make ends meet while doing it, you'll be happier and end up doing better creative work. In any event, don't worry too much about how most people say they do things. Do what feels right to you. Good luck!

EvanMiller··on The Rust standard library no longer has any scheduling baked into it
The Erlang VM is a successful implementation of N:M. It addresses some but not all of the issues you mentioned.

* IIRC each new process consumes 284 bytes so it's not "spending lots of memory on creating stacks for your green threads"

* Synchronization is a non-issue as everything is message-passing.

* Disk I/O is blocking, but the VM has dedicated disk threads.

* Dealing with 3rd party code is a problem, which is why relatively few libraries exist for Erlang. However, there is some hope now that NIFs (native interface functions, i.e. C code) can interact with the scheduler, i.e. report how much time has been used and return control if necessary.

You are right that "Building a correct highly concurrent scheduler is no easy task." The Erlang scheduler was not an easy task, and has received a ton of work from extremely talented engineers over the course of decades. Definitely worth a look to see how N:M can be made to work well.

EvanMiller··on My Hardest Bug
While we're swapping war stories and tall tales, here is the chronicle of my hardest bug:

http://www.evanmiller.org/winkel-tripel-warping-trouble.html

EvanMiller··on Open Big Data Computing with Julia
I like Julia because in many ways, it's a better C than C:

* You can inspect a function's generated LLVM or (x86, ARM) assembly code from the REPL. (code_llvm or code_native)

* You can interface with C libraries without writing any C. That let me wrap a 5 kLoC C library with 100 lines of Julia:

https://github.com/WizardMac/DataRead.jl

* You can use the dynamic features of the language to write something quickly, then add type annotations to make it fast later

* Certain tuples are compiled to SIMD vectors. In contrast, the only way to access SIMD features in C is to pray that you have one of those "sufficiently smart compilers".

* Like C, there's no OO junk and the associated handwringing about where methods belong. There are structures and there are functions, that's it. But then multiple dispatch in Julia gives you the main benefits of OO without all the ontological crap that comes with it.

For me, Julia feels like it's simultaneously higher-level and lower-level than C. The deep LLVM integration is fantastic because I can get an idea for how functions will be compiled without having to learn the behemoth that is the modern x86 ISA. (LLVM IR is relatively simple, and its SSA format makes code relatively easy to follow.)

Anyway, I only started with Julia recently, but I'm a fan. I should also mention that the members of the developer community are very, very smart. (Most are associated with MIT.) BTW I am starting a Julia meetup in Chicago for folks in the Midwest who want to learn more: http://www.meetup.com/JuliaChicago/

EvanMiller··on Linux memory manager and your big data
A "50/50" rule such as the one described is a clear signal that the memory manager's cache eviction policy is an ad-hoc hack. This sort of problem would benefit from some old-fashioned operations research-style statistical and mathematical modeling. What's the probability that a page will be needed? What's the cost of keeping it around? A proper model would be able to answer these questions. A "recency cache" and "frequency cache" is missing the bigger picture.
EvanMiller··on [dead]
The provided link is broken. (You probably want http not https.)
EvanMiller··on The Future of Programming
It is satire, not farce.
EvanMiller··on Why are 28/30 of the Top articles about the same thing?
Because for the first time, there is documentary evidence of massive government spying on the United States population. Until this week it was all hearsay from whistleblowers and unnamed sources.

Because we care about individual rights and believe that no one should have access to our private communications without our permission.

Because we're at a critical juncture in world history, and one of the roads ahead leads to the unprecedented concentration of power in just a few hands.

EvanMiller··on DNI Statement on Recent Unauthorized Disclosures of Classified Information
Not the same document.
EvanMiller··on Statistical Formulas For Programmers
Agreed 100%! As the introduction says this is supposed to be a "cheat-sheet" and a jumping-off point for further discovery... I'd love to see (and write!) more posts on when to apply things.
EvanMiller··on Statistical Formulas For Programmers
Thanks! I have relabeled it. (In a previous draft the entry was for unbiased variance and the inaccuracy slipped in during the transition.)
EvanMiller··on Statistical Formulas For Programmers
Uhhhh, Evan Miller here. Not sure why my name is in the submitted title, but whatever.

The current selection on that page is somewhat limited, but I hope to grow it over time. The stuff at the beginning is pretty basic (e.g. standard deviation), but things get pretty gnarly by the time you get to the Kiefer equation. At some point I'll add some more references on how to implement things, e.g. find successive zeros of Bessel equations. For now it should be a good jumping-off point. Enjoy!

EvanMiller··on The next frontier for big data is the individual
Ah yes, "the power of a society based on 10 times as much data". The best examples that the authors could come up with are:

* A professor receiving an automatic notification about when his flight was delayed. This saved him approximately 2 minutes as compared to checking the flight number before he left for the airport.

* Stephen Wolfram figuring out what time he likes to send emails.

Big data proponents need to tell a better story about how data will empower the individual. So far it seems like it's just about large corporations engaging in cross-merchandising, advertisers getting you to click on things, and spy agencies building dossiers without all that pesky legwork.

← PreviousPage 2 of 4Next →