HNHacker News
TopNewBestAskShowJobs

andy_wrote

436 karma · joined November 14, 2013

Software engineer at Flatiron Health
submissionscomments
andy_wrote··on S&P 500 to exclude Snap after voting rights debate
Although the immediate reason BRK/B ended up in the S&P 500 was because they acquired BNI (Burlington Northern, a railroad company) for stock in 2010. But fair enough that it wasn't in there beforehand and is still around now. I don't remember why, probably some other S&P decision about the big-numbered, less liquid BRK/As.
andy_wrote··on Why Sonic the Hedgehog is 'incorrect' game design
I definitely agree with your point, though I'd recast it slightly away from "twitchy"/"speedrun"/map memorization and towards the thrill of things being fast and disorienting. [1] Sonic involved lots of bright lights, fast animation, and noises. It was fun because everything ran out of control sometimes, or at least seemed to, and I think this, aside from considerations of wanting to optimize for efficiency of solution, was a large part of its joy. (I say "seemed to" because I remember elements of levels that blasted you through loops and walls at dizzying speed, but in practice served as just a rail that took you from point A to B.)

I didn't realize before I read this article that the design of Sonic (even excluding Sonic Spinball) was driven by pinball, though it makes tons of sense in retrospect. Pinball's got the same thrills; you're constantly in an unstable state, not really learning a puzzle so much as trying to manage a high-degree of chaos before you inevitably get thrown off the horse. It's a different brand of fun than the article's "correct" game design, less/more appealing to different people and less/more appealing when you're in the wrong/right mood.

[1] "Ilinx" in the taxonomy of Roger Callois's "Man, Play, and Games," though a personal caveat that I didn't really enjoy that book.

andy_wrote··on Why isn't everything normally distributed?
This. Technically the radii can't even have a normal distribution, because the support of the normal is (-Inf, Inf) but the radii are bounded below by 0.

So you might say that the radii = r + e, where r is constant and e is normal around 0 with variance way, way smaller than r. But then the volume of the ball bearings = K * (r+e)^3, and because r >> e, the largest source of randomness term is going to be 3Kr^2e, which makes the whole thing look pretty normal.

andy_wrote··on Show HN: A Set of Dice That Follows the Gambler's Fallacy
Interesting - just perusing those links, it sounds like a multi-armed bandit problem, in which you reason that if something has worked out before, you should tilt your bets more in that direction. In the context of the urn model, you'd return more balls of the same color for every successful draw. In the context of medicine, you can balance between proving or disproving a treatment effect and actually supplying that treatment to the test subjects who need them.

Relatedly, there's a Bayesian interpretation to overweighting successful past draws. A model where you return one extra ball of the same color to the urn gets you a Dirichlet-multinomial distribution, which is a die-roll distribution where the weights to each face are not known for sure, but are given a probability distribution and revised with observed evidence. In other words: here's an n-sided die, I don't know its weightings, but as I observe outcomes I'll update my beliefs that the sides that come up are more favorably weighted. The number of balls in the urn you start with correspond to your priors; only 1 ball of each color means a very weak belief that it's a fair die, 1000 balls of each color means a strong belief, unequal numbers mean that you start off believing it's weighted.

andy_wrote··on Show HN: A Set of Dice That Follows the Gambler's Fallacy
I've never played in real life with the deck, but I first came across this version on a software version of Catan I had (I think Xbox?). I really enjoyed it and my guess is that in real life it adds a substantial angle to negotiating placement of the robber. In particular, bribing people to put/not put the robber in a given location - if the deck is 7-heavy or your good squares are already drawn, no big deal.
andy_wrote··on Show HN: A Set of Dice That Follows the Gambler's Fallacy
There's a probability model called the Pólya urn where you imagine an urns containing numbered balls (colored balls in a typical example, but to draw the comparison with dice we can say they're numbered 1-6), and every time you draw a ball of a certain color, you put back more balls according to some rule. A few probability distributions can be expressed in terms of a Pólya urn, see https://en.wikipedia.org/wiki/P%C3%B3lya_urn_model.

A fair 6-sided die would be an equal number of balls numbered 1-6 and a rule that you simply return the ball you drew. You can get a gambler's fallacy distribution by, say, adding one of every ball that you didn't draw. I read the code as a Pólya urn starting with 1 ball 1-N and doing that on each draw plus reducing the number of balls of the drawn number to 1.

Also related, in 2d space, is the idea of randomly covering the plane in points but getting a spread-out distribution, since uniformity will result in clusters. (If you're moving a small window in any direction and you haven't seen a point in a while, you're "due" to see another one, and vice versa if you just saw a point.) Mike Bostock did a very nice visualization of that here: https://bost.ocks.org/mike/algorithms/

andy_wrote··on The Mandelbrot Monk (1999)
Fun story! But look at the date at the bottom...
andy_wrote··on Dark Futures: What happens when literary novelists experiment with Sci-Fi
I've never read Lessing, but Le Guin is one of my favorite authors. For me, _The Left Hand of Darkness_ is easily on a par with other canonical "great" 20th-century literature I've read. I've always felt that if she were born later, now that sci-fi has become a little less fringe and a little more respectable, that she'd be more properly grouped this way.
andy_wrote··on [dead]
Just saw it starting to happen to me a few minutes ago, at which point the status page still listed 100% app server availability, but it's since ticked down to 95%. Edit: just the web interface, I can still push code.
andy_wrote··on Ask HN: What does your production machine learning pipeline look like?
At Custora (YC W11), analyzing consumer behavior and presenting the insights through a web-based interface.

We have an in-house system for modeling DAGs of statistics, perhaps like Airflow (I've only read about Airflow, never used it, so can't comment more there). Computed values at different nodes in the DAG can be cached by their arguments. So for example, a fitted model would be a very logical node to cache, as you'll have many other potential statistical requests that would depend on the model.

Refitting the model would entail either clearing that node; if you changed the inputs to the fit, a separate stat would be cached and other dependent stats would know the difference and point to the right stat depending on the arguments you passed.

A lot of our models are Bayesian in nature, so generating predictions is typically a two-part process: training parameters, which is slow, can happen infrequently and which need not critically include all of the latest data, and applying the posterior update, which is faster and which needs to be redone on every data update. (We import data in batch.) Retraining is somewhat ad-hoc right now, although we've got an active project on the docket to systematize this and produce streamlined before-and-after comparisons.

Computation work is dispatched to EC2 instances by a Redis-based job manager (qless, developed by Moz, formerly SEOmoz). We do the stats work in R, in-memory. For things like order and transactional data this is feasible even for relatively large retailers, but we're gradually looking at involving Spark more so that we have the capacity for larger analyses. (We already do use Spark for some non-ML tasks, like customer data import.)

andy_wrote··on An Ad Hoc Affair: Jane Jacobs’s Clear-Eyed Vision of Humanity
I'm in the middle of Death and Life right now. I'm enjoying it for sure, and I find myself nodding my head at a lot of what she says, particularly stuff about border vacuums and mixed-usage areas, both of which have figured into my own past decisions about where to live.

One thing I do find lacking is a discussion of assessing tradeoffs, particularly when her design principles bump up against natural pressures (i.e. driven by the market, not by fiat city planners). For example, she advocates a mix of new and old city buildings. Enforcing this may generate housing and office shortages, if a neighborhood becomes attractive at a rate exceeding that which we can maintain the desired old-new mix. How do we trade that off; at what price do we value incremental old-new mix?

She does cover some of this in her chapter on self-destruction of diversity, but her solution of zoning for diversity was a little tautological for me, just stating that diversity should be enforced. What I'm interested in is how we weigh that against the lost utility of individuals and businesses who would prefer to move to the neighborhood in question.

It's definitely a tough question! While I'd prefer some more head-on discussion of tradeoffs, I think (unfortunately, decades later) that there's still plenty of room for the ideas she advocates to be recognized at all.

andy_wrote··on Unlearning descriptive statistics
Thanks for writing it! It's one of my favorite math blog posts floating out there.
andy_wrote··on Unlearning descriptive statistics
For readers who are OK with some math, I recommend John Myles White's eye-opening post about means, medians, and modes: http://www.johnmyleswhite.com/notebook/2013/03/22/modes-medi... He describes these summary descriptive stats in terms of what penalty function they minimize: mean minimizes L2, median minimizes L1, mode minimizes L0.

A single-number statistic is _going_ to leave things out, so if you must boil things down into one number, or even a few numbers, you're going to lose something that you had in the raw data. This is why I find claims along the lines of "statistics don't tell the whole story" a little bemusing - of course they don't, the very definition of a statistic is a summarization of data that is easier to work with. The question is what data is kept or lost, or more generally what importance we place on different aspects of the raw data such that it's reflected in our descriptive statistics.

The lessons for non-technical people who want to communicate with descriptive statistics are to recognize that summarization is inherent in the nature of any descriptive statistic, that they are thereby opinionated in some way in terms of what they've preserved and what they've left out, and to recognize whether those opinions are appropriate for your purpose.

andy_wrote··on Why UBS Is Hiring So Many Quants
The only time I've actually seen a bat on the trading floor was as a branded piece of swag from the BATS exchange: https://www.batstrading.com/
andy_wrote··on Why the High Cost of Big-City Living Is Bad for Everyone
From the article: "One hardheaded answer is to build more housing. An increasing supply of housing would theoretically put downward pressure on prices. The reality, unfortunately, is that almost all urban construction happens too late."

That last sentence does not sound like an argument against this approach, but an additional argument for it. If there is a secular trend of urban rents outpacing inflation, it sure seems like a "better late than never" situation.

andy_wrote··on Ask HN: What are your favorite board games?
I've never played regular Werewolf, unless you are referring to what I know of as the informal party game Mafia (several nights, one killer, several townspeople, maybe one or two additional roles who get special abilities).

In that case, yes much better:

- Everyone is involved start to finish, and a free phone app serves as the emcee.

- People receive cards with (usually) unique roles that furnish them some special information in their own way. I think some Mafia variants have minor special roles, but the makers of ONUW really did a good job thinking of creative ones.

- Many roles involve moving cards around at night, so when you wake up you usually don't know for sure whether you're still the role you thought you were, or whether a card you saw at night is still where it was when you saw it.

andy_wrote··on Ask HN: What are your favorite board games?
I'll second Hanabi. In addition to being inherently very fun, collaborative games are a great way to bring in people who aren't as into board games and are worried about getting stomped on by experienced players.

An interesting and much more challenging variation for experienced Hanabi players is to disallow people from saying a number or color. You point to a set of cards in another player's hand and that's it. Those cards share some attribute, and all the other cards in the hand don't have that attribute, whatever it is.

andy_wrote··on Ask HN: What are your favorite board games?
I totally agree with your point about the comparatively low conflict in ONUW. In addition to the points you mentioned that drive this, I'd add uncertainty about your own team membership, and the fact that "good guys" may also be telling lies, so it's not necessarily true that a liar is a bad guy.

These factors prevent people from identifying too heavily with one team or the other, and that issue, for me, is at the heart of why there is less conflict in this game than other Mafia-style games. Right before getting into ONUW we had played The Resistance a bit, a similar game which had some fun moments, but in which we found the conflict potential to be high. Some people would be just dug in on one view, in direct conflict with others, throughout almost the whole game.

andy_wrote··on Ask HN: What are your favorite board games?
I can't believe that I'm the first to mention One Night Ultimate Werewolf, which has been the consensus favorite amongst my friends for a while now.

It has relatively simple rules but very solid strategy (rewards logic and duplicity). The games are short, so you never feel "stuck" on a long game and when you're a beginner, you can rapidly absorb new lessons and strategies and apply them to the next round. The replay value is tremendous.

I have observed/heard about the game not "clicking" for some players the first few times. You _can_ reason out substantial amounts of information by sharing claims and thinking hard about what you personally know, and you _can_ tactically disrupt other people. I think if you have a crowd of new people, it helps to have an experienced player sit out one round and emcee, encouraging certain lines of thought and discouraging others. One of my friends said he only really "got" it after the third round, when he saw me spin a story from start to finish so that I could pin a wolf on someone when I in fact was a wolf.

I also love Dominion, which others have mentioned. (That's my personal favorite; Werewolf is my friend group-favorite.) It is in a very different genre, but it also has fast-cycling games, deep strategy vs. simple rules, and huge replay value, which are three aspects of board games that I really value.

andy_wrote··on Nintendo re-releases NES as mini console
Towards the end of its life, my NES wouldn't work if I pushed the cartridge all the way in. I had to push it just the right depth in, such that the top of the cartridge (the part facing you when you inserted it) would lightly scrape the edge of the slot as I pushed it down into the machine. It was an art. Fun times!
andy_wrote··on Don’t invert that matrix (2010)
In R there is a little syntactical nudge in this direction in that the matrix inversion function is actually called `solve`. `solve(A, B)` gives you the solution to AX = B, and B's default argument is the identity, so `solve(A)` just gives you the inverse, and any instances of `solve(A) %*% B` in your code should stand out as red flags. As rcxdude mentions this is something that you tend to care more about in statistical applications, and that's what R is for.
andy_wrote··on The Energy Startup Conundrum
It sounds like the problem Fong faces is that the industry she wants to work in is very capital-intensive. In that case, does the startup/VC/equity model make sense? Wouldn't it make more sense to be attached to a large company with big balance sheets, access to debt markets, and a good credit rating to get this large capital expenditure funded? (Without having any first-hand knowledge of these places, I'm sort of thinking of places like Google X or Bell Labs.)

Perhaps the issue is a lack of corporate/financial structures that allow people who want to work in capital-intensive industries to enjoy some of the nice aspects of working at a small startup: casual atmosphere, lack of corporate hierarchy, openness to change and innovation. I don't know whether that's something that can be wholly remedied - you can't be "move fast break stuff" when you're building something like a bridge. But I'm not convinced that the usual "startup" model is the best way here.

andy_wrote··on Tennis for Two was an electronic game developed in 1958 on an analog computer
If you live in NY, you can play this at a currently ongoing exhibit about the history of computing in New York at the NY Historical Society.

(Review: I would say the exhibit was decent, though not great - might not be much new for people already familiar with computing history, and it's very clearly sponsored by IBM, though you could argue that this is not unfair for a history of NY-area computing. There's also a fun little exhibit on comic book heroes at the museum, plus I'd never been there before, maybe all added together it's worth a ticket.)

andy_wrote··on Conceptual Debt Is Worse Than Technical Debt
I think they don't have to be related. The article's example of tags and folders is a good illustration. "Tag" and "folder" models might have a clear enough implementation in the code; a programmer might wonder why they're both implemented but may have no trouble supporting both in principle. But the existence of both may be very confusing to the user.

The cost inflicted on the coders is not coders' conceptual debt, but users's conceptual debt - certainly costly to coders as they have to support multiple patterns, but I think there is a difference. I guess when I think of coders' conceptual debt, I'd think of something that may be abstracted away at the UX level such that the user doesn't see it, but inflicts pain on implementors due to counter-intuitive patterns.

andy_wrote··on Conceptual Debt Is Worse Than Technical Debt
The article, which I liked, is very end user-oriented, but I think that the ability of other programmers to interact with your conceptual models is just as important. Other people have to work with your code, and their ability to develop new features will very much depend on how easily they can work with your conceptual models. Onboarding programming hires into a poorly organized conceptual model will be a mess, and disheartening to the new person.

Good conceptual models can be extended, interacted with, and built upon, and that which an end user may regard as a "feature," including features not yet thought of, flow naturally from the best models. So there are some knock-on benefits to the end user from good conceptual models under the hood, even if the end user doesn't see them or have an opinion about how intuitive they are.

andy_wrote··on Conceptual Debt Is Worse Than Technical Debt
I agree totally re your point of not understanding of the domain, it can be really hard to make good conceptual models until you've actually tried throwing some code at the walls.

I guess I'd rephrase your point about "one to treat as a throwaway" as saying that you should realistically expect to be moving back and forth between architecting the conceptual model and implementing it. I find it useful to first think about the conceptual model, then write some code for a while, then revisit the model and see what difficulties I've hit and/or what new good ideas I've thought of in the process, etc.

andy_wrote··on Leó Szilárd: A Forgotten Father of the Atomic Bomb
I was going to say - I believe historians regard "The Making of the Atomic Bomb" as the definitive work on the subject, and Szilárd appears in the first sentence of that book. (I've only read the first chapter or two, been meaning to get around to this someday...)

Like this article, the book opens with a description of Szilárd thinking about a nuclear chain reaction while walking about London.

andy_wrote··on The Grammar of Data Science
In light of this issue, dplyr's tbl_df structure (a light but helpful wrapper around data.frame) actually has different drop defaults, for example

  > x <- data.frame(foo=1:5, bar=1:5, baz=1:5)
  > dim(x[,'foo'])
  NULL
  > dim(x[,c('foo','bar')])
  [1] 5 2
  > dim(x[,'foo',drop=FALSE])
  [1] 5 1
compared to

  > x <- dplyr::data_frame(foo=1:5, bar=1:5, baz=1:5)
  > dim(x[,'foo'])
  [1] 5 1
Although I think these are more reasonable (I've got multiple commits at work with messages bemoaning drop=FALSE), this can ironically also mess you up if you got used to the old defaults :)
andy_wrote··on The Grammar of Data Science
Wickham et al's R packages are great, especially dplyr, and I think should be taught to new R users pretty much right off the bat. I find R's syntax to be a big hangup for new learners, especially on indexing and apply-to-each (sapply, mapply, just plain apply...), but dplyr really makes life much easier.

The %>% operator alone (which to be fair was originally from magrittr) is a great help. Not sure if this is my personal biases, but I always find it easier to read calls chained postfix-style.

andy_wrote··on Pinterest open sources Pinball – a flexible data workflow manager
Are there people who have more experience with comparative workflow managers who can quickly see the pros and cons of Pinball vs. Luigi? Perhaps someone at Pinterest who tried out other systems, as was mentioned in the post? (Though maybe Luigi wasn't available to the public when this comparison happened.)
← PreviousPage 2 of 3Next →