HNHacker News
TopNewBestAskShowJobs

sl8r

287 karma · joined May 29, 2014

submissionscomments
sl8r··on Diffusion Without Tears
I left the TENET references on the cutting room floor.

I too found it really surprising that the reverse-time equation has a simple closed form. Like, surely breaking a glass is easier than unbreaking it? That’s part of what got me interested in this stuff in the first place!

If you haven’t seen it yet, highly recommend the blogs of Sander Dieleman & Yang Song (who co-invented the SDE interpretation).

sl8r··on Diffusion Without Tears
The production of tears is left as an exercise to the reader. /s

Thanks for reading. The 2D simulation section might be more interesting on a first read — it makes the math less mysterious, I hope!

sl8r··on Diffusion Without Tears
Hey, I’m the post author — thanks for reading.

Theta represents all the model params — all the weights in the neural network. The convention is to write theta for the “learned” score function and omit theta for the “true” score function.

sl8r··on Diffusion Without Tears
Hey! I’m the post author. Highly recommend Sander Dieleman’s blog for alternative interpretations https://sander.ai/2023/07/20/perspectives.html

I personally find the SDEs the most intuitive, and the deterministic ODE / consistency models / rectified flow stuff as ideas that are easier to understand after the SDEs. But not everyone agrees!

sl8r··on I wish GPT4 had never happened
> This has been the promise over and over again, for centuries, and it has consistently not paid off. Where's the predicted society where automation allows us all to work for two hours a day, and spend the rest at leisure?

In the “Sad Irons” chapter of Caro’s LBJ biography, he talks about the pre-electrification lives of Texas farmers. In comparison with that, our whole day is leisure.

Similarly: As late as 1900, the poor in Europe were so severely malnourished that growth stunting was common. Look at Our World in Data’s charts of height over time. Or Robert Fogel’s “The Escape from Hunger and Premature Death.”

Etc. Etc.

sl8r··on Are there limits to economic growth?
https://astralcodexten.substack.com/p/your-book-review-the-w...
sl8r··on Ask HN: How are you preparing for the incoming recession?
I'm not saying that crises aren't predictable — I'm pointing out that institutional failure isn't always a leading indicator (as in OP's argument).
sl8r··on Ask HN: How are you preparing for the incoming recession?
what major institution failed in the early stages of [1] the japan asset bubble or [2] the dot com bubble? Or farther back, the 1840s railroad mania or the south sea bubble?

credit bubbles often pop when some institution can't cover its obligations, but asset bubbles don't seem to need such a failure — and can deflate on their own.

sl8r··on Excel World Championship Finals
Interestingly, they’re actually used all the time in LBO models. Because the default is to sweep all FCF to pay down debt, but then the interest expense is dependent on FCF, which depends on the interest expense… Sort of a trivial example b/c you could solve it by being more granular with time periods, but in practice people just use the circular ref.
sl8r··on Kelly Criterion
I made a streamlit app about Kelly last year, showing how to bet when you have an "edge" over a toy market of coin flippers: https://kelly-streamlit.herokuapp.com/

Other references I found interesting:

  - Cover and Thomas's "Elements of Information Theory" shows some interesting connections between Kelly betting and optimal message encoding.
  - Ed Thorp, the inventor of card counting, has a nice compendium of papers on this in "The Kelly Capital Growth Investment Criterion".
sl8r··on A Mathematical Theory of Communication (1948) [pdf]
If you're interested in this, check out Cover's Information Theory textbook — the rabbit hole goes much deeper. One of the most interesting examples, is that when you're betting on a random event, Shannon entropy tells you how much to bet & how quickly you can compound your wealth. Cover covers (heh) this, and the original paper is Kelly: http://www.herrold.com/brokerage/kelly.pdf
sl8r··on Socialist Millionaire Problem
Won't you possibly get some information about some passengers by doing this, if you know / can figure out the distribution of the random numbers?

E.g., say that the random numbers are uniformly distributed between -200 and 200. If somebody says a number like 425, then I know they weigh at least 225 pounds. And the probability they weigh more than 225 - k pounds is 1 - k/400.

sl8r··on Ask HN: Favorite Nonfiction Books of 2019?
* "the dream machine" is an fascinating, joyful account of jr licklider's work in interactive computing

* "hard landing" makes the airline industry seem tumultuous and exciting

* "a man for all markets" is the autobiography of ed thorpe, father of card counting and quant hedge funds

* "unix: a history and a memoir" is a mischievous first hand account from brian kernighan (of unix / c / awk fame)

sl8r··on How category theory is applied
FWIW, I think the best "applications" (if they can be called that) come from domains like algebraic topology and algebraic geometry, where some concepts come up so frequently that it's useful to formalize them in category theory. The homology proof of Brouwer's fixed point theorem [1] is one good example. Things like Eilenberg-Maclane spaces are another (really is most natural to think of them the thing that represents some functor).

I wouldn't say that you can't do this stuff without category theory, but it does make it easier / clearer. (Similarly, you can do a lot of geometry without coordinates, but coordinates definitely make some stuff easier / clearer.)

[1] https://www.wikiwand.com/en/Brouwer_fixed-point_theorem#/A_p... [2] https://www.wikiwand.com/en/Eilenberg%E2%80%93MacLane_space

sl8r··on How category theory is applied
Although, to play devil's advocate, you can prove results with set theory that you'd care about even if you weren't super interested in foundations, usually by playing with different cardinalities. E.g.:

Call a real number "algebraic" if it's a zero to some polynomial with rational coefficients. (e.g. \sqrt{2} is algebraic since it's a zero for x^2 - 2). Claim: There exist non-algebraic ("transcendental") numbers. Proof: There are only countably many polynomials, and so there are only countably many algebraic numbers, but there are uncountably many reals. Similarly, there are numbers that aren't Turning-computable. Etc.

sl8r··on Compare career levels across companies
I mean... the parent post is right, the site I linked seems to be calculating CA state tax incorrectly.

But: 42% is the effective tax rate (not marginal, average). And this is pretty close to 45%. For a hypothetical person making $500k or more as regular income in CA (as many of the posted salaries above would be), they would indeed pay about $210k in taxes plus Medicare plus Social Security.

sl8r··on Compare career levels across companies
Some of these reported salaries would come close to that, if you include social security and medicare:

For a $500k income in CA

    29.59% Federal Income Tax
  + 10.48% CA Income Tax
  +  1.59% Social Security
  +  1.99% Medicate
  ----------------------------
  = 43.65% total tax rate.
(Source: https://smartasset.com/taxes/california-paycheck-calculator)
sl8r··on Thomas Bayes and the crisis in science
Late to the party, but:

> Both the historical and the control bucket used version A of the website, and they are consistent in their 2.0% conversion rate. Version B is different, and it appears to have a different conversion rate of 2.5%. So why should it not have a future conversion rate close to 2.5%?

It's all a matter of degree. You'd model B's rate as closer to 2.5%, but probably not centered around 2.5%. As you observe more data, the prior becomes less important. E.g., with 10k samples as in the original example, if you used Beta(2+1,100-2+1) as your prior, your posterior would be Beta(252+1, 10100-2+1) as your posterior, which is centered at 2.495%. But if you only had 1000 samples (and 25 conversions), you'd get a distro centered at 2.45%. And if you only had 200 samples (and 5 conversions), you'd get a distro centered at 2.33%. Etc.

> Let's replace the website with a 6-sided die. Historically, the probability of throwing a 3 was 1/6. Now you replace your die with a different die and throw it 10,000 times; the 3 comes up 2560 times. If I had to guess how many times the 3 comes up the next 10,000 throws, I certainly would bet that it's closer to 2560 times than to 1667 times.

In the case of a die where you believe any weighting of the faces is equally likely, this would be true. So this may be an appropriate model in this case. But in the case of the website, I don't think the conversion rates are equally likely, even for a new, un-tested site. If the historical conversion rate is 2.0%, and I'm forced to bet on the most likely conversion for a new (never before seen) variant B, I'd much rather bet on a number near 2.0% than a number like 99%.

> Case B: The historical version A of the online shop did not have any influence on the conversion rate during the testing of version B (compare the dice example above). Then both ranges are equally plausible.

This is exactly what I'm claiming is not true. It's not that A influences B, it's that A tells you something about the likely range of A and B (in this specific case of an e-commerce site). (The reason I chose the ranges [2.0%, 2.5%] vs [2.5%, 3.0%] is that if you model B independently, you'd be indifferent between these ranges; but if you use A to inform a prior, you'd prefer [2.0%, 2.5%].)

sl8r··on Thomas Bayes and the crisis in science
> Why would that be wrong?

The issue is that modeling B with a distro centered around 2.5% ignores what we know about the historical conversion rate (2.0%) and the control bucket's conversion rate (also 2.0%). If our goal is to make the best estimate for the future that we can, we should take this data into account when evaluating B. As a thought experiment, imagine that you have A at 2.0% and B at 2.5% conversion for Week 1, with a historical conversion rate of 2.0%. Someone says they'll pay you $100 if you correctly guess what B's conversion rate will be next week, either (i) in the range 2.0% to 2.5%, or (ii) in the range 2.5% to 3.0%. I'd prefer to bet on (i) than on (ii).

> What would a Bayesian conclude instead?

One simple approach would just be to start with a more informative prior, like Beta(2+1,100-2+1) instead of Beta(1,1). This would pull bucket B's posterior distribution closer to 2.0%. Another approach is to use a hierarchical model [1], which will fit the individual buckets' priors for you.

[1] Here's something I wrote on this a couple years ago, more focused on solving multiple comparisons problems but with the same proposed solution: http://normal-extensions.com/2014/07/16/ab-testing-hierarchi...

sl8r··on Thomas Bayes and the crisis in science
Part of the issue is that you can't not assume a prior; it's unavoidable. The Bayesian POV just makes this assumption clear / explicit, while many frequentist methods (if naively applied) amount to choosing a uniform prior.

E.g., imagine you have an e-commerce site that has, historically, had a 2% conversion rate (landing page to purchase). Now you run an A/B test with two variants, a control (A) and a treatment (B). Both buckets get 10,000 landings, of which A converts 200 of them and B converts 250. How can you tell if A is better than B? Cutting to the chase, the frequentist approach (if applied naively) would be to model B as some distribution centered around 2.5%, for example N(2.5%, 0.15%) or Beta(251, 9751).

A Bayesian would say that this assumes a uniform prior - but that this is probably a bad prior because it ignores what we know about the historical conversion rate of 2.0%. Said another way, the above amounts to saying (before we run the test) that we think it's just as likely for B to have a conversion rate of 2.5% as it is to have a conversion rate of 100%. Clearly we don't actually believe this.

sl8r··on General Thinking Tools: Mental Models to Solve Difficult Problems
This is probably because Howard Marks popularized the idea under that name in "The Most Important Thing". He looks like FS's main source in the detail piece (https://www.fs.blog/2016/04/second-level-thinking/).

I think he does mean it as an analogy to the math concept, like ~x being first order and ~x^2 being second order for x near zero; x dominates but x^2 (the "second order") gets you closer to the truth. And similarly for x^3, x^4, ..., x^n.

sl8r··on Nassim Taleb, Absorbent Barriers and House Money
One interpretation of Kelly is maximizing the e.v. of the log return. Another (equivalent) interpretation is maximizing the expected IRR -- which is where the "in the long run" comment comes in, since "in the long run" you care more about the internal rate of return than the expected value of any individual bet.

E.g., imagine a bet that costs $1 to play and pays out $2 with probability 60% and $0 with probability 1%. How much would you bet? The expected value of the bet is $1.2 per dollar you bet, so for a single bet, you might wager 100% of your bankroll. But "in the long run" you'll loose all your money doing this. Instead, Kelly would recommend that you bet only 20% of your bankroll. "In the long run", you'll make infinite money doing this. (Not only that, but there's no other strategy that will make you money faster.)

sl8r··on Twitter explores subscription-based option
Twitter had $2.5B of revenue in FY '16. The average user with more than 10k followers probably only has slightly more than 10k followers (because the distribution of followers/user has most of its mass near 0), so the average rev. per 10k+ user is almost certainly less than $2/month.

So for this model to get TWTR near its current revenue, Twitter would need to have 104 million accounts with more than 10k followers. Twitter only has about 320 million MAUs.

If you think that the % of MAUs with 10k+ followers is closer 0.1%, then the avg. monthly charge would need to be $7,800/month.

sl8r··on Huffington Post's criticism of 538's election forecast
""" As a financial analyst at an investment bank, or a research analyst at an economic consulting firm, your job would be in serious jeopardy if you produced 538’s model output without a clear explanation of how those fat tails that represent an inordinate number of close to impossible scenarios could actually occur. A model like that just isn’t client-ready. Time to re-think those assumptions! """

This line of reasoning irks me -- If the outputs don't agree with with my preconceived conclusions, the model must be wrong. (Rather than: Maybe my preconceived conclusions are wrong.)

I don't have an opinion on the 538 model, not having analyzed the internals. But I don't like the idea of criticizing a model because you don't agree with its results.

On the bright side, the election will shortly be over and we'll have at least some measure of how accurate each model (538, HuffPo, etc.) actually was.

sl8r··on A Guide to Employee Equity
Several comments allude to liquidation preference / stock-participation for preferred stock holders.

While both of these terms do bite into the common share stake, as others have pointed out, their effect is model-able; so you can see where you stand under different exit scenarios. Ask your employer, or get access to something like Pitchbook, to find out:

1. The amount of $ raised from investors, along with the % of equity in common vs preferred shares.

2. The liquidation preference (1.0x is common) for preferred.

3. Whether the preferred stock participates (not doing so is common).

4. What % of common equity your grant represents.

You'll then be able to draw a payout diagram (like this[1], written by Andy Rachleff) showing how much you'll make for different exit prices. Be aware that under some complexities, like the fact that later rounds will cause dilution and that certain exit scenarios scenarios (like acquisitions) may trigger a different liquidation preference for preferred stock.

[1] https://blog.wealthfront.com/wildly-different-financial-outc...

sl8r··on Tesla Makes Offer to Acquire SolarCity
Others have already made this point[1] but to me, SolarCity has some aspects that make it look like more of a of a tax-arbitrage business than a solar panel retail business. It reminds me a bit of the ethanol blending tax credit, which led, in some cases, to companies mixing ethanol with petrochemicals solely for the tax benefit.

[1] http://www.newsmax.com/BradleyBlakeman/solar-kroll-subsidy/2...

sl8r··on Why Do Nigerian Scammers Say They Are from Nigeria? (2012)
Cf. The Dark Lord of the Internet [1] and the relevant HN discussion [2], which give an interesting peek into the psychology of internet scams.

[1] http://www.theatlantic.com/magazine/archive/2014/01/the-dark...

[2] https://news.ycombinator.com/item?id=6972139

sl8r··on Ask HN: How to learn machine learning on my own?
What are you looking for in particular?

If you're looking to gain functional familiarity / put in practice reps with classical classification & regression algorithms, I think that running through the online tutorials for scikit-learn is the best bet.

If you're looking for the theory behind the above, I think the book by Peter Flach is the best intro; "Elements of Statistical Learning" is the classic tome, but much more mathematically motivated.

If you're looking for more specialized subjects, each has its own resources. Bayesian modeling? Gelman's BDA3 and Cam David Pilson's github book. Gaussian processes? Rasmussen. Etc., etc. for neural networks, reinforcement learning, etc.

As a random recommendation: David Mumford's "Information theory" is eclectic and fun, but disconnected from the mainstream.

sl8r··on Ask HN: How to earn over $250K a year?
Tactically, for the tech industry:

1. Earning potential is higher in management, if you can make it to a director-level or vp-level position at a major public tech company; take a look at the relevant comp figures for Facebook, Google, Microsoft, etc. on Glassdoor.

2. Being a sales rep at an enterprise software company with an aggressive comp plan is another option, if you can do the work. As @kasey_junk noted, this is because sales is compensated on a percent of revenue basis.

You might also look at technical consultancies, or try to become a technical expert for pe/vc firms.

Gaining wealth through investing your own money is hard; even at a market-beating 12% a year, it would take you more than 20 years to turn $100k into $1M. You can make money this way, but you need to manage a fund. The entry point is an analyst role at a hf/pe/vc firm.

I think @kasey_junk is 100% right; ultimately, you need to tie your work to some kind of economic value. We as humans are bad at assessing indirect value creation, so the roles that tend to achieve this tied-to-economic-value comp in practice are [1] sales, [2] investment professionals, [3] substantial equity owners.

sl8r··on What happens when private equity buys your competitor?
I think this is more about distributions than it is about expected values. If VC and PE generate roughly the same returns to their investors (say 20% to 25% IRR), what does this mean for the companies they invest in? Fred Wilson notes (http://avc.com/2009/03/what-is-a-good-venture-return/) that with a five year average horizon, he expects roughly three buckets of outcomes:

1. 1/3 of investments go to zero, i.e. blow up and lose substantially all of investors' money.

2. 1/3 return 1.0x to 1.5x (on average across the bucket).

3. 1/3 return 7.5x (on average across the bucket).

That is wide distribution of outcomes. In PE, on the other hand, you'd get a much tighter distribution of outcomes around 2.7x returns. A single investment (much less 1/3) going to zero would destroy the fund, so PE funds want to prevent that from happening. The conclusion is that Vista is pretty sure it can get a 2.7x outcome or better, and it's also pretty sure it wont zero its investment.

So should Ping's competitors rejoice after the Vista buyout? It really depends on how quickly they're growing and how much market share they think they can win. Do they believe that either [1] Vista will fail in 2.7x-ing Ping, or [2] they can succeed even in this 2.7x world? If they believe either of these things, then the buyout is probably good news for them; if they don't, then it's probably bad news for them.

Page 1 of 2Next →