HNHacker News
TopNewBestAskShowJobs

tfehring

3,230 karma · joined February 5, 2018

Hi! I'm Tom. I work on quantifying AI risk at the Artificial Intelligence Underwriting Company (aiuc.com). I've previously worked across software engineering, data science, and actuarial roles at organizations including OpenAI, Airbnb, Allianz, and the Harvard-Smithsonian Center for Astrophysics.

Email: [firstname]@fehri.ng

submissionscomments
tfehring··on California homeowners to fund half of high-risk insurer's $1B 'bailout'
There's no need to draw a discrete line and say it's inappropriate to be on the wrong side of it. People should be able to live where they choose. But people should also pay actuarially fair insurance premiums and market-driven prices for energy and other services, based on those risks and the associated cost to serve them. That's not happening today, because market distortions created by government entities force people in lower-risk areas to cross-subsidize expenses for people in higher-risk areas.
tfehring··on I tasted Honda’s spicy rodent-repelling tape and I will do it again (2021)
It's like the opposite of clickbait. The author did, upon information and belief, taste Honda's spicy rodent-repelling tape, and made a strong case that she will in fact do it again unless someone stops her. Truly giving the people what they want.
tfehring··on Ask HN: Physics PhD at Stanford or Berkeley
I live in Palo Alto and previously went to grad school in Berkeley, and I can corroborate. Frankly you wouldn't need to spend more than a couple hours in either location to pick up on those vibes.
tfehring··on Caltrain's electric fleet more efficient than expected
Electrification wasn't just an efficiency thing, it also made the service much better because of the faster acceleration - IIRC my travel time to SF on weekends went from 75 to 50 minutes, which opens up a lot of use cases.
tfehring··on Caltrain's Electric Fleet More Efficient Than Expected
BART is better than it was but still bad. Caltrain is totally fine though.
tfehring··on Is the world becoming uninsurable?
The 6.0% margin (for UnitedHealthGroup as a whole) already includes that. UnitedHealthcare (the subsidiary health insurer) had a slightly lower operating margin of 5.6% in Q3. https://www.unitedhealthgroup.com/content/dam/UHG/PDF/invest...
tfehring··on Ask HN: Is it a bad idea to make an email domain with an uncommon TLD?
My primary email uses a .ng domain, currently through Fastmail. I’ve had exactly one site reject it and zero spam filter issues that I know of. The only real issue has been that it’s a pain to tell people in person, so I usually end up just giving my old Gmail in those situations.
tfehring··on SF Purity Test
* Co-worked from HanaHaus

* Taken a meeting at Buck's

* Received a check or term sheet at Buck's

* Ate at Zareen's more than 3 times in one week

* Taken an Uber/Lyft from San Francisco Caltrain station to South Bay after missing your train

* Complained about how Waymo doesn't go to South Bay

* Reminisced about Antonio's Nut House

* Googled Antonio's Nut House after hearing someone reminisce about it

* Went to a Stanford talk

* Went to a Stanford talk and actually understood the material

* Made a LinkedIn connection from pickleball

* Said "I'm gonna move to the city" because you can't pull on the peninsula

But a lot of the originals still work for South Bay, even though people probably haven't done as many of them

tfehring··on Why R is the best coding language for data journalism
If you parse e.g. a json file containing a scalar to an R object, that scalar will be represented as a length-1 vector. I get what you’re saying, but “it doesn’t let you access scalars” would suggest to me that it can’t parse or represent values that are defined elsewhere as scalars at all.
tfehring··on Google says AI weather model masters 15-day forecast
I’m by no means an expert in weather forecasting, but I have some familiarity with the methods. My understanding is that non-“AI” weather models basically subdivide the atmosphere into a 3d grid of cells that are on the order of hundreds to thousands of meters in each dimension, treat each cell as atomic/homogeneous at a given point in time, and then advance the relevant differential equations deterministically to forecast changes across the grid over time. This approach, again based on my limited understanding, is primarily held back by the sparse resolution and the computational resources needed to improve it, not by limitations of our understanding of the underlying physics. (Relatedly, I believe these models can be very sensitive to small changes in the initial conditions.) It’s not hard to imagine a neural net learning a more efficient way to encode and forecast the underlying physical patterns.
tfehring··on Debanking (and Debunking?)
I think he only made that claim referring to the profits that the banks earn from their relationships with the crypto companies, not the total profits made by the crypto companies themselves (though both could be true).
tfehring··on Y Combinator and Power in Silicon Valley
It wouldn’t have to scale, but I’m glad that it did. It’s likely that there’s been somewhat more innovation and economic growth in the world over the last ~decade than there would have been if YC stayed smaller.
tfehring··on Numpyro: Probabilistic programming with NumPy powered by Jax
One way I've seen this done in practice is to construct an offline model that produces an initial set of posterior samples, then construct a second, online model that takes posterior samples and new observations as input and constructs a new posterior. This probably wouldn't make sense computationally in a high-frequency streaming context, but (micro)batching works fine.

I've seen lots of other approaches proposed for this over the years, here's a recent Stan forum thread with some links: https://discourse.mc-stan.org/t/updating-model-based-on-new-...

tfehring··on Numpyro: Probabilistic programming with NumPy powered by Jax
PyMC can use NumPyro as a backend. PyMC's syntax and primitives for declaring models are much nicer than (Num)Pyro's, as is the developer experience overall. But those come at the cost of having to deal with PyTensor (a fork of a fork of Theano), which is quite bad IMO, instead of just working with Numpy or PyTorch.
tfehring··on Trump wins presidency for second time
I'm a data scientist, and my impression is that the growth of data science as a profession over the last ~decade has enabled companies to price more efficiently than they used to. That in turn was enabled by technical improvements like cheaper storage and compute and commoditized data infrastructure. I don't have a strong opinion on how much of the inflation this explains, but directionally I'm very confident that companies have gotten significantly more efficient at pricing over that time period, and pretty confident that that would lead to price increases for a lot of businesses.

Supply chain and price shocks during COVID probably accelerated this trend quite a bit - McDonald's would have eventually figured out that the profit-maximizing price of a burger is closer to $4 than $1, but COVID shocks gave it license to raise prices much faster. The good news is that I think of this largely as a one-time shock: once companies have perfectly set profit-maximizing prices, there's no room for more price-optimization-driven inflation, except to the extent that consumers get richer or less price-sensitive over time.

Quoting Matt Levine, "a good unified theory of modern society’s anxieties might be 'everything is too efficient and it’s exhausting.'"

tfehring··on Trump wins presidency for second time
I think it is stupid to vote based on how a politician talks rather than the expected impact of their proposed policies, though of course I realize that that’s how elections have been decided in practice for as long as representative governments have existed.
tfehring··on Bitcoin has made a new all-time high price
It’s not really about dividends. If you buy a share of Amazon stock, you’re getting a cut of the value that a million+ people are producing every day by writing code or moving stuff around. The value most workers create for their employers’ shareholders is much bigger than their rent check, and capturing that value is even easier than being a landlord. Dividends are just moving money from one pocket to the other, the value created by people’s work is where the returns come from. Financial returns without value creation are an anomaly, though not necessarily a rare or short-lived one.
tfehring··on Amazon CEO denies full in-office mandate is 'backdoor layoff'
https://archive.ph/0QcOK
tfehring··on Federal investigators probe Tether
Reserve requirements are completely different from capital requirements. US banks are required to hold about $108 in assets for every $100 they hold in deposits, and their actual holdings are typically in the range of around $110 to $115 in assets per $100 of deposits. Central bank reserves are one type of asset that commercial banks can hold; a 0% reserve requirement just means that commercial banks can hold all of that ~$108 in other assets if they choose to.

In contrast, Tether has historically admitted to having as little as $100.20 in assets per $100 in liabilities [0], with a significant fraction of it in crypto and other assets that effectively wouldn't even count toward banks' capital requirements. It has probably dropped below $100 in assets per $100 in liabilities - i.e., been insolvent - at some point, and even taking its latest audit [1] at face value, it has far less capital than would be needed for a bank with the same asset profile in the US or other developed countries.

[0] https://assets.ctfassets.net/vyse88cgwfbl/1np5dpcwuHrWJ4AgUg...

[1] https://assets.ctfassets.net/vyse88cgwfbl/6h4YWqZOXbwtBaPtYg...

tfehring··on CRLF is obsolete and should be abolished
At least for CSV, there's a divergence between usage in practice and the spec. The spec requires CRLF, but all of the commonly used tools I've encountered for reading and writing CSVs can read files with CR, LF, or CRLF line endings, and when writing CSVs they'll default to either LF or platform-specific line endings. (Even Excel for Mac doesn't default to CRLF!) I think this divergence is bad and should be fixed.

But IMO the right resolution is to update the spec so that (1) readers MUST accept any of (CR, LF, CRLF), (2) writers MUST use one of (CR, LF, CRLF), and (3) writers SHOULD use LF. Removing compatibility from existing applications to break legacy code would be asinine.

tfehring··on An Uber, Lyft Loophole Denys NYC Drivers Millions in Pay
Many contracted projects are contracted on a longer timescale, often weeks or months, compared to ridesharing where the outputs are expected within minutes. Contractors in other fields can often choose when to work on a day-to-day basis, but their next project will only materialize if the demand is there.
tfehring··on Y Combinator Traded Prestige for Growth
Just from public information my impression is that prestige was never a goal while scaling was in the plans from fairly early on.

Anecdotally, some of the best recent founders I know are opting not to apply, which I think is a bad sign. But their reasons have nothing to do with the scale, competitiveness, or prestige of the program.

tfehring··on React for R
One concrete example: R has 5 distinct, actively maintained class systems, at least 3 of which are somewhat commonly used for new projects. I.e., there are 3+ reasonable ways to declare a class, and the class will have different semantics for object access, method calling and dispatch, etc. depending on which one you choose.

Another: R can’t losslessly represent JSON because 1 and [1] are identical. That’s a float (well, float vector) literal by the way, the corresponding int literal is 1L, though ints are very prone to being silently converted to float anyway.

tfehring··on Chai-1: Decoding the molecular interactions of life
Manifold [0] has markets on this sort of thing, but it primarily uses fake money. (They're working on a real-money "sweepstakes" thing, which I'm not super familiar with.) If you're outside the US and looking for a real-money market, Polymarket [1] is probably your best bet. In the US, real-money prediction market contracts are regulated by the CFTC in the US, so availability of contracts is pretty limited; Kalshi [2] would be the most likely option, but I doubt they have anything on this topic.

[0] https://manifold.markets

[1] https://polymarket.com

[2] https://kalshi.com/

tfehring··on What is the longest known sequence that repeats in Pi? (homelab)
Isn't this functionally the same thing that the author already did, just with N=10 and filtering on the first digit instead of the last?
tfehring··on DOJ sues realpage for algorithmic pricing scheme that harms renters
The comment you replied to seems to be targeted at managers trying to advocate for their teams, not at individual contributors trying to advocate for themselves. I agree that reaching out to the compensation team as an IC is generally not going to be an appropriate or effective way to get a raise. But for managers, working with HR on issues like that is just part of the job.
tfehring··on Laid-off California tech workers are sick to death of LinkedIn
Honestly, if you just block or ignore the feed, the other features - profile, social graph, messages, job postings - are basically fine. My biggest complaints on the job seeking side are the mediocre recommendation system and the lack of high quality level and compensation data for roles (like levels.fyi has); on the hiring side, a good recommender would also be helpful, as would better filtering tools. But if you want to replace LinkedIn as the primary sourcing tool for many roles, solving those problems is the easy part - the hard part, like you said, is building up a two-sided network at anywhere LinkedIn's scale.

Would you sign up, make a profile, and respond to messages on a new LinkedIn alternative that has no first-party job listings and no recruiters for companies you'd want to work for? Most people's answer seems to be no, based on all of the failed LinkedIn competitors that have been tried over the years.

tfehring··on Bayesian Statistics: The three cultures
I think each category of Bayesian described in the article generally falls under Breiman's [0] "data modeling" culture, while ML practitioners, even when using Bayesian methods, almost invariably fall under the "algorithmic modeling" culture. In particular, the article's definition of pragmatic Bayes says that "the model should be consistent with knowledge about the underlying scientific problem and the data collection process," which I don't consider the norm in ML at all.

I do think ML practitioners in general align with the "iteration" category in my characterization, though you could joke that that miscategorizes people who just use (boosted trees|transformers) for everything.

[0] https://projecteuclid.org/journals/statistical-science/volum...

tfehring··on Bayesian Statistics: The three cultures
An implicit shared belief of all of the practitioners the author mentions is that they attempt to construct models that correspond to some underlying "data generating process". Machine learning practitioners may use similar models or even the same models as Bayesian statisticians, but they tend to evaluate their models primarily or entirely based on their predictive performance, not on intuitions about why the data is taking on the values that it is.

See Breiman's classic "Two Cultures" paper that this post's title is referencing: https://projecteuclid.org/journals/statistical-science/volum...

tfehring··on Bayesian Statistics: The three cultures
The author is claiming that Bayesians vary along two axes: (1) whether they generally try to inform their priors with their knowledge or beliefs about the world, and (2) whether they iterate on the functional form of the model based on its goodness-of-fit and the reasonableness and utility of its outputs. He then labels 3 of the 4 resulting combinations as follows:

    ┌───────────────┬───────────┬──────────────┐
    │               │ iteration │ no iteration │
    ├───────────────┼───────────┼──────────────┤
    │ informative   │ pragmatic │ subjective   │
    │ uninformative │     -     │ objective    │
    └───────────────┴───────────┴──────────────┘
My main disagreement with this model is the empty bottom-left box - in fact, I think that's where most self-labeled Bayesians in industry fall:

- Iterating on the functional form of the model (and therefore the assumed underlying data generating process) is generally considered obviously good and necessary, in my experience.

- Priors are usually uninformative or weakly informative, partly because data is often big enough to overwhelm the prior.

The need for iteration feels so obvious to me that the entire "no iteration" column feels like a straw man. But the author, who knows far more academic statisticians than I do, explicitly says that he had the same belief and "was shocked to learn that statisticians didn’t think this way."

← PreviousPage 3 of 27Next →