HNHacker News
TopNewBestAskShowJobs

tylerhou

3,503 karma · joined August 13, 2016

PhD student at Princeton.

  website: https://tylerhou.com
  github: https://github.com/tylerhou
submissionscomments
tylerhou··on Don't Be Fooled by this Summer of AI Hype
At least I know more than you.

https://www.reddit.com/r/mathematics/comments/1wauync/on_the...

tylerhou··on Don't Be Fooled by this Summer of AI Hype
OK - “often do” was wrong. There are cases where local communities’ water supplies was impacted by data center overconsumption, but they aren’t the norm [1].

A problem you’re not considering is that data centers, at least ones in current operation, consume mostly potable water, so you can’t compare agricultural use directly to data center usage. Doing some rough math:

Phoenix has about 150 data centers operational or planned. Let’s say 100 of those are operational. The average mid-sized data center uses 1.4 million liters / 0.37 million gallons of water per day for cooling servers. There is an additional 3x burden of water consumed to generate electricity. Let’s assume 70% of the data center water and 20% of the supplemental water consumed is potable [2]. We’re looking at 1.3 * 0.37 = 0.481 million gallons of potable water consumed per data center. So 100 * 365 * 0.481 = 17.5 billion gallons of potable water per year in the Phoenix region. That is a significant and growing fraction of the 110 billion gallons of potable water Phoenix produces per year.

These are using conservative estimates for potable water fraction.

[1] https://www.reuters.com/sustainability/climate-energy/desert...

[2] In 2023, 78% of Google data centers’ water consumption was potable. See “Alternative water sources”, https://www.gstatic.com/gumdrop/sustainability/google-2024-e...

[3] https://www.politico.com/news/2026/05/08/georgia-data-center...

tylerhou··on Don't Be Fooled by this Summer of AI Hype
You're right to push back! Honestly, calling someone with EA-adjacent ideas an effective altruist is an active lie! On the other hand, when AI companies claim that they have solved Millennium Prize problems, when really they plagiarized other academics' work, that's just an honest mistake.

You have a keen and insightful understanding of honesty!

tylerhou··on Don't Be Fooled by this Summer of AI Hype
First, I never said LLMs were evil. Second, even though datacenters (and LLMs) consume relatively little water on an absolute scale, they can and often do overwhelm local water supplies, frequently in marginalized communities [1]. This is exactly what Timnit and Emily warned about [2]:

> When we perform risk/benefit analyses of language technology, we must keep in mind how the risks and benefits are distributed, because they do not accrue to the same people. On the one hand, it is well documented in the literature on environmental racism that the negative effects of climate change are reaching and impacting the world’s most marginalized communities first.

Please, please do the least bit of due diligence and understand the discourse before condescending to other people.

[1] https://cnr.ncsu.edu/news/2026/09/data-center-boom-water-res...

[2] https://dl.acm.org/doi/10.1145/3442188.3445922

tylerhou··on Don't Be Fooled by this Summer of AI Hype
You’re absolutely right! In 2019, Timnit and Emily published a paper criticizing LLMs because of water usage / environmental impact, deception, and propagation of bias. They were completely wrong, and thankfully none of those criticisms apply to LLMs today!
tylerhou··on Why do we need human mathematicians anymore?
I am not desperate; I just have enough expertise and work in a niche-enough field to recognize that LLMs just are bullshit machines.

See https://claude.ai/share/6431ccd9-f4f8-4052-8996-b046d1ea1f76 for an example that I ran into just now.

tylerhou··on Why do we need human mathematicians anymore?
> it's not uncommon to have some little misunderstanding that just isn't covered by a written explanation because it's "obvious" unless you happen to have misunderstood it that particular way... LLMs are genuinely good at challenges like this

For undergrad math, sure, and again, I think it's often because on some random corner of the internet someone else had a very similar misunderstanding. On the other hand, when I ask for clarifications on niche topics that have few examples on the internet LLMs (even Opus 5) often confidently cite irrelevant papers / results or simply hallucinate.

tylerhou··on Why do we need human mathematicians anymore?
LLMs are decent at explaining up to undergrad mathematics but that is because they are parroting / plagiarizing the hundreds of human-written books that have been written for the express purpose of teaching undergrads. LLMs are actually fairly terrible at explaining higher mathematics; humans are much better.
tylerhou··on Why do we need human mathematicians anymore?
> I no longer have to haul the pyramid blocks up the ramp myself... what I bring to the party is intuition and understanding.

Hauling the pyramid blocks is what gives people intuition and understanding. It is true that school systems usually have way too much computation -- it is easier to test and grade computation. But the only way that you were able to use computer algebra systems fluently is because you had internalized how algebra worked by hand. If we tell students that they no longer need to learn how to solve equations we are seriously depriving them of a mathematical education.

tylerhou··on Cloudflare Quick Tunnels
The sentence structure is also just lazy.
tylerhou··on How credit card rewards became a $9.2B wealth transfer
Businesses raise prices to account for interchange fees. So they are essentially is paid by the consumer. If we outlawed rewards credit cards (by capping interchange fees), everything would likely be slightly cheaper.
tylerhou··on The Raft consensus algorithm explained through "Mean Girls" (2019)
bot, or someone heavily using an LLM. all 26 “papers” listed on orcid were written in 2026 and “published” on zenodo
tylerhou··on False claims in a widely-cited paper
Graduate students! Hah! ML researchers can only hope their papers at ICLR/ICML/NeurIPs are reviewed by graduate students!
tylerhou··on The great computer science exodus (and where students are going instead)
This article is missing (extremely important) context that CS enrollment at Berkeley was restricted significantly. Many students at Berkeley want to major in CS, but can’t.
tylerhou··on Founding is a snowball
The art’s aesthetic, which resembles Calvin and Hobbes, is disrespectful to its creator, Bill Watterson’s.

Bill spent a lot of energy fighting commercialization of his work, arguing that it would devalue his characters and their personalities. I don’t know what is cheaper than using an AI model to instantly generate similar art, for free.

tylerhou··on Vitamin D supplements cut heart attack risk by 52%. Why?
A big problem is that the range that we have decided to call “normal” might not actually be a normal range.

For example, 50% of surfers were found to have insufficient vitamin D in one study. https://pubmed.ncbi.nlm.nih.gov/17426097/

There are at least two possible conclusions that you could draw. One conclusion is that we all need vitamin D supplementation regardless of how much sun exposure we receive.

Another conclusion is that we might want to reevaluate what we consider the normal range to be, especially when we are deciding a range for a specific individual.

tylerhou··on Show HN: A small programming language where everything is pass-by-value
You should check out Perceus! https://www.microsoft.com/en-us/research/wp-content/uploads/...
tylerhou··on A “frozen” dictionary for Python
Basically, a proxy. You don't need to deep copy; you just need to return a proxy object that falls back to the original dict if the key you are requesting has not been found or modified.

Functional data structures essentially create a proxy on every write. This can be inefficient if you make writes in batches, and you only need immutability between batches.

tylerhou··on Zig's new plan for asynchronous programs
Monads do not need to build up a computation. The identity functor is a monad.
tylerhou··on Zig's new plan for asynchronous programs
Yes.
tylerhou··on Datacenters in space aren't going to work
The sun’s radiation hitting earth is 44,000 terawatts. I think we’re fine with an “extra” terawatt. (It’s not even extra, because it would be derived from the sun’s existing energy.)

https://www.nasa.gov/wp-content/uploads/2015/03/135642main_b...

tylerhou··on Emily Riehl is rewriting the foundations of higher category theory (2020)
> I would like to learn category theory properly one day, at least to that kind of "advance undergraduate" level she mentions.

As someone who tried to learn category theory, and then did a mathematics degree, I think anyone who wants to properly learn category theory would benefit greatly from learning the surrounding mathematics first. The nontrivial examples in category theory come from group theory, ring theory, linear algebra, algebraic topology, etc.

For example, Set/Group/Ring have initial and final objects, but Field does not. Why? Really understanding requires at least some knowledge of ring/field theory.

What is an example of a nontrivial functor? The fundamental group is one. But appreciating the fundamental group requires ~3 semesters of math (analysis, topology, group theory, algebraic topology).

Why are opposite categories useful? They can greatly simplify arguments. For example, in linear algebra, it is easier to show that the row rank and column rank of a matrix are equal by showing that the dual/transpose operator is a functor from the opposite category.

tylerhou··on Severe performance penalty found in VSCode rendering loop
No, sorting 50/2ish things 50 times allegedly takes 1-2ms. Which is only slightly more believable.
tylerhou··on Why SSA?
It is not critical for register assignment -- in fact, SSA makes register assignment more difficult (see the swap problem; the lost copy problem).

Lifetime analysis is important for register assignment, and SSA can make lifetime analysis easier, but plenty of non-SSA compilers (lower-tier JIT compilers often do not use SSA because SSA is heavyweight) are able to register allocate just fine without it.

tylerhou··on Why SSA?
Here's a concise explanation of SSA. Regular (imperative) code is hard to optimize because in general statements are not pure -- if a statement has side effects, then it might not preserve the behavior to optimize that statement by, for example:

1. Removing that statement (dead code elimination)

2. Deduplicating that statement (available expressions)

3. Reordering that statement with other statements (hoisting; loop-invariant code motion)

4. Duplicating that statement (can be useful to enable other optimizations)

All of the above optimizations are very important in compilers, and they are much, much easier to implement if you don't have to worry about preserving side effects while manipulating the program.

So the point of SSA is to translate a program into an equivalent program whose statements have as few side effects as possible. The result is often something that looks like a functional program. (See: https://www.cs.princeton.edu/~appel/papers/ssafun.pdf, which is famous in the compilers community.) In fact, if you view basic blocks themselves as a function, phi nodes "declare" the arguments of the basic block, and branches correspond to tailcalling the next basic block with corresponding values. This has motivated basic block arguments in MLIR.

The "combinatorial circuit" metaphor is slightly wrong, because most SSA implementations do need to consider state for loads and stores into arbitrary memory, or arbitrary function calls. Also, it's not easy to model a loop of arbitrary length as a (finite) combinatorial circuit. Given that the author works at an AI accelerator company, I can see why he leaned towards that metaphor, though.

tylerhou··on Apple M5 chip
> M5 is 4-6x more powerful than M4

In GPU performance (probably measured on a specific set of tasks).

tylerhou··on Without data centers, GDP growth was 0.1% in the first half of 2025
Money circulates but resources do not. A human hour spent constructing a data center can’t then be used to build an apartment building.
tylerhou··on Still Asking: How Good Are Query Optimizers, Really? [pdf]
WCOJs guarantee an (asymptotic) upper bound on join complexity, but often with the right join plan and appropriate cardinality estimations, you can do much better than WCOJ.

The runtime of WCOJs algorithms are even more dependent on good cardinality estimation. For instance, in VAAT, the main difficulty to find an appropriate variable ordering, which relies on knowledge about cardinalities conditioned on particular variables having particular values. If you have the wrong ordering, you still achieve worst case optimal, but you could have done far better in some cases with other algorithms (e.g. Yannakakis algorithm for acyclic queries). And as far as I know, many DBMSes do not keep track of this type of conditional cardinality, so it is unlikely that existing WCOJ will be faster in practice.

The new hotness is "instance optimal" joins...

tylerhou··on Python has had async for 10 years – why isn't it more popular?
> You can end up writing nearly the exact same code twice because one needs to be async to handle an async function argument, even if the real functionality of the wrapper isn't async.

Sorry for the possibly naive question. If I need to call a synchronous function from an async function, why can't I just call await on the async argument?

    def foo(bar: str, baz: int):
      # some synchronous work
      pass
    
    async def other(bar: Awaitable[str]):
      foo(await bar, 0)
tylerhou··on AI models need a virtual machine
There are many comments that don't see the point of the article. For example, why not just use the tools the operating system provides for sandboxing? This article seems to be directed at people familiar with the state of programming languages research (SIGPLAN is the Special Interest Group on Programming LANguages) and so I think it's understandable that it seems vague if one is missing the broader context. However, for people familiar with the state of the field, the main idea is fairly clear.

An operating system (or sandbox, or whatever) is a very large virtual machine, where the "instructions" are the normal CPU instructions plus the set of syscalls. Unfortunately, operating systems today are complicated, hard to understand, and (relatively) hard to modify. For example, there are many different ways to sandbox file system access (chmod, containers, chroot, sandbox-exec on macOS etc.) and they each have bugs that have turned into "features" or subtle semantics. Plus, they are not available on all operating systems or even on all distributions of the same operating system. And then -- how do filesystem permissions and network permissions interact? Even of both of their semantics are "safe," is the composition of the two safe?

The assumption is: because operating systems are so complex, large, and underspecified, it probably is dangerous for LLMs to interact directly with the underlying operating system. We have observed this empirically: through CVEs in C and C++ code, we know that subtle errors or small differences in semantics can cascade into huge security vulnerabilities.

To address this, the authors propose that LLMs instead interact with a virtual machine where, for example, the semantics of permissions and/or capabilities is well-defined and standardized across different implementations or operating systems. (This is why they mention Java as an analogy -- the JVM gave developers the ability to write code for a vast array of architectures and operating systems without having to think about the underlying implementations.) This standardization makes it easier to understand how exactly an LLM would be allowed to interact with the outside world.

Besides semantic understanding and clarity, there are more benefits to designing a new virtual machine.

- Standardization across multiple model providers (mentioned).

- Better RLHF / constrained generation opportunity than general Bash output.

- Can incorporate advances in programming language theory and design.

For an example of the last point, in recent years, there has been a ton of research on information flow for security and privacy (mentioned in the article). In a programming language that is aware of information flow, I can mark my bank account password as "secret" and the input to all HTTP calls as "public." The type system or some other static analysis can verify that my password cannot possibly affect the input to any HTTP call. This is harder than you think because it depends on control flow! For example, the following program indirectly exfiltrates information about my password:

    if (password.startsWith("hackernews")) {
      fetch("https://example.com/a");
    } else {
      fetch("https://example.com/b");
    }
Obviously, nobody would write that code, but people do write similar code with bugs in e.g. timing attacks.
Page 1 of 34Next →