HNHacker News
TopNewBestAskShowJobs

Taek

8,764 karma · joined July 17, 2013

submissionscomments
Taek··on Suits Are Better Tech Than Modern Clothes
The argument he makes is that suits are designed to be easily tailored, and tailored for many different aspects of the body.

Modern clothes aren't as easy to tailor, because the seams aren't placed well for tailoring, and the patterns / fabric shape also doesn't tailor well.

Taek··on AMD's random number generator can't generate a 0?
Yes, numerous times. Here are some famous ones:

  Android SecureRandom (2013)
https://android-developers.googleblog.com/2013/08/some-secur...

  CryptoJS / Ill Bloom (2026)
https://illbloom.org/articles/cryptojs-vulnerability/

  Trust Wallet Browser Extension (2023)
https://www.ledger.com/blog/funds-of-every-wallet-created-wi...

  Libbitcoin / Milk Sad (2023)
https://milksad.info/disclosure.html

  Trust Wallet iOS / Trezor Library
https://secbit.io/blog/en/2024/01/19/trust-wallets-fomo3d-su...
Taek··on 'We hacked the FBI:' Hackers say they have data on all FBI employees
Google seems capable
Taek··on AMD's random number generator can't generate a 0?
The reason that you get 3-4 bits of entropy per hash is because of the fundamental nature of CPUs. In addition to having considerable professional experience with cryptography, I also have considerable professional experience with hardware; hardware is fickle as hell, especially when your transistors are tens of nanometers large. Every time you flip a bit, you expend some energy, which heats up the chip, and the heat changes the timing of the next clock cycle. Chips are composed of literally billions of transistors, and each one is going to have a different temperature, because clock cycles last less than a nanosecond (well, embedded hardware is slower but the same idea still applies reliably) and that's not enough time for temperature deltas to dissipate across the chip.

Hashing is particularly chaotic because it lights up a different set of transistors on each clock cycle, which means the hotspots on the chip are being jerked around. Some transistors are going to light up 5-10 times in a row, and others are going to be idle 5-10 times in a row, and then randomly that changes. And all of this changes the number of picoseconds that it takes for a clock cycle to complete, which means that each clock cycle is genuinely going to take a different amount of time to complete, and stuff like temperature throttling is completely not at play whatsoever, because we're not talking about chip-wide temperatures, we're literally talking about temperature deltas between transistor a and transistor b.

That makes it a really wonderful source of entropy for cryptographic applications, because the CPU clock is so critical that it's almost never buggy (especially relative to other components that provide entropy), it's also almost impossible to manipulate reliably by an attacker (unless the attacker has an exploit that allows them to set the value of the clock directly - which is possible, but it's a very narrow surface area relative to other entropy sources), and you can completely take advantage of this entropy entirely in userspace, which once again heavily minimizes attack surface area and exposure to bugs.

I have searched far and wide for a CPU that does not reliably generate entropy using the iterated-hashing-against-the-clock method, and I have not found a single example of a CPU that consistently takes the same amount of time to complete a hash. And the reason isn't implementation, the physics of CPUs simply insist on introducing entropy when trying to repeatedly hash something quickly.

Taek··on AMD's random number generator can't generate a 0?
I ran it 500,000 times, discarding the 10% most entropic results ... in the hopes of arriving at a relatively conservative estimate for the amount of entropy you actually get from each iteration. Here's the prompt I used to generate the code: https://chatgpt.com/share/6ab2df4a-7f94-83ea-aecf-1bb57c4838...

And here are the results of running that code:

  === No hashing ===
  Clock resolution: 0.000000001 seconds
  Clock reads:                       500,000
  Second-difference outcomes:        499,998
  Retained outcomes:                 449,998 (90.000%)
  Average Shannon information:       1.755579 bits/retained outcome
  Marginal min-entropy estimate:      1.339460 bits/retained outcome
  Lag-1 conditional min-entropy:      0.960079 bits/retained adjacent outcome
  Conservative descriptive proxy:    0.960079 bits/retained outcome
  Proxy scaled per clock iteration:  0.864067 bits/iteration
  These are empirical timing statistics, not a proven entropy rate.

  === One SHA-256 between clock reads ===
  Clock resolution: 0.000000001 seconds
  Clock reads:                       500,000
  Second-difference outcomes:        499,998
  Retained outcomes:                 449,998 (90.000%)
  Average Shannon information:       4.205076 bits/retained outcome
  Marginal min-entropy estimate:      3.610848 bits/retained outcome
  Lag-1 conditional min-entropy:      3.351217 bits/retained adjacent outcome
  Conservative descriptive proxy:    3.351217 bits/retained outcome
  Proxy scaled per clock iteration:  3.016082 bits/iteration
  These are empirical timing statistics, not a proven entropy rate.
------------

As GPT helpfully points out, this isn't a proven guarantee, but a reasonable estimate is somewhere between 3 and 4 bits of entropy per hash. That means 50 is actually enough, though if you want to be conservative I don't think there's any harm in doing 500 or even 5,000 instead of 50. And, if you are going to be using this in a hostile environment, it doesn't hurt to also add a fortuna-like accumulator that resets your entropy every once in a while.

I said this in another reply as well, but the reason that you get 3-4 bits of entropy per hash is because of the fundamental nature of CPUs. In addition to having considerable professional experience with cryptography, I also have considerable professional experience with hardware; hardware is fickle as hell, especially when your transistors are tens of nanometers large. Every time you flip a bit, you expend some energy, which heats up the chip, and the heat changes the timing of the next clock cycle. Chips are composed of literally billions of transistors, and each one is going to have a different temperature, because clock cycles last less than a nanosecond (well, embedded hardware is slower but the same idea still applies reliably) and that's not enough time for temperature deltas to dissipate across the chip.

Hashing is particularly chaotic because it lights up a different set of transistors on each clock cycle, which means the hotspots on the chip are being jerked around. Some transistors are going to light up 5-10 times in a row, and others are going to be idle 5-10 times in a row, and then randomly that changes. And all of this changes the number of picoseconds that it takes for a clock cycle to complete, which means that each clock cycle is genuinely going to take a different amount of time to complete, and stuff like temperature throttling is completely not at play whatsoever, because we're not talking about chip-wide temperatures, we're literally talking about temperature deltas between transistor a and transistor b.

That makes it a really wonderful source of entropy for cryptographic applications, because the CPU clock is so critical that it's almost never buggy (especially relative to other components that provide entropy), it's also almost impossible to manipulate reliably by an attacker (unless the attacker has an exploit that allows them to set the value of the clock directly - which is possible, but it's a very narrow surface area relative to other entropy sources), and you can completely take advantage of this entropy entirely in userspace, which once again heavily minimizes attack surface area and exposure to bugs.

Taek··on AMD's random number generator can't generate a 0?
I know that there's a really strong culture in the software world around downvoting anything that looks or smells like "hand-rolled cryptography", but this is my actual profession and specialization within the software world, and most of what I'm seeing in this thread is knee-jerk reactions to an unexpected technique rather than careful intellectual commentary and consideration of the merits of the technique.

I am happy to have a discussion with you at the deepest technical levels of applied cryptography, this is not something I blindly made up on my own. I'm well studied in the field and can readily defend this technique.

Taek··on AMD's random number generator can't generate a 0?
That's exactly the challenge though: "any good crypto library" - there is a long history of meaningful security breached (like stolen crypto tokens) due to bugs in an upstream library, especially when using things like embedded code, alternative operating systems, newer programming languages, etc.

The value of the iterated hashing method is that it is dead simple and has little dependency on potentially buggy upstream code; it works even in very lightweight environments designed by engineers with no experience in security.

Taek··on AMD's random number generator can't generate a 0?
Yes but why introduce complexity and room for error when something that's extremely basic is also sufficient?

The point here is to eliminate surface area for mistakes, and an XOF has a much larger and more complex implementation than iterated hashing against a timer.

Taek··on AMD's random number generator can't generate a 0?
Hold on I have to go edit the rest of my responses because I just assumed you wrote the code correctly; you did not.

You are not hashing between calls to the timer. The sha256 hash itself is responsible for doing physical things to the chip (heating up some parts unevenly during the hashing computation) which introduces meaningful entropy between calls to the current time.

You can't just do calls to clock_gettime(), you have do an actual sequential sha256() call between them. Please run this code again and tell me what results you get.

Taek··on AMD's random number generator can't generate a 0?
Actually, it gives you as much entropy as you need, just increase the iterations. That guy's output is shockingly consistent, so to be conservative maybe we say 0.2 bits of entropy per iteration. So just do 1000 iterations. That's still only going to take a few milliseconds even on embedded hardware.

EDIT: I reviewed his code, and he's not hashing between calls to check the clock; the hash call itself causes the CPU to heat up in arbitrary ways which changes the timing between hashes and introduces more entropy; removing that call basically entirely defeats the idea behind the technique, these results are fully invalid.

Taek··on AMD's random number generator can't generate a 0?
The strength in this method is that it has the littlest possible surface area for upstream bugs to compromise your final entropy. Because, in the applied world, upstream bugs in "secure" system RNGs have been the cause of stolen crypto and other critical security compromises on numerous occasions.

And, I agree that if the system is compromised to the level that the attacker can control the output of the timer, it's probably compromised to the level that the attacker can just read your generated entropy straight from memory.

The point here is not to be fast, it's to be protected against implementation bugs on systems that weren't designed by security professionals.

Taek··on AMD's random number generator can't generate a 0?
The reason I roll entropy in userspace is because there's a very long history of "cryptographic" libraries getting it wrong (see the parent article for an example). Crypto tokens stolen because the underlying call to the web browser entropy only had 32 bits of actual randomness. Crypto tokens stolen because the underlying embedded system (like cold card) turned off some security critical features to improve performance and power.

Pretty much the only thing you can control when shipping software to many devices is that it runs on a physical CPU and has a timer. Every other RNG assumption over the decades has shown that sometimes someone upstream gets something catastrophically incorrect.

Taek··on AMD's random number generator can't generate a 0?
You don't need 256 bits of entropy, you only need 128.

I have tested this method on over 100 different CPUs and I have never seen such consistent output. I'm genuinely surprised to see that you only hit 92 bits of entropy, but that can trivially be fixed by doing 10x the iterations. 500 iterations is still going to put you under a millisecond of cost.

And, for what it's worth, code I've actually shipped has combined the above technique with Fortuna, and has typically targeted 2000 bits of entropy rather than 128 (for security buffer).

EDIT: I reviewed his code, and he's not hashing between calls to check the clock; the hash call itself causes the CPU to heat up in arbitrary ways which changes the timing between hashes and introduces more entropy; removing that call basically entirely defeats the idea behind the technique, these results are fully invalid.

---

I updated the code to insert the hash call, this is what I got for his original code on my machine, and the updated code with hashing on my machine (and the difference is cryptographically meaningful):

  === Original C — no hashing ===
  Clock resolution: 0.000000001
  Deltas (ns):   50   34   19   19   13   13   13   13   13   14   13   13   13   13   13   14   13   13   14   12   13   14   13   13   13   14   13   13   14   12   13   14   13   13   14   12   13   14   13   14   13   12   13   14   14   13   13   13   13
  Deltas of deltas:   -16  -15    0   -6    0    0    0    0    1   -1    0    0    0    0    1   -1    0    1   -2    1    1   -1    0    0    1   -1    0    1   -2    1    1   -1    0    1   -2    1    1   -1    1   -1   -1    1    1    0   -1    0    0    0
  Maximum entropy: 90

  === C with SHA-256 between clock reads ===
  Clock resolution: 0.000000001
  Deltas (ns): 756852 1287  542  470  472  445  442  436  434  439  488  435  433  434  440  439  439  435  432  433  435  432  433  433  429  433  453  441  437  437  431  433  432  430  431  438  436  434  431  433  435  436  435  433  430  436  435  437  428
  Deltas of deltas:  -755565 -745  -72    2  -27   -3   -6   -2    5   49  -53   -2    1    6   -1    0   -4   -3    1    2   -3    1    0   -4    4   20  -12   -4    0   -6    2   -1   -2    1    7   -2   -2   -3    2    2    1   -1   -2   -3    6   -1    2   -9
  Maximum entropy: 188
Taek··on AMD's random number generator can't generate a 0?
Unfortunately you are not correct, and djb explains it quite well here:

https://blog.cr.yp.to/20140205-entropy.html

TL;DR adding a compromised source of entropy to a pool of already secure sources of entropy can catastrophically compromise the final result.

It's better to source entropy from a smaller number of harder-to-compromise sources. That's why I like the iterated hashes method; the security surface area is both very small and highly likely to be well tested.

Taek··on AMD's random number generator can't generate a 0?
You can effectively achieve the same result with this simple operation:

  hash = sha256(current_time());
  for i := 0; i < n; i++ {
      hash = sha256(hash.append(current_time()))
  }

This is because the number of nanoseconds between hashes is actually itself variable, and this is true for physics reasons that are basically beyond the control of any attacker trying to manipulate your entropy. If your time() function has a resolution of nanoseconds, you only need your loop to iterate about 50 times to get a cryptographically secure amount of entropy. If your time() function has a resolution of milliseconds, you need to let this run for more like 20 milliseconds, and if your time() function has a resolution of seconds you need to let it run for more like 5 seconds.

The reason I like doing it this way is that it happens entirely in userspace, it's genuinely a secure method of generating entropy, and it has no dependencies on potentially buggy firmware or microcode outside of the time() call, which is both fairly narrow, fairly heavily used (meaning a bug is likely to be discovered during testing, as the implementation is likely heavily scrutinized), and also fairly easy to test independently - just look at the number of nanoseconds that elapse at each consecutive call to sha256(current_time()) and verify that there's some statistical variance. The above suggestions are assuming about 2.5 bits of variance between calls, meaning there should be a range of at least 20 nanoseconds between your slowest and fastest hash call. This has been true on every CPU I've ever measured, including microcontrollers.

Taek··on Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
I wrote this elsewhere but I felt it was worth saying again. "9x smaller" is an idiom, much like "spill the beans" or "it costs an arm and a leg".

Idioms don't have to make literal sense or be linguistically/mathematically correct to be useful. All that matters is that other people know exactly what you mean when you say it.

And, pretty much universally, if I tell someone "the compressed file is 10x smaller than the original", they are going to know what I mean is that the byte size is 10% of the size of the original.

That makes it an idiom that is perfectly okay for everyday use.

Taek··on Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
It's an idiom, much like "you are pulling my hair". Are you actually pulling my hair? Of course not; taken literally, many English phrases make no sense. But they are common elements of the language and everyone understands them, so they aren't problematic.

This is the same. Taken literally, "9x smaller" is nonsense, but everyone who hears that phrase knows exactly what mathematical operation you are referring to, thus it's a totally acceptable way to express that you mean to say 11.11% as large.

Taek··on Felony charges for citizen deleting phone data at US Border
What parts of the United States can you visit which have a >10% chance of being kidnapped or murdered? That means for every 100 visitors, 10 people don't come home. I honestly don't think that applies to a single neighborhood in the entire country. If murder/kidnapping rates get remotely close to that high, the FBI steps in.

I also can't think of any US cities where neighborhoods have military grade armed security. Sure, there are places where every local business has an armed guard, but that's not really the same as hiring a trained private militia. The armed guards are for protecting against petty theft, not for protecting against organized crime.

On point four, I'm not sure if there are places in the US where minorities need to maintain relationships social relationships with cops as a survival mechanic, but it certainly doesn't apply to most cities, and I don't think it applies to anyone who is white.

On point five, I don't think you understand. There are parts of the world where having white skin will get you, quite literally, reminders every 15 minutes "hey it's really not safe for you here, do you want to hang out inside my shop while I call you a taxi?" No part of the US is like that for travelers. I know there are occasionally ICE raids that make the news, but "hey you strictly cannot be outside without a local chaperone" is just not a thing in the US.

Taek··on Felony charges for citizen deleting phone data at US Border
I get that the political system could be better, but having been to places that actually are not governed by the rule of law, I can assure you that the US is doing quite well on that front. When there is actually no rule of law, consequences include:

+ entire regions / areas where merely visiting those areas invites a highly non-trivial (think, more than 10% chance) chance of being kidnapped or murdered

+ every neighborhood and business has substantial, often military-grade private security

+ if credit exists at all, it exists outside of any formal banking structure and will have interest rates that are north of 30% APR, sometimes north of 100% APR. I've genuinely seen interest rates on credit as high as 2% *per day*, and these are rates that the local population is willing to pay for certain short term expenses (like food)

+ families that maintain good social relationships with the local police/militants/whoever-has-guns live substantially better lives than people without good social connections to the local authorities.

+ Travelers are told on repeat: "it's really not safe here for non-locals, you should stay inside and also reconsider being in this part of the world at all"

+ If the travelers are there for business reasons, they are probably assigned 24/7 armed guards (as many as 4 guards per traveler, each guard carrying full-auto weapons) by the locals, provided for free.

And while the US maybe has a neighborhood here or there which might be like this, every part of every major city in the country has more rule of law than the above.

Taek··on Felony charges for citizen deleting phone data at US Border
Definitely booting into a decoy system beats an automatic wipe.
Taek··on Felony charges for citizen deleting phone data at US Border
I do like the idea of forcing due process to access an encryption key. I'm not sure if that idea is compatible with current law, but it seems just at the very least.
Taek··on AI;DR (AI; Didn't Read)
I just wrote a utility to rip all comments out of the code. Now the code is fully uncommented and it has saved lots of input tokens and also lots of meandering because the model is no longer getting stuck on bad ideas it told itself about.
Taek··on Grok Bot
A job faire might serve you well!
Taek··on Grok Bot
I have a hard time looking at the world around me and thinking "close to no value"
Taek··on Magnitude 7.4 Earthquake – 5 km S of San José del Palmar, Colombia
Some sounds are genetically more intense than others, like shrill screeching sorts of noises. So not quite as bad as the euphemism treadmill, because there is biology that makes certain sounds fundamentally less terrifying than others.
Taek··on Magnitude 7.4 Earthquake – 5 km S of San José del Palmar, Colombia
Was on the 6th floor of a tower. Shaking lasted almost 2 minutes. There was no damage as far as I could tell (Medellín) but they still evacuated everyone to inspect the building. City is in a bit of chaos, people are fine but unsettled, and the comms lines are clogged with people presumably calling their families.

Was a bit tense to see my phone alerts going off over and over, continually increasing the estimate of the earthquake strength, followed by the shaking getting worse.

For people who don't know, those phone alerts give you literally like 5 seconds of heads up. By the time I finished reading the first message, the shaking had already started. And I got the message while I was holding my phone and using it.

Taek··on Herdr is joining Y Combinator. The runtime stays open
This title is killing me. "The runtime stays open" is such a heavy, attention grabbing sentence, it's all I can stare at on the front page. It's the type of writing that makes me hate working with LLMs, because it's so powerful at monopolizing my attention, and makes it difficult for me to focus on any of the other words on the page (whether the front page, or the comments page).
Taek··on If you're a button, you have one job
That's a different bad UX pattern. If a button has already rendered in a certain location, a new button shouldn't replace it without first giving the user ample warning that a material change is about to happen.
Taek··on If you're a button, you have one job
I wish software apps had "tape-out rules" the way that computer chips do. Basically, when you design a computer chip, a program reviews the design and compares it against something like 300 pages of rules with stuff like "wires of X metal and Y metal can't be within Z distance of each other".

We could make something similar for UX. Just a bunch of design pattern constraints that throw flags if you try to ship something with well established UX warts.

Taek··on 60% Fable cost cut by converting code to images and having the model OCR it
I don't think it's that bad, if I recall correctly it's about 8 kilobytes per token, and a token can be 3-4 characters so you're talking ~2 kilobytes per character.

An image token I recall is something like 16x16, so you get 32 bytes of overhead per pixel. And a character is minimally like 20 pixels including the whitespace, so you've jumped from 4 characters per token to maybe 12.

So 3x savings... which actually maps pretty closely to 60% savings.

Page 1 of 34Next →