HNHacker News
TopNewBestAskShowJobs

cgearhart

2,101 karma · joined August 25, 2010

chris@gearley.com
submissionscomments
cgearhart··on Driver Ticketed for No Insurance Just Because Flock (YC 2017) Said She Didn't
This has been going on for awhile. There was a fella that got flagged by facial recognition in a casino. He was hauled back by security and arrested by the police for trespassing. Even though he gave police his license. They thought it was more likely a fake id than than the “99% match” from the software was wrong. And no accountability afterwards; no change in policy or training.
cgearhart··on Chopping up books when they're physically too big
I hadn’t seen this before, but he’s near the top of my list as an example of this. One the common characteristics I see of people with this mindset is that they have high risk tolerance. Things are replaceable, and they feel safe exploring that trade space.
cgearhart··on Chopping up books when they're physically too big
That’s mostly where I’ve wound up as an adult. I think a lot of it comes down to financial stability. I spent a lot of my own money on books when I was younger. They embodied that cost in perceived value. But I learned that they might have more value than their cost if you’re willing to make them more useful.
cgearhart··on Chopping up books when they're physically too big
We studied the book in order starting from chapter 1 in the first calc class up to the last chapters in calc 4. I was in the last class when he gave me the idea to cut it, so I cut out the last few chapters and carried a very thin volume of what I needed for the last class.
cgearhart··on Chopping up books when they're physically too big
20ish years ago I complained to my college roommate that my math book was too big. He asked if I was gonna sell it back, and I said I couldn’t—they had a new edition out, so mine was worthless. Cut it, he said.

It was something I had never considered. Not for a moment. Books were sacred to me. Cut one? Blasphemy.

But he was right. And so I cut it. And it fixed all of my complaints.

What I learned is that some folks have a different mindset; they see the world as something to be shaped around themselves, rather than seeing themselves as something that should be shaped to fit the world. That’s been a useful lesson sometimes.

cgearhart··on A 386 PC for Your RP2350
Back in the mid 90’s I got it to “run” on a 386, but it was miserable. It was basically wondering if it was hung or just really slow on every command or mouse click.
cgearhart··on Muse – Meta’s personal AI agent
Out of curiosity, do you have other ideas or things you want it to do? It just seems like they always converge on the same set of executive tasks because they’re a nail that is exceptionally well suited to this particular hammer.
cgearhart··on I spent $266 and four AI models to own my tablet. GLM-5.3 finished it in a day
I understand why “prompt kiddie” feels accurate, but I don’t think it is. Expertise is _amplified_ with LLM agents. The same $300 of tokens given to my plumber—who is an _excellent_ plumber—is unlikely to produce the same outcome.
cgearhart··on AI isn’t outthinking mathematicians, it’s out-remembering them
This is the exact opposite of what I’ve been dealing with for awhile. LLMs absolute cannot work on something without an understanding unless they can outsource the understanding to a verifier. If you’ve got an easy to check function to measure progress then “keep going” is all the prompt you need. But if you need it to figure out “I pushed the up button and it moved up and left” then it’ll find the same bug five ways without realizing it’s just one bug in the underlying math.
cgearhart··on The New AI Superpowers: Focus and Followthrough
Zero dependencies should not be the goal. It’s just a modern retelling of “not invented here” syndrome. It’s an anti-pattern because now you own maintenance on something that is not your core problem. In every case I’ve seen so far, the problem is that the person running the LLM doesn’t understand what problems the library solves and so they erroneously think the dependency and their generated solution are equivalent.
cgearhart··on The New AI Superpowers: Focus and Followthrough
This is exactly what I’m talking about. I’m seeing the same pattern. And the most useful proxy I’ve found so far for “we don’t understand the problem” is taking zero dependencies. It means that you haven’t bothered to understand what the library does and you think it can be replaced with something AI one-shots.
cgearhart··on The New AI Superpowers: Focus and Followthrough
I’m seeing kinda the opposite. At some point it becomes obvious what to build and there’s no friction to building it anymore. So you get a flood of low effort copies of the same thing and no one wants to depend on anyone else, and the old value function was that ownership matters. I think the new value function needs to be much more focused on outcomes—and ownership isn’t an outcome.
cgearhart··on The New AI Superpowers: Focus and Followthrough
This seems very related to a trend I’m seeing as my company goes all in on AI: everyone thinks that every problem is “a couple hours” with AI now, and they all want zero external dependencies because they can move faster alone. As a result, we’re now in an even worse “yet-another-…” age where everyone has built approximately the same (but somehow incompatible) versions of all the same beginner-level software, and (ironically) while they want no external dependencies they’re also pushing for org-level mandates to require everyone else to use their solution. Meanwhile, no one wants to do the slow/bottleneck part that cant easily be automated or scaled; they just throw an “agent” at it and call it done—but there’s nothing _there_. You can trust the agent on easy tasks and you can’t trust it on hard ones, but you can’t tell which ones are easy or hard. Improvements in foundation model tech move thresholds of the problem but can’t eliminate it.

Long story short, I think we’re in a phase where the organizational value function is lagging behind the tech. A “proof of concept” used to be correlated with “proof of work” and some amount of domain understanding, but I think now what we need is a focus on “proof of understanding” or else you’re probably just wasting tokens on a baby version of the problem. A decent proxy right now is that if you have zero external dependencies then your solution is probably a toy.

cgearhart··on Steam Machine: Between 12k and 15k Units Sold per week
I like my steam deck and got tired of waiting for the steam machine. I asked one of the AIs for something similar and it told me to install Bazzite. Took me an hour and I got the Steam-like UI that I wanted. I did not explore other options because my problem was solved.
cgearhart··on Steam Machine: Between 12k and 15k Units Sold per week
A few months ago I got tired of waiting for the steam machine and built my own. Geekom box on sale (note: would NOT buy from them again) and then a quick hour or so to get Bazzite running. The hardest part was purchasing a thumb drive (had 3 in a row fail to deliver from amazon—that’s never happened to me before). Despite all that, Bazzite has been amazing. And for the kind of gaming I do, this little machine is more than enough. The Steam machine is likely overkill for me, honestly.
cgearhart··on A software engineering interview question I like: computing the median
We used to use this, but it was a broader conversation around tradeoffs to meet different constraints. If the expected array is small, then sort + index is probably fine. If it’s big (bigger than main memory?) and latency is the most important then maybe you want median-of-medians. If it’s a stream and you want to keep memory fixed then you might want a sketching algorithm. If I suggest that we can bound the error of the median estimate with constant additional space and the same complexity, would you believe me? (Just track the mean and standard deviation.)

Honestly, when I ran this interview I didn’t care much about the specifics of what you memorized beforehand. I care if you can read and write code a bit. I care more whether we can have a productive conversation. If you learn something new from me or the problem, how does that look and feel? If I make a mistake, how do you react? Are we able to communicate technical ideas to each other? Are we able to productively work through conflict?

We’re not computing many medians day-to-day, but we’re doing all those other things constantly.

cgearhart··on Big Tech Has Suddenly Flipped on the AI Jobs Wipeout Scenario
My read has been that a lot of leaders were trying to drive “being early” as the catalyst for future success. At the complexity scale of big orgs you’re mostly fiddling with the incentives that the system self-aligns toward. Firing a bunch of people does create an incentive to use AI, if you think it’ll help.

The more pernicious effect I’ve been seeing is that we’re living in the golden age of LLMs, but eventually that’ll fade. Tokens are subsidized and cheap, model capabilities leap forward regularly, and there’s competition driving it all. But even now there’s stories about frontier models suddenly becoming less capable, or providers switching to usage-based billing, and new model releases feel a bit more sluggish and less dramatic. (Fable/Mythos notwithstanding.)

Eventually the models are going to settle into a rut of being just “good enough” to earn a living rather than all this hoopla. A lot of people will be re-hired. And we’ll do it all again for the next wave.

cgearhart··on New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]
Yes, that’s what I think at this point. There is no effect of the study group except as a support group. (That’s all it was for me when I was a student and joined the self-organized study group.)
cgearhart··on New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]
I used to TA a graduate level CS math class at Georgia Tech. We regularly saw that the students who self-organized study groups did dramatically better in the course than average. One semester they told us to put everyone in study groups to see if it helped. The effect disappeared. Turns out that it was the self-selection of the most engaged students into a small group that mattered, not the study group itself.
cgearhart··on What it feels like to work with Mythos
I’m starting to realize that LLMs are really good at building low-stakes projects. Your questions mostly presume that the stakes are higher. The software will last a long time; the requirements will evolve; we can’t tolerate mistakes; etc.

The trick to getting good at using LLMs for software is to learn how to make _all_ projects low-stakes.

cgearhart··on Siri AI
I just tested this myself. I wrote “flip the reduce white point toggle accessibility option in the settings app” and it worked perfectly. Run once to set it and run again to disable it.
cgearhart··on Refusal in Language Models Is Mediated by a Single Direction
Spreading out the refusal encoding shouldn’t be effective as a countermeasure. Even if it were smeared across the vector space, as long as it’s in a subspace that doesn’t span the entire domain then you should be able to either null out the entire subspace spanned by the refusals or run some kind of clustering on the generated samples to identify the dominant directions and nullify all of them. I think an effective defense would either need to spread them to span the entire domain—basically “encrypting” the refusal so it can hide anywhere, or you’d need a very large number of independent refusal circuits in the model so that simple hacks in the vectors themselves don’t matter, or maybe you could make other circuits depend on proper functioning of the refusal circuits… hmmm… is that along the lines of what you’re saying they’ve done already? (Any references or links to modern techniques?)
cgearhart··on There Will Be a Scientific Theory of Deep Learning
A much earlier major win for deep learning was AlexNet for image recognition in 2012. It dominated the competition and within a couple years it was effectively the only way to do image tasks. I think it was Jeremy Howard who wrote a paper around 2017 wondering when we’d get a transfer learning approach that worked as well for NLP as convnets did for images. The attention paper that year didn’t immediately dominate. The hardware wasn’t good enough and there wasn’t consensus on belief that scale would solve everything. It took like five more years before GPT3 took off and started this current wave.

I also think you might be discounting exactly how much compute is used to train these monsters. A single 1ghz processor would take about 100,000,000 years to train something in this class. Even with on the order of 25k GPUs training GPT3 size models takes a couple months. The anemic RAM on GPUs a decade ago (I think we had k80 GPUs with 12GB vs 100’s of GBs on H100/H200 today) and it was actually completely impossible to train a large transformer model prior to the early 2020s.

I’m even reminded how much gamers complained in the late 2010s about GPU prices skyrocketing because of ML use.

cgearhart··on Gender Equality and Work
I agree, it would be nice if we could prioritize basic human needs rather than treating them like burdens caused by bad luck or poor choices.
cgearhart··on FreeCAD v1.1
Slightly unrelated to this story, but I’m curious if anyone has good resources for learning FreeCAD. I have quite a lot of experience with SolidWorks, AutoCAD, OnShape, and similar software, but FreeCAD has always been hard for me to pick up.
cgearhart··on An opinionated take on how to do important research that matters
Eh. I think my point is that the OP is presented as a “how to” (literally: “how to do important research”) and then it immediately dodges the question by saying “have good taste”. That does not help anyone do important research or improve the quality of the research they do; it’s a cop out.

If I wrote about “how to paint great art” or “how to cook great meals” or “how to build great things” then it would be silly to say “have good taste”—even if that’s part of the answer. It won’t help anyone else to improve in any of those endeavors.

cgearhart··on An opinionated take on how to do important research that matters
That seems even less actionable, and somewhat misaligned with the OP article. “Taste” implies an ability to distinguish between a good example and a bad one. If it’s only recognizable in retrospect then it’s just another name for survivorship bias.
cgearhart··on An opinionated take on how to do important research that matters
I often find this kind of advice too vague to really be useful. “Have taste” in the problems you work on isn’t very actionable. (Unless perhaps you list examples of good and bad taste.)

I’ll admit that I may just be immature at research as almost all my experience has either been attempting to replicate research or to put it into practice in production systems.

cgearhart··on Qwen3-Coder-Next
Any notes on the problems with MLX caching? I’ve experimented with local models on my MacBook and there’s usually a good speedup from MLX, but I wasn’t aware there’s an issue with prompt caching. Is it from MLX itself or LMstudio/mlx-lm/etc?
cgearhart··on Anthropic's original take home assignment open sourced
I think this is the actual “bitter lesson”—the scalable solution (letting LLMs bang against the problem nonstop) will eventually far outperform human effort. There will come a point—whether sooner or later—where this’ll be the expected norm for handling such problems. I think the only question is whether there is any distinction between problems like this (clearly defined with a verifiable outcome) vs the space of all interesting computer programs. (At the moment I think there’s space between them. TBD.)
Page 1 of 15Next →