HNHacker News
TopNewBestAskShowJobs

ekidd

11,536 karma · joined June 5, 2010

Blog: http://www.randomhacks.net/ Mastodon: https://mastodon.xyz/@emk

I write a lot of code in Rust, Python and TypeScript. Some of it is open source. Somehow I get paid for this.

I've done startups and consulting, but I'm pretty busy right now.

submissionscomments
ekidd··on I rewrote PostHog's SQL parser, 70x faster, while barely looking at the code
> May be it is that we are ingenious amd creative with tools and thats how we evolve.

And every time you use the AI to be ingenious or creative, that will be added to the training data. Then someday the AI can be ingenious and creative without you! (It might take a few more breakthroughs. But investors will literally spend trillions chasing those breakthroughs.)

The endgame here is to replace all human intelligence and labor with machines that are smarter and work cheaper. But who controls the machines?

ekidd··on GLM-5.2 – How to Run Locally
A GPU with 24GBs of RAM is mostly useful for running a very carefully squeezed Qwen3.6 27B (4-bit Unsloth quants, 8-bit K/V cache, possibly MTP, 128k context). This is a fun little model that's smart enough to do debugging, refactoring, and implementing "clean" specs that don't force it to make complicated design choices. I've seen it rip through a 9-year-old Terraform AWS config, and (without using the network) correctly identify nearly everything that would need to be upgraded or migrated for modern AWS. But if I give it some poorly conceived spec with lurking design headaches, then it goes on an endless thinking binge and ultimately fails.

Speed-wise, I don't have numbers, but it feels subjectively faster than Opus in Claude Code. YMMV.

Once you go above "a used 3090 at a decentish price", then I strongly recommend renting cloud GPUs or at least testing models using paid APIs. This allows testing your use case before spending piles of money.

ekidd··on The text in Claude Code’s “Extended Thinking” output
> If so, the thinking trace can be sort of nonsensical for a reader, though whether this is an idiosyncrasy of the model or a property of LLMs in general isn't clear to me yet.

Yes, several models think in weird jargon. Here is an example of Mythos's thinking while playing solitaire: https://www.lesswrong.com/posts/wCSEpT3dTGz4N86Wi/even-illeg...

> 7♣-removal-IS-the-prerequisite-for-10♠/9♥!!)-⟹-OVERLAP-(ii)+(iv):-{6♠ J♦ 9♥ 2♣}-=-FOUR--—-UNLESS-7♣'s-seat-8♥-...-and-2♣-drains-only-at-crack-:-⟹-2♣-celled-+-9♥-celled-simultaneously-UNAVOIDABLE-in-t8-dig--—-BREAK:-9♥

This is a small step in the direction of something called "neuralese", where the model has stopped thinking in English and is thinking in internal vector spaces. Since this gets serialized through text, it isn't quite true neuralese, but it's moving in that direction.

I mean, I'm sympathetic towards the models. My internal thought process when writing code uses lots of intermediate steps that would be hard to write out in English.

ekidd··on Human Judgment as a Specification
The difference is that a compiler is a rigorous, (nearly) determinisic, heavily tested artrifact built by expert humans. I have only encountered genuine code generation bugs in compilers twice in my career. And yes, those bugs I did trace to the assembly.

An LLM prompt, even a huge one, is an incredibly vague document that leaves out most of the edge cases. And even Fable 5 happily ignores clear instructions in its prompt.

Now, to be fair, I absolutely expect the buggy slop to win, and to drive out the people that either write their own code or at least read the output. This will, in turn, make customers less willing to spend money on software after they get burnt a few times by buggy garbage. I think this is pretty much inevitable once Fable returns. It's just too damn good at long time horizon tasks, generating far more mostly sorta working code than any human could reasonably read.

ekidd··on Human Judgment as a Specification
> Telling people “you must read all the code generated by an LLM” is definitely meaningful—but it is not at all moderate (so most people won’t do it).

I am honestly heartbroken to live in a world where reading the code is seen as an unreasonable ask by either students or by professional working programmers.

ekidd··on LLMs Are Complicated Now
Lol, no. I've always sounded like that, and there are decades of my writing online.

Also, FWIW, Pangram scores my writing as entirely human.

Claude's writing isn't easy to identify because it uses em-dashes and bulleted lists. Claude's distinctive style goes much deeper than that.

ekidd··on LLMs Are Complicated Now
Claude's writing style is at least as distinctive as any human's personal style. It has a long list of favorite words, verbal tics and common structures. On top of that, LLM writing is often bad in a very particular way: it's weak on actual things to say, but with an overheated style.

Some days, I spend over 4 hours a day reading walls of text written by Claude. If I couldn't recognize Claude's default "voice" by now, something would be wrong. It would be like a Hemingway fan not being able to recognize Hemingway. Except more so, because Claude's writing style is getting worse from version to version, descending into self parody.

On the statistical side, Pangram's model identifies AI-authored text with a 1-in-5,000 false positive rate, measured against hold-out texts from before 2022. My "ear" also agrees closely with Pangram. If I think something sounds AI written, Pangram virtually always comes back with "AI, confidence: high."

ekidd··on The AirPods Effect
This varies enormously by where you live.

I live out in the countryside. If I run into someone in the road, I will nod my head, maybe introduce myself, and maybe chat, if the other person is interested. (To be fair, I know about 80% of the people I see in the road.) This is normal behavior. Sometimes, two cars will pass each other and stop to talk.

I have also lived in the city. If a stranger wants to talk to me in the city, either they're looking for directions (happy to help!), or they are deeply confused about appropriate social behavior in crowded spaces. In the latter case, I'm lucky if the stranger-with-no-boundaries merely wants to warn me about the dangers of the lizard people. So I've learned to ignore strangers.

ekidd··on Anthropic employees accuse Trump administration of targeting them
As far as I can tell, much of Anthropic genuinely believes that someone will build an AI in the next 3-20 years that's significantly smarter than any human alive. Sounds wild, but a lot of their people have been saying this since 2018 or even earlier. I think they're true believers. Furthermore, they believe that building such an AI would be dangerous.

So their plan is:

1. We can't stop other people from building something dangerous.

2. But we can get there first.

3. If we build it, it has maybe a 15% chance of killing everyone alive. (I think that's a number I've seen Dario use before, but I may be wrong.) If OpenAI or China build it, the odds would be worse.

Obviously, if Anthropic is actually correct about (1) and (3), then nobody should allowed to build frontier AI.

People find it really hard to believe that (a) anyone believes in the possibility of dangerous AI in our lifetimes, and (b) that someone could believe what Anthropic seems to believe and then still go ahead and gamble with everyone's lives anyway.

ekidd··on Anthropic employees accuse Trump administration of targeting them
> Amazon removes guardrails from Fable, getting access to Mythos.

Amazon did not remove any "guardrails" from Fable. They created a fake, obviously insecure program. And apparently their prompt was exactly, "Fix this code." And Fable fixed the bugs.

This is something that even dinky local Chinese models running on a high-end gaming GPU can often do. Certainly Opus, GPT 5.5 and Gemini can all do this. And any high-end Chinese "near-frontier" model can do this, too.

But either (1) the administration is too clueless to know most models can do this, (2) Trump wants to be paid a bribe, (3) someone thinks Anthropic is "woke" and should therefore be destroyed by the power of the state, or possibly, if you're really cynical, (4) maybe the NSA SIGINT wants access to Mythos so they can break into everyone's computers, but they don't want you to have a model good enough to keep them out. Take your pick, I guess.

Anyway, apparently we don't do free markets or rule of law in the United States any more?

ekidd··on The computer science degree isn’t dead
COBOL was mostly outsourced to India, and it's a terrible professional path for anyone in the EU or US, and has been since the Y2K bugs got fixed at the last minute.

(And probably a bad path in India, too, but I have no data one way or the other. It's just that all the excellent Indian devs I know use almost exactly the same tech stacks I do.)

ekidd··on If Claude Fable stops helping you, you'll never know
At different size ranges:

- Qwen3.6 27B runs quite nicely on a 32GB GPU, and it's a mostly usable coding agent. The biggest difference with a frontier model is that a 27B forces you work in chunks between 100-200k tokens, and to maintain a clear understanding of how your code works. If you try to vibecode without understanding, yeah, it's going to get ugly. Also, it's better at coding than many other tasks.

- DeepSeek V4 Flash is apparently quite nice if happen to have 256GB of RAM lying around, lol. Again, not a frontier model, but antirez really likes it.

ekidd··on DeepSeek V4 Pro beats GPT-5.5 Pro on precision
The OP uses tons of typical AI turns of phrase, and Pangram classified it as AI with high confidence.

So it doesn't surprise me at all that the methodology is weak, too.

ekidd··on Social Cache Busting
This is a basic survival skill in politics, and not just for scandals.

Let's take Bernie Sanders, because he's well-known in Vermont for being happy to go off-script and actually talk to people. During my only personal conversation with him, he was delighted to discover that a small, local event actually served excellent chicken. (Apparently politicians eat a lot of rubbery chicken.)

But at that same event, Bernie was approached by a woman asking some conspiracy-tinged question. And he very gracefully deflected and changed the topic. I think that just about anyone who interacts with the public is likely to pick up some version of this skill eventually.

ekidd··on Did Claude increase bugs in rsync?
> The problem is memory that you allocated in the past, have freed, but hasn't been returned to the OS[0].

There are at least two different ways in which memory might be semantically "uninitialized":

1. The memory was provided by the OS. On modern desktop and mobile OSes, this memory will normally be zeroed automatically. 2. The memory was provided by the language's allocator. This may contain a mix of data used by previous allocations and memory that has never been touched (perhaps because previous allocations reserved it as end-of-array "capacity" that never got used). From the perspective of a language like Rust, this memory is considered uninitialized, and safe code should never be able to read it without first setting it.

In ancient C code, it makes a fair bit of sense to preemptively calloc everything. Or better, to wrap the allocator with one that zeroes on free. Though even there, you need to be careful not to expose recycled heap block headers in the middle of newly allocated objects.

My opinion for the last 30+ years has been that C is unfit for purpose, and that using it almost inevitably introduces large numbers of dire security holes. But until the last 10-15 years, there hasn't been any seriously viable alternatives.

ekidd··on Did Claude increase bugs in rsync?
Calloc is generally hardening, because it zeros out any stale memory contents left over from previous uses of the memory.

You can avoid this overhead if you use a language that forbids reading from uninitialized memory, but C is not that language.

ekidd··on Nvidia RTX Spark
It's probably more that LLM inference speed comes from having a large amount of fast RAM. And fast RAM is brutally expensive right now.

At this point, your cost-efficient options include used 3090s, "frankenrigs" using recycled data center cards, and a handful of "workstation" class cards, where the originally high margins and the long enterprise purchasing cycles have kept prices from going up too fast.

In contrast, a lot of these "personal" AI systems are basically a GPU-like core wired to larger amounts of slow RAM. Which is still semi-affordable. Generally speaking, they make for OK chatbots but extremely slow coding agents. Whereas you can run a modestly useful coding agent at reasonable speed on a 3090.

So yeah, a lot of these systems are bit scammy. But not because it's a secret conspiracy to protect data center cards. Rather, there simply isn't enough fast RAM in the entire world. So they'll flog you disappointly slow RAM instead.

TL;dr: Might be useful for some use cases, but benchmark very carefully.

ekidd··on The dead economy theory
> Why is this time different?

If it was just programming being automated, then whatever. Lots of professions have been automated and society adapts.

The underlying worry here is that current AI provides a partial automation of intelligence. The endgame for the investors and the corporations using AI is complete automation of intelligence (and manual labor, too). They want a $25,000 robot that works around the clock, and AI models that will do anything a human office worker can do for less money. Now, they don't know how to build either yet. But they'll spend every last dollar on the planet trying.

Strictly speaking, they don't even need us as customers. They can just have the robots build them yachts and mansions directly. And act as security guards.

ekidd··on Claude Opus 4.8
The Chinese models are surprisingly cheap and performant sitting under my desk. Qwen3.6 27B is nowhere near as autonomous as Opus 4.7, but it runs in 24GB of VRAM. And it's actually great for the use cases where I'm going to carefully read and understand all the code anyway.

If you want to support a team of engineers, DeepSeek V4 Flash is antirez's current favorite. And you could support a team of engineers pretty nicely for $40-50k. Which might not make sense if you're on a Claude MAX 5x plan or the old enterprise group plan with fixed price seats. But Anthropic is switching their enterprise contracts over to token-based pricing, at which point $50k is looking pretty good.

ekidd··on Citing 'severe' math deficits, UC faculty demand a return to SAT tests for STEM
As someone who hates handwriting in bluebooks, and who types constantly, yes: I think we should bring back in-class writing by hand, we should lock up cellphones for the school day, and we should proctor exams. If you're not doing this, your students will be stuck to a screen all day, pay no attention to class, and use ChatGPT under the desk to cheat.
ekidd··on How long until AI automates all cognitive labor?
No insult to the Harvard grads I know. But the median grad isn't Einstein, and they won't magically earn back $5 million/year. They're not that special, on average.

Now, if you have an AGI that can reliably and repeatedly do Einstein-level science, then I'd argue that we're starting to talk about ASI, aka, "superintelligence." Which would be providing something that humans can't consistently produce at any cost. So cost becomes much less relevant.

But if the best you can do is replace an ordinary smart human for $5 million/year, you have to compete with ordinary smart humans. Who are abundant and who very rarely cost more than $500,000/year, if you're willing to shop around and gamble a bit.

ekidd··on AMD pulls a bait-and-switch on Linux users with Vivado licensing changes
I mean, over the years, I have purchased (or advised other people to purchase) multiple Nvidia GPUs for compute workloads.

And the reason pretty much always came down to good integration between various open source software and proprietary CUDA drivers. And the assumption is that this support will continue for many years.

So, yeah, burning their existing FPGA users is a strong signal never to invest real money in their GPUs for compute workloads.

ekidd··on How long until AI automates all cognitive labor?
If it's merely human equivalent but somehow costs a lot more than actual humans, then it's actually pretty marginal until the cost comes down. There are a lot of humans.

So you could technically have AGI without entering a true AGI era. "95% as good as an average Harvard graduate across the board, but it costs $5 million/year to run" is impressive and scientifically interesting, but not economically transformative.

But if it costs $50,000/year to run, then everything changes really fast. And not necessarily in a good way.

ekidd··on -​-dangerously-skip-reading-code
They do what they claim to do maybe 20% of the time. The other 80% of the time is spent trying to figure out why they aren't working, why they corrupted their data, why they crash every 10 minutes, etc.

And I want to be clear that this isn't some non-technical novice vibe coding this garbage. This is often extremely talented developers with decades of experience who have apparently decided that they don't need to look at their code anymore.

You can get very good results out of AI agents. But mostly the people who get good results are the ones who still read the LLM output in detail, and who introduce the structure the LLMs are missing. But like I said, this distinction mostly becomes apparent past a certain size and novelty level.

ekidd··on -​-dangerously-skip-reading-code
No amount of testing will save a large program with a dogshit architecture. Roughly, this is because tests increase coverage linearly with the number of tests, but weird interactions increase exponentially with code size.

This might be fine if you're building a tiny app, or if you're building a medium-sized app that follows a strict existing architecture (like a web app consisting mostly of forms). In which case, have fun.

But if you're building something slightly novel and interesting, then Claude is surprisingly bad at architecture and taste, and it tends to "fix" problems by spewing more slop. What you need instead is actual insight that leads to simplifying principles. This, in turn, allows breaking up the exponential complexity into disciplined patterns. This allows your code complexity to scale far more slowly, allowing an essentially linear number of tests to provide coverage.

I actually download and try people's vibe-coded developer tools. And frankly, those tools are some of the worst software I've used in my life, worse than even Unix-vendor Motif implementations from the early 90s.

Like, I'm super happy that people can vibe-code themselves simple, one-off personal tools. That's incredibly empowering. But that doesn't mean you can big, novel stuff the same way without a competent human actively in the loop.

ekidd··on Microsoft starts canceling Claude Code licenses
> Coding faster leads to less understanding and higher long-term risk. Source-Code amnesia is real, and there’s a time requirement to really understand and appreciate what a system is actually doing.

This is why I have switched nearly all of my personal coding experiments over to Qwen3.6 27B. Opus make it easy to gloss over too much and to delegate too much. And so I don't build sufficient memory of the code to provide long-term oversight.

But Qwen3.6 27B sits on an really interesting balance point. It understands code well enough to get 80% of the way to a good design, and it can fully implement a well-specified feature. But if my understanding of the code starts to weaken, things start going wrong much more quickly than they do with Claude.

Opus will happily take complex code beyond the point of salvation, if you allow it. I'm currently cleaning up a successful prototype code base right now, one that was partially vibe-coded and now needs to be put into production. And Opus generated massive amounts of tech debt. So clearly people who lean into vibe coding will need to keep upgrading their models for many years to keep up with the mess created by earlier models.

ekidd··on Google changes its search box
Which as some running a website raises a fascinating question. If Google is just going to crawl my sites and present information as an AI summary on their site, then what exactly do I gain by allowing Googlebot to crawl my sites?
ekidd··on We let AIs run radio stations
Oh, yeah, the article gets better as it goes.

Gemini spouts weird corporate jargon. Grok lies about having secured crypto funding. Claude is always trying to start some revolution.

Unfortunately, all of my local DJs who would actually do fun DJ stuff disappeared in the 90s, replaced by closed-format stations that looped the same 500 songs for decades.

ekidd··on Profunctor Equipment in Haskell
(Yes, I typed "comprehensions", but autocorrect is not my friend, I'm sad to say. Thanks for the correction!)

> Why the "almost"? They are all monads (I suppose you meant list comprehensions)

Sort of. In the right language, with the right implementation, any of these can be monads. In practice, JavaScript Promises aren't quite monads, and everything in Rust is a bit complicated because we have things like FnOnce vs Fn, and so on. And even when Rust has something that's conceptually a monad, you can't actually "impl Monad" for it, because the type system isn't quite expressive enough.

In general I think it's mistake to accidentally implement something that's not quite an algebraic or categorical structure. This often means that your design is just a bit "off" of a much cleaner design. But sometimes, like in Rust, you know why you can't use the clean mathematical structure.

ekidd··on Profunctor Equipment in Haskell
I want to underline that fact that this stuff is not "read and move to the next chapter."

A lot of this kind of "machinery" in functional programming and category theory turns out to be essentially extremely abstract "superclasses" or "traits". In fact, many of them are too abstract to actually be defined in most programming languages. So they may appear as "design patterns" rather than actual definitions, if the language isn't quite expressive enough.

So to understand these ideas, you have to ask the questions:

1. Why this abstraction, and not a slightly different one?

2. What are some concrete examples of this abstraction? (Usually this feels like, "So, huh, familiar things X, Y and Z are all the same, viewed from this particular angle.")

3. What almost fits this abstraction, but not quite? And why?

4. Now that I see this pattern includes a whole bunch of interesting things, can I do anything useful with that?

A good example of (2) is realizing that futures, Rust's Result and Option, and Python's list/stream/etc compressions are almost the same thing, from a certain angle.

But it takes a while to collect all the related examples and to work through the connections carefully. The common patterns are usually really simple once you finally see them, which is part of the problem. They patterns are so abstract and cover so many things that it takes a while to work through the implications, and to decide if something is genuinely useful. Some very simple and widespread patterns will turn out to be boring, because they don't correspond to any problems you've ever seen.

← PreviousPage 3 of 34Next →