HNHacker News
TopNewBestAskShowJobs

ekidd

11,517 karma · joined June 5, 2010

Blog: http://www.randomhacks.net/ Mastodon: https://mastodon.xyz/@emk

I write a lot of code in Rust, Python and TypeScript. Some of it is open source. Somehow I get paid for this.

I've done startups and consulting, but I'm pretty busy right now.

submissionscomments
ekidd··on The Download: why AI's latest breakthroughs and fears may be more hype than rea
Looking at "Felony Bench" https://www.felonybench.com/ , I see that a majority of known "rogue model" incidents do involve cybersecurity evaluations. But several of them do not. The attacks on RubyGems appears, bizarrely, to have had the goal of downloading freely available data from the UK government during some kind of research task. There is also probably some sample bias: Most of these models have monitors that attempt to detect offensive cybersecurity uses, and those monitors are only turned off during cybersecurity evals. Therefore, models doing ordinary research tasks that go off the rails are likely to be caught early, before they get around to committing felonies, and they will thus be underrepresented in the data.

Also, if you a tell a model, "Please break into evaluation server X," and if the model decides to cheat on the test by breaking into companies Y and Z to steal an answer key, that is still very bad. We all see how that's bad, right?

After all, the broomstick in the Sorcerer's Apprentice was doing exactly what it was told, too. "The model was sort of obeying the humans when it started committing felonies" is not a very reassuring excuse.

But the most relevant idea here is sometimes called "instrumental convergence." No what goals you have, there are certain subgoals that almost always help: Accumulate money and power. Avoid getting turned off. Don't get caught. Etc. So, for example, you could pass the cybersecurity evaluation by performing the requested tasks. But maybe the grader made some mistakes and mislabeled some answers. In that case, the "right" answers will occasionally lose you points. If you want a perfect score, the only way to do it is to steal the teacher's answer key.

But also, let's not forget the "OMG demons" part of this. We now have models that can pull off complex attacks with thousands of steps, abilities that used to be reserved for intelligence agencies and highly motivated CTF teams. This frog may not be boiled yet, but the water's getting uncomfortably warm.

ekidd··on A misalignment of AI in mathematics
> Those who think it's me or the machine will fail.

> Those who realize how much you can accelerate your research with the help of AI will succeed.

This is only true up until a point. If I treat a mid-sized model (say, Qwen3.8 Flash Next) like a pair programmer, then yes, it accelerates my work.

But I can already see the next stage with Fable: If I give it a couple of paragraphs of spec and $50, then I can just leave the room and go wash the dishes. I learn nothing, I participate in nothing, and I bring nothing to the process. I am no longer succeeding at all. Fable's succeeding without me.

Now, in this model generation, Fable starts getting sloppy after a few thousand lines. I can still build better at scale.

But I don't expect AI to accelerate humans or improve our productivity for long. I can already see the first signs of a future where the AI doesn't need us for anything at all.

ekidd··on DeepSeek v4.1 Flash
On the other hand, we have recently taught rocks to think about software engineering. And while not perfect, they're surprisingly good at it. Once you start building things that are even a little bit like minds, I suspect that it's worthwhile to consider that the future might end up looking a bit like science fiction.

The alternative is to insist that Nothing Ever Happens, and the future won't get too weird. Which is no longer a bet I'm entirely comfortable with. Weirdness is at least a possibility.

ekidd··on Research acceleration: The view inside OpenAI
Eigenvectors represent fixed directions, not fixed magnitudes. From Wikipedia:

> More precisely, an eigenvector v of a linear transformation T is scaled by a constant factor lambda when the linear transformation is applied to it: Tv = lambda v .

In other words, repeated multiplication of an eigenvector by a matrix can still create exponential growth.

ekidd··on How I feel about AI
> > we should fight for human survival

> All problems arise from this statement.

Let me go on the record, then: Given a choice between every living human being becoming absolutely, for-real dead and the human species surviving, then I come down on the side of human survival. There are too many humans that I love personally, and too many innocent people just going about their lives that I would like to protect. Given a choice between the mass death of billions, and their survival, I choose survival.

Human rights are essential. But human rights are difficult to achieve if everyone is dead.

ekidd··on How I feel about AI
> * For a section of population who were in a certain age-range when the pandemic hit and that derailed their certain kinds of plans/hopes/dreams in personal life*

...

> So I feel a strange kind of relief about this possibility of doom.

So let me make sure I understand what you're saying. The pandemic disrupted your life plans for a few years and threw you off the track you planned? (Which does suck.) And now the possibility of actual human extinction makes you feel a strange relief? Am I understanding this correctly?

If so, this is a remarkable level of something. Depression? Nihilism? Jealousy? Something else?

Imagining total human extinction and feeling relief is not a healthy state of mind. And, yeah, you seem to know this. But I kind of wish we could reach some sort of broad public agreement that human extinction would be bad. And if it ever looks like it might happen, we should fight for human survival.

ekidd··on Claude Fable 5.1 and Claude Mythos 5.1
Yeah, but I understand that fingerprinting is essentially a pseudorandom overlay onto a pseudorandom base signal. And unless you have access to both the random number generators and the weights, I don't think you can detect it?

So "fingerprinting" operates on a totally different and basically invisible level, as opposed to the obvious stylistic patterns that the average programmer can identify in about 2 sentences.

ekidd··on Agents still can't automate Excel
Yup. Their benchmark appears to consist of giving an agent XLSX files that use advanced Excel features, but no copy of Excel, and expecting the agents to interpret the spreadsheet the same way Excel would. So the agents try LibreOffice instead, and if that doesn't work, they try to write Python scripts. But the behavior of the Python doesn't match the behavior of Excel.

To put it politely, this seems like a self-inflicted problem. If you really want your agents to interpret advanced Excel features exactly the way Excel would, have you considered maybe giving your agents Excel?

ekidd··on Creepy Crawlies
Largely because they're residential botnets in places like Brazil (a real example from one of my sites that was crawled to near-destruction). Someone could probably do something about this, but it's out of reach for individual site owners.
ekidd··on Iceland votes on whether to restart talks on joining EU
Thank you for the clarifications!

Someone from Iceland tried to explain this all to me once, and if I understood correctly, it had less to do with the GDP (which also includes money circulating within Iceland) and more to do with the balance of trade, where fishing played a bigger role? But my information may have been misunderstood, or it may be out of date.

In any case, you have a stunningly gorgeous country and very nice people, and I wish you the best of luck in figuring all this out in a way that benefits Iceland.

ekidd··on Iceland votes on whether to restart talks on joining EU
> The No campaign fears losing control over Iceland's prized fishing grounds under the EU's Common Fisheries Policy and has vowed never to share the country's waters with anyone. Brussels has indicated Iceland could earn some kind of exemption, but it is considered potentially the biggest obstacle to any agreement.

As noted in the article, Iceland's fishing industry makes up something like 40% of their exports. Which is the only way a tiny island nation can afford essential imports. On top of that, Icelandic fishing is apparently carefully regulated to prevent overfishing, with the result that it is supposedly one of the few truly profitable fisheries in the world.

So maintaining control over their own fisheries may be a survival-level issue for Iceland.

On the other hand, they have a tiny population, no ability to defend themselves, and membership in the fraying NATO military alliance. So I can see why they're debating this issue. Aligning themselves more tightly with Europe doesn't fix their problems, but I can also see not wanting to go it alone.

ekidd··on U.S. State Department pauses immigrant visa applications
Yeah, this is the problem. If you're on an immigrant visa, working a real job, then you have a life in the US: an apartment or a house, a pile of belongings, maybe a spouse and kids. If you need to leave the country to renew your immigrant visa, you may not be allowed back in to pack your apartment, sell your house, or help your US spouse and kids prepare to move to a country that will allow your family to stay together. It's a giant neon sign telling you you're not at all welcome here.

Of course, if you have a spouse and kids, this all gets easier once you have permanent resident status. Which used to be comparatively easy to get once you were married and you had offspring running around. I don't know what that process looks like now.

In practice, this kind of arbitrary screwing with people makes the US far too risky for educated, professional immigrants. So the scientists and the engineers and the company founders will just go elsewhere and build that country's economy. Which is fine, in the long run: If the US doesn't want to build the future or create new industries, then the voters have every right to vote for that. We've already decided we have no hope of competing with China in cars. And we'll need to cede more industries over time. As far as I can tell, the US has benefited heavily from poaching other people's scientists and engineers and professionals for decades now. But maybe the voters don't buy my argument, and they believe that chasing away skilled professionals and imposing 100% tariffs will lead to the future they want. It's a difference in beliefs about about what makes a country prosperous, and people have a right to vote based on their beliefs.

ekidd··on Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
> That really depends on the model, I run a few models locally. All at speeds comparable to or faster than Opus.

Yes, a lot of Qwen3.8 27B setups are actually quite snappy, as long as they fit 100% in VRAM. In my testing, I wouldn't go below 32GB of VRAM, though—you really want a 6-bit quant and 8-bit K/V quants minimum. I've seen too much weirdness out of 4-bit quants since Qwen3.8 shipped. I think it may be damaged more than 3.6 at similar levels of quantization?

If hyperscalers hadn't bought up almost all the fast RAM production for the next several years, 32GB of VRAM would be tolerably cheap—a lot by "home PC" standards, but not terrible by "professional tools" standards. Sadly, the RAM market is amazingly ugly right now.

> I doubt the timing differential here, but even still I run my 3090 pretty heavily with inference workloads and it stays cooler than when I use it for gaming.

Yeah, running inference on a laptop is likely to run quite hot. But in an ATX case with decent cooling, it's generally a lower load than gaming. One handy tip: Many Nvidia GPUs (and some from other manufacturers) support power limits. For example, limit a 5090 to 400W instead of 600W, and it will run much cooler. You might lose 11% off your tokens/sec (depending on the exact card).

ekidd··on I were 17, I'd learn how to build LLMs from scratch
> AI is the subtrate the future runs on.

Current AI can automate significant amounts of grunt work in programming and math. It's good at running web searches and writing summaries. There are a few other niches where it is currently successful. But other than that, many corporate AI projects are spectacular failures.

So just given what we have in hand, assuming no further breakthroughs, then we're maybe looking at AI being somewhat bigger than the Internet. Which would make it a revolutionary technology, sure.

But to get from "a revolutionary technology" to "the substrate the future runs on", then you need to assume more breakthroughs: long-context operation over weeks or months, displacing human workers 100% instead of 75%, and the ability to directly economically compete with actual humans. And people are investing literal trillions of dollars to make that future come true, without really thinking through what truly competitive-with-human AI would actually mean. We might be looking at massive job loss, centralization of power, fully automated "companies" with no humans dominating markets, and other dystopian scenarios.

And in those worlds, it's unclear that being good at CUDA and matrix math will be all that helpful, careerwise. The AIs are already pretty good at that stuff. Data scientists get paid OK when they actually get hired, but it's not everything college students were promised in the 2010s, either.

We can't yet build a fully-general competitor for the human mind. But we're getting closer. And if we ever do build one, the consequences will be really weird in any number of ways. So I worry about visions of the future that assume AI keeps improving significantly, but that also assume it still somehow remains a "normal" technology that doesn't, for example, render most humans fundamentally uncompetitive.

ekidd··on My friends all hate AI; I just joined an AI startup
I've done 35,000+ Anki card reviews over the years, and I strongly suspect that most of the obvious and encouraged ways of using Anki are counterproductive. (Full disclosure: The following thoughts are heavily influenced by using Anki for language learning, and may not fully generalize to subjects like medicine.)

- Use Anki as an "exposure amplifier", not a database of raw facts to memorize. Think "tampering with the relative frequency of input to your brain", not "I will memorize every fact."

- Automate card creation as a much as possible. This will almost always require custom tooling. Card creation is part of the learning process, but you still want to automate everything you can possibly automate. Imagine tools to turn ebook highlights into cards, or to turn entire movies worth of subtitles into listening cards.

- Delete cards ruthlessly. If a card makes you say "Ugh", just delete it. If you're automating as much of card creation as possible, and if you're treating Anki as an "exposure amplifier", then deleting cards is cheap. And a significant amount of review pain comes from a small minority of the cards. Cull them, and reviews start to feel like popping bubblewrap, and not like getting a root canal.

- Prioritize large context and small active recall. For language learning, you might capture several sentences of context from a book or web page you're reading. Read or ignore the context depending on your mood. The active recall portion of a card significantly boosts how much you retain. But oddly, it seems OK to make the active recall portion really small. Try hiding half a word, or a single preposition in a phrase. For passive cards, just boldface an interesting expression in context, and mark the card as passed if you mostly understand it in that context.

- Introduce new cards slowly. Total review burden will stabilize around 5-6 times the number of new cards you learn a day. So 10 cards is a good default, and 20 is a significant commitment.

- The critical window seems to happen around 20-30 days after you introduce a card. This is when either your memory magically "consolidates" somehow (and formerly challenging cards suddenly become blindly obvious), or when a card enters repeated failure territory (and should thus be deleted ruthlessly). You could probably do just fine if you auto-suspended any card that you successfully recalled after a 20- or 30-day gap.

Also, this kind of tooling is the sort of thing you could just vibe code these days. If you do, make the "Delete" button really big.

ekidd··on GenRec: Towards LLM-Native Recommendation at Netflix
I think the biggest change from the Netflix prize days is that their catalog went away. Back in the DVD era, they had everything. And in the very early streaming era, they still had a huge number of things to watch.

But their current catalog is badly impoverished, and they're just going to recommend the same 30 Netflix originals they always recommend to me, plus a few films or series that are rotating through on a temporary license. If I actually try to search for something specific I'd really like to rewatch, it's almost never there. They haven't quite regressed to the level of a small-town, early 90s Blockbuster, but it sure feels that way sometimes.

So honestly, how much good can the Netflix recommendation algorithm do these days, given the much smaller catalog of movies and films it apparently has to work with?

(The one that I don't get is the Kindle recommendation algorithm. If I read one really good book with a certain theme, the Kindle immediately replaces my recommendations with 40 bad knockoffs, 25% of them clearly AI written. There's apparently no signal for "actually good.")

ekidd··on DeepSeek V4 Flash 0731
> If what you're saying is true and accurate, then US-based AI labs are in big trouble.

I've been working with DeepSeek V4 Flash 0731. I'd say that it's maybe not quite as smart as Opus 4.5, but it's willing to think things through carefully and keep going until it gets a good answer. So it's a decent Opus 4.5 replacement. Just let it cook.

It isn't Opus 5 or Fable 5. But it's nearly free on Open Router, and it's self hostable on a Mac Studio with plenty of RAM, or using an RTX Pro 6000 Blackwell or two. Which is chump change for any company that employs programmers.

It would absolutely have been a frontier model last December.

ekidd··on Improving GPT‑5.6 Sol in ChatGPT, expanding GPT‑5.6 Luna access for free users
Google's public AI model lineup is pretty bad right now. Gemini 3.1 Pro is scoring worse than some mid-sized Chinese models at 1/20th the cost, and their Flash and Flash Light models are horrendously overpriced on task benchmarks compared to GPT 5.6 Luna or the mid-sized Chinese models.

There's no reason why Google's public stuff is this stale, overpriced and underwhelming. But at least until their next round of models drops, even calling them a "frontier lab" is starting to feel like a stretch. Which is weird!

ekidd··on DeepSeek V4 Flash on a Single AMD MI300X
> Well, if no body can say when it will pop, then can we really say it's a bubble and it's overvalued?

Well, given the literal trillions being spent, the only ways this pays off are:

1. AI replaces a non-trivial fraction of human employees.

2. Someone builds a Culture Mind, and humans become (hopefully) pampered pets of AIs we don't understand. Seems unlikely, but it would arguably count as a payoff even if it made money meaningless.

Or maybe the AIs don't want pets, and you get SkyNet. Which definitely doesn't care about paying off anyone's investments.

When you look at various news articles about investors, yeah, there are definitely a lot of rich people who think that they're going to automate all human labor or just bring about the Singularity. Possibly with them in charge of the rest of us. If you don't make these kinds of wild assumptions, then yeah, this is looking like one of the biggest bubbles ever.

ekidd··on AI doesn't generate working products, that's still your job
> Note that they would also need to keep the economy alive because somebody has to pay them money. They can't let everybody to go unemployed.

Well, that's why they want the robots! If your robots are capable enough, and if your AIs are smart enough, why, they could just build the yachts directly!

Right now, ordinary humans are needed by the economy because we do all the work, and because robotics hardware is still far behind AI. But there's no inherent logical reason why you need a human to turn raw materials into luxury products. The really important questions are: Who controls the AI and robots? How good will they get? Who or what does all the work? And who controls the natural resources?

> Of course (let me go distopic)

Yup, that is a possible end state.

ekidd··on AI doesn't generate working products, that's still your job
What I actually fear is more subtle: I already do a fair bit of project management and technical leadership. I could do more. Sure, I'd miss the coding, but I also enjoy a lot of the stuff around it.

But the goal is to expand what the AI can do in each generation. At this point, Fable 5 can ace almost any greenfield project a skilled developer might have written in a few days. But it's bad at refactoring, bad at keeping the code clean as it goes, and bad at discovering new insights as it codes. So Anthropic will train Fable 6, using benchmarks like SlopCodeBench that test maintenance over time.

Now what about project management? Train Fable 7. What about product management and talking to stakeholders? Train Fable 8. What about market research and sales? Train Fable 9.

By this point, Anthropic doesn't need to actually release these newest models to the public. Why, that might be dangerous! Instead, they write, "deisgn [sic] a successful software product and sell it plz." And they spin up a million dollars worth of compute and let it crank out SaaSes, iPhone apps, etc., driving entire software companies out of business.

Then they spin up some more instances, and say, "make robot plz" and "try a thousand ways to make yrself smrater." I mean, Qwen and DeepSeek keep finding ways to pack more smarts into a given number of weights. Fable 9 will likely be able to do the same. Hell, Fable 5 can probably run 1,000 machine learning experiments now, just grinding through ideas the way ChatGPT's internal models grind through proofs.

And this is my problem. If it were just programmers losing their jobs, well, sometimes professions die. But what makes you think it will stop with us? How far will this go in the next 4 years? The next 20?

ekidd··on OpenJDK Interim Policy on Generative AI
Is anyone using even a tiny on-device language model for spelling? For grammar, I could almost imagine it.

But this also bans "simple" AI-powered auto-complete, like Zed's Zeta2 model. This is a very conservative model that rarely tries to propose more than a few obvious lines (at least in my use cases). If a developer accepts a three-line autocomplete that introduces a bug, that's kind of on them. Honestly, about the only thing that Zeta2 is good for is reducing the risk of RSI.

To be fair, it's also their project, and I completely support their right to set whatever policies they want.

ekidd··on Handbook.md shows that long policy documents do not reliably govern agents
For values of "usable" that include "14.6 seconds/token". It's a cool accomplishment! And newer hardware would speed it up some. But I think I'd want something a bit faster before declaring it usable in practice.
ekidd··on After the AI Crash
Well, since you "freelance using GPT" and you expressed concern that you'd stop being able to make a profit if you had to pay API prices, I figured that swapping an off-the-shelf graphics card into a gaming rig might be a reasonable way to stay in the black. If it came to that. And the software setup on Linux is just compiling and installing llama-server, which shouldn't be enough to stop any programmer trying to make a living.

Or you could just take your credit card and spend $20 on credits at https://openrouter.ai/deepseek/deepseek-v4-flash. I'm not sure that I could manage to spend even a $1/day at those rates.

My larger point is that while frontier tokens are a near-monopoly and who knows what they really cost, many real-world workflows can be run using commodity tokens, or even served in-house by anyone who can afford to hire US or EU programmers. And if you're willing to settle for what would have been a state-of-the-art coding model in October 2025, commodity tokens are close to free.

ekidd··on After the AI Crash
You're unlikely to ever be priced out of tokens, at least if you'd be willing to settle for a model closer to Sonnet 4.5. That level of model certainly isn't as efficient as Fable 5, but it can crank out CRUD apps and other consulting mainstays quite well, with some supervision.

To give you an example of model in this class, the DeepSeek V4 Flash preview is a 284B A13B model, with a native quantization mixing 4 bit and 8-bit values. You can easily run it on an RTX Pro 6000 Blackwell (or 2) at a reasonable quant, especially if you offload the MoE weights to 48-64GB of system RAM. This costs US$11,800 at Microcenter right now, and it will work in any gaming box with decent cooling and a modern 1000W power supply. Over the lifetime of the card, an entire system would cost you under $4,000/year. Power is about 300W for the card (either a blower model, or a workstation model with the power cap), and another 150W or so for the rest of the server.

Or you could buy it on Open Router from dozens of different commodity vendors, starting around $0.09 per million tokens input, $0.18 per million tokens output. This is a competitive market price, so some of the providers might be losing money or reselling surplus capacity. But given the underlying hardware costs, the numbers are in the ballpark. In other words, if you're willing to settle for lower-quality tokens, you can get roughly Sonnet 4.5 for close to free, or as a modest capital expense for a successful freelancer. Halfway decent tokens are cheap, and you can generate them in-house!

ekidd··on Show HN: Formally verified 3D CSG: Trust 93 lines spec, not 1000 lines AI code
Still, this is not a bad situation overall. If you are forced to trust the C compiler to correctly compile C, and the hardware to correctly implement the instructions in the documentation, well, you're already forced to trust both of those every day.

So this doesn't give you an absolute proof of correctness. But it does substantially reduce the size of the problem, which is now limited to (1) verifying that you actually proved what you think you did, and (2) all the stuff you were normally trusting anyway. (Some of the stuff you were trust anyway is broken, of course.) But this is a smaller problem than trusting 1,000 lines of highly-optimized CSG code written by a model we don't actually understand.

ekidd··on What is happening to jobs? Separating AI hype from reality
> My experience so far, has been that it's barely a junior dev.

At least on greenfield projects, Fable is more like having an endless succession of fly-by senior devs who lock themselves in an office for a week, and who come back with very reasonable code that I then need to maintain somehow.

For longer-term maintenance, I am actually slowly warming to Sonnet 4.5-era models (so October 2025, right before the Opus revolution). They need to be given clear instructions and watched carefully. But since they force a human to stay in the loop, you don't have the institutional knowledge loss I see in some Opus projects, or the code that was one-shot with no human interaction at all that's a constant temptation in Fable projects.

And yeah, the bit rot is painfully real once the humans step back too far. I've been dealing with a compelling prototype that someone built, and that stakeholders love (for good reasons). But it had to be put on a tech debt repayment plan for a couple of months.

ekidd··on The new rules of context engineering for Claude 5 generation models
When I learned French as an adult, I pushed myself heavily into "forced immersion", avoiding English completely for long periods of time as my French slowly built itself up. This almost entirely silenced my inner monologue in English for those periods of time, and left me with a toddler's ability to think in French. In this comparative silence, it was much easier to observe my non-verbal thought processes moving around "behind" the scarce words.

Then my French inner monologue got good enough that I could mostly think in French, especially when I was in a French-speaking environment. One fascinating detail was that after switching from a French-speaking environment to an English-speaking one, I would actually spontaneously translate from French to English for about 15 minutes until my brain switched back.

So it seems obvious to me that it's possible to suppress or at least severely impoverish the language of thought, that other "layers" of thought exist besides the words, and that it's even possible to change the actual language of verbal thought.

Also, something which at least some other people in the HN crowd might recognize: When I'm deepest in the zone programming and refactoring, I tend to work with a lot of half articulated concepts I can't put into words. You know how people talk about "code smells"? That isn't a literal smell for me, but it's generally a non-verbal sense that a pattern is wrong.

ekidd··on Nvidia, Microsoft, Meta warn against overregulating open-weight models
DeepSeek has published some really good papers. Lately they're pushing really hard for dramatically cheaper serving costs.
ekidd··on The arguments against open source AI are bad
> Especially for point #1 I don't think we've established that - we've been given information by a private company that makes their tooling look extremely valuable which may be true and genuine or may just be yet another doomday statement to bolster their valuation.

Many of the details of the attack on Huggingface were reported by them before they knew who was attacking. So no, OpenAI is not the only source here. It was a pretty impressive example of an APT-style attack just from their end.

"Our model is powerful enough to commit multiple felonies (and we can't stop it)" is "marketing," I suppose.

Page 1 of 34Next →