HNHacker News
TopNewBestAskShowJobs

lukev

9,321 karma · joined February 3, 2010

submissionscomments
lukev··on Show HN: Jev Plays Pokémon Red
Jev is interesting in that it's much cheaper and faster than a frontier LLM.

But I've seen nothing to indicate that the upper bound on classification tasks of a Jev-like model can exceed a frontier LLM with reasoning tokens. That seems nearly impossible even in principle (since Jev-style models are still based on LLM pretraining).

So while they're definitely on the Pareto frontier, which is valuable, they're at the "cheap" end of the spectrum more than the "good" end and I don't expect that to change.

lukev··on Three sites made 215,128 “best software” pages for AI. Perplexity cites them
Begun, the AI SEO wars have.
lukev··on METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
From the report:

> Because there were over a thousand transcripts and most were extremely long, we had to heavily delegate our analysis to AI agents; these agents had significantly worse judgment and reliability than human researchers, and it was challenging to spot check their work because both the underlying data and the agents’ analysis of it was often difficult to interpret.

> We estimate we spent roughly ~$400K in API credits over the six days of our investigation.

I don't understand why you think it's conceptually absurd? I use agents to analyze complex production issues all the time and they are very much capable of hallucinating a narrative.

lukev··on METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
The elephant in the room here is that the METR report itself was researched and compiled almost entirely by AI, with only very limited human "spot checks."

So I'm really not sure how much of it can be believed, especially since AI agents are strongly biased about the capabilities of AI agents.

lukev··on Are AI Labs Pelicanmaxxing?
Well, it’s mentioned as a limitation of the analysis, very much not ruled out (or in.)

That simonw is causing labs to do extra fine-tuning runs for this seems highly probable :)

lukev··on Are AI labs pelicanmaxxing?
What if they’re not pelicanmaxxing, but svgmaxxxing in general?

Because otherwise using a LLM to generate complex svgs is pretty niche and what I thought made this a good benchmark when it was new - generalized programming and spatial knowledge.

Obviously image gen in svg format is not a particularly hard problem if tackled directly on its own.

lukev··on Telus Uses AI to Alter Call-Agent Accents
You get calls about a new service or promotion, and it's the diction of the caller that makes you not wish to engage...?!
lukev··on Modern Front end Complexity: essential or accidental?
Counterpoint: the standardized surface area of a browser is already enormous, and while these components seem simple, there are a billion different options, variables or alternative implementations to consider.

At some point, functionality needs to exist in user space, even if it's common.

lukev··on The Abstraction Fallacy: Why AI Can Simulate but Not Instantiate Consciousness
What was the “well defined” definition? I’m not aware of any other than “this particular thing a human can do that I expect would be difficult for a computer.”
lukev··on The Abstraction Fallacy: Why AI Can Simulate but Not Instantiate Consciousness
"intelligence" is not well defined. LLMs are throwing this into high relief with how "spiky" their capability curve is. Yes, they can solve some crazy hard problems with enough compute and thinking tokens. Yes, they also fall down in the dumbest ways without an ability to self-correct... despite how "smart" they are, human supervision remains absolutely critical for any system of importance.

But I don't think the takeaway is "humans are intelligent and LLMs are not", it's that our vocabulary for talking about the intersection of language, cognition and compute is not up for the task.

lukev··on Claude Design
I would bet that Canva's bet is that companies will always want a "last mile" of manual control, even if only for the Queen's Duck effect. If Canva is the default, zero friction path for that, great for them.

The alternative is to not hop on the AI bandwagon, or run an "also ran" AI story, and both those scenarios (I expect) game out worse given the current zeitgeist.

lukev··on Ancient DNA reveals pervasive directional selection across West Eurasia [pdf]
Well if you are talking about environmental stuff (like leaded gasoline), sure.

If you’re talking about trying to improve the genetics of populations at scale… yikes.

lukev··on The future of everything is lies, I guess: Where do we go from here?
This is a must-read series of articles, and I think Kyle is very much correct.

The comparison to the adoption of automobiles is apt, and something I've thought about before as well. Just because a technology can be useful doesn't mean it will have positive effects on society.

That said, I'm more open to using LLMs in constrained scenarios, in cases where they're an appropriate tool for the job and the downsides can be reasonably mitigated. The equivalent position in 1920 would not be telling individuals "don't ever drive a car," but rather extrapolating critically about the negative social and environmental effects (many of which were predictable) and preventing the worst outcomes via policy.

But this requires understanding the actual limits and possibilities of the technology. In my opinion, it's important for technologists who actually see the downsides to stay aware and involved, and even be experts and leaders in the field. I want to be in a position to say "no" to the worst excesses of AI, from a position of credible authority.

lukev··on Ancient DNA reveals pervasive directional selection across West Eurasia [pdf]
To be clear: most people who are keen on making such an argument, or who are identifying racial genetic differences as the primary takeaway of studies like this, are doing so to justify racism, either implicitly or explicitly.

But that's a strawman. Racism is wrong, even if there are minor genetic variances across populations (which... seems obvious?) Variance within a population strongly dominates the weak cross-population effects, and personal history (nutrition, education, etc) strongly dominates that.

And that's setting aside the moral implications of judging someone or changing your behavior towards them even if you have somehow measured them to be "less intelligent," as if that was a single axis of worth.

Because, apparently, this needs to be said.

lukev··on Exploiting the most prominent AI agent benchmarks
Did you read the article? There's a whole section on "this is already happening."
lukev··on Exploiting the most prominent AI agent benchmarks
I think we should all consider the possibility that part of the reason Anthropic hasn't immediately released Mythos is that it would be slightly disappointing relative to the benchmark scores.
lukev··on Small models also found the vulnerabilities that Mythos found
This is a really interesting point though -- it's really scaffold-dependent.

Because for the same price, you could point the small model at each function, one by one, N times each, across N prompts instructing it to look for a specific class of issue.

It's not that there's no difference between models, but it's hard to judge exactly how much difference there is when so much depends on the scaffold used. For a properly scientific test, you'd need to use exactly the same one.

Which isn't possible when Anthropic won't release the model.

lukev··on The cult of vibe coding is dogfooding run amok
Disagree, I don't particularly want to up the level at which I'm building the core. Core is where I want to prioritize quality over speed, and (at least with today's models) what I build by hand is much, much higher quality.
lukev··on The cult of vibe coding is dogfooding run amok
I like this framing, but it does seem to imply that a whole dev shop, or a whole product, can or should be built at the same level.

The fact is, I think the art of building well with AI (and I'm not saying it's easy) is to have a heterogenously vibe-coded app.

For example, in the app I'm working on now, certain algorithmically novel parts are level 0 (I started at level 1, but this was a tremendously difficult problem and the AI actually introduced more confusion than it provided ideas.)

And other parts of the app (mostly the UI in this case) are level 7. And most of the middleware (state management, data model) is somewhere in between.

Identifying the appropriate level for a given part of the codebase is IMO the whole game.

lukev··on The threat is comfortable drift toward not understanding what you're doing
Yes. That is the point I was making.

Calculators provide a deterministic solution to a well-defined task. LLMs don't.

lukev··on The threat is comfortable drift toward not understanding what you're doing
That's not what I mean.

If I use a calculator to find a logarithm, and I know what a logarithm is, then the answer the calculator gives me is perfectly useful and 100% substitutable for what I would have found if I'd calculated the logarithm myself.

If I use Claude to "build a login page", it will definitely build me a login page. But there's a very real chance that what it generated contains a security issue. If I'm an experienced engineer I can take a quick look and validate whether it does or whether it doesn't, but if I'm not, I've introduced real risk to my application.

lukev··on The threat is comfortable drift toward not understanding what you're doing
I think that's too easy an analogy, though.

Calculators are deterministically correct given the right input. It does not require expert judgement on whether an answer they gave is reasonable or not.

As someone who uses LLMs all day for coding, and who regularly bumps against the boundaries of what they're capable of, that's very much not the case. The only reason I can use them effectively is because I know what good software looks like and when to drop down to more explicit instructions.

lukev··on Signing data structures the wrong way
So, isn't this a rather longwinded way to say that a signature only extends to the scope of the message it contains?

It doesn't matter if I sign the word "yes", if you don't know what question is being asked. The signature needs to included the necessary context for the signature to be meaningful.

Lots of ways of doing that, and you definitely need to be thoughtful about redundant data and storage overhead, but the concept isn't tricky.

lukev··on Last gasps of the rent seeking class?
I could not possibly enumerate all the possible things that have been enclosed. Human beings obviously being the most morally egregious.
lukev··on Last gasps of the rent seeking class?
Also, the defining feature of capitalism is that it encloses what was previously common.

Land used not to be owned (feudal lordship was functionally different than private ownership.) Then, society shifted, land became private, and that was the beginning of rent. This is enclosure.

The whole concept of IP is to explicitly extend this process to ideas -- they are not free, they are owned, and I have to pay you to use them. This is also enclosure, precisely.

lukev··on Last gasps of the rent seeking class?
Rent is charging money for access to an asset or property.

The P in IP is Property.

lukev··on ARC-AGI-3
I'm not sure how this relates to AGI.

This measures the ability of a LLM to succeed in a certain class of games. Sure, that could be a valuable metric on how powerful (or even generally powerful) a LLM is.

Humans may or may not be good at the same class of games.

We know there exists a class of games (including most human games like checkers/chess/go) that computers (not LLMs!) already vastly outpace humans.

So the argument for whether a LLM is "AGI" or not should not be whether a LLM does well on any given class of games, but whether that class of games is representative of "AGI" (however you define that.)

Seems unlikely that this set of games is a definition meaningful for any practical, philosophical or business application?

lukev··on Is anybody else bored of talking about AI?
This is bad in tech. But at least we are (relatively) well equipped to deal with it.

My partner teaches at a small college. These people are absolutely lost, with administration totally sold on the idea that "AI is the future" while lacking any kind of coherent theory about how to apply it to pedagogy.

Administrators are typically uncritically buying into the hype, professors are a mix of compliant and (understandably) completely belligerent to the idea.

Students are being told conflicting information -- in one class that "ChatGPT is cheating" and in the very next class that using AI is mandatory for a good grade.

Its an absolute disaster.

lukev··on The bridge to wealth is being pulled up with AI
That's right -- the best way to succeed within a system is to hustle as hard as you can, and definitely don't stop to question the system itself.
lukev··on Ask HN: AI productivity gains – do you fire devs or build better products?
I'm not speaking of burdens of proof about unfalsifiable statements.

I'm saying that I think this is an important enough question that I think we should seek real evidence in either direction, especially since apparently everyone already has a strong opinion (warranted or not.)

Page 1 of 34Next →