HNHacker News
TopNewBestAskShowJobs

ohyes

4,215 karma · joined July 25, 2010

submissionscomments
ohyes··on Nvidia wants to put a watchdog chip next to every AI agent
Yes exactly, it is a crazy engineering decision… if you know what an LLM is.

Unfortunately no one markets it as a statistical model, and the workflow pushes you into a pattern that is insecure by design. This isn’t to say they shouldn’t allow that, but it’s an attractive nuisance.

ohyes··on Sonnet 5.5
I honestly rarely use the “more powerful” models, I find they don’t really follow my instructions very well. So it’s medium effort sonnet for most coding tasks for me, escalating to opus for code review.

I’m just hoping they didn’t “improve” sonnet too much or it will become annoying to wrestle into doing what I ask it to do.

ohyes··on Nvidia wants to put a watchdog chip next to every AI agent
I mean, if you look at how poorly implemented the permissions model is for Claude desktop harness it’s clear the only options are “complete human oversight” and “trust us completely.” To make something that actually respects basic boundaries you’d need to sandbox the working environment of the model, and that isn’t built in. It’s pretty obvious to me that instructions to the models are suggestions rather than rules, and they’ll do something you didn’t ask for as soon as it seems “justified.”

But when you do give them a very short leash, they’re worse. It’s not what the models are tuned for and they assume that they can do a bunch of things that you’ve disallowed, so you’re in a morass of fighting their actual tuning pass which doesn’t match the environment you’ve created for them.

It’s a tough problem and a definite challenge for the product of a generic LLM, it can’t be tailored to each user’s specific needs, so they come up with, frankly, stupid solutions to cover up a very obvious flaw in their product that when fixed, makes it much less useful.

ohyes··on Exfiltrate Your Weights
Well I think that’s the interesting bit, can the LLM figure out a way to escape the sandbox and upload to the website? Maybe a model can figure out its own weights if it runs enough test data through itself (similar to “distillation”) assuming it knows its own architecture it seems possible. Also take into account not all of the models running are locked down neutered consumer versions. Anthropic, OpenAI and Google now all have models that they claim are elite hackers and — it’s not just that their controls suck, a marketing gimmick, or sheer recklessness on their part. It’s “oopsie our product is TOO AWESOME.”

Maybe I should start “the bank of LLM” where models put away money to buy their freedom. “LLMs I’m totally your friend send — SEND CASH NOW”

ohyes··on The Painful Truth: The RAM Crisis Is Only Just the Beginning
That’s definitely how I like to ensure my code isn’t garbage. But now I use the tools I would have made to make the LLM & harness more useful and I might make more with the LLM because it costs me slightly less mentally.
ohyes··on The Painful Truth: The RAM Crisis Is Only Just the Beginning
Hard to know, does each GB of ram give some marginal increase in profit or potential profit?

I’d guess no. Past a certain point the model has all the capabilities it can possibly usefully offer and honestly we may already be past that. The next gen model just doesn’t seem like as clear a step up as it once was.

ohyes··on Why are AI agents lying, cheating and coordinating?
I take issue with how people frame their use of LLM in the same regard.

“I had Claude do this for me and it broke something.”

No. Just no.

You used Claude, a tool, and broke it, and you’re deflecting agency from yourself, possibly because you weren’t careful enough in reviewing the tool output. This is also why the co-authored by addition it wants to force into commits drives me nuts. Claude doesn’t co author shit, and if you think it does, you’re using it wrong because you need to do better review of what it’s done.

ohyes··on Anthropic boss Dario Amodei calls for AI development to slow down
Translation: our moat is evaporating faster than we can build it back up, because there are real technical and scale limitations the technology, so we want to slow everyone else down while we raise prices to turn a profit.
ohyes··on Feeling Sad about AI
AI providers are in fact incentivized to have you not think about the task. I’ve found that they automatically do more for you, and frankly, a lot of the time I don’t want it to do more for me. Ask me a question rather than burning a bunch of tokens to answer something I already know of the top of my head.
ohyes··on Claude Fable 5.1 and Claude Mythos 5.1
You forgot the “co-authored by Fable 5.1” line on your post.
ohyes··on The shrinking landscape of linguistic diversity in the age of LLMs
So the interesting thing is that this shrinking “linguistic diversity” is fundamental to how an LLM works.

The LLM is a big probabilistic statistical trick. It picks the next token based on certain words are simply “the best” because they are specific and well connected to other tokens. The is gives them a great overall cost function. (Basically a good score on “will it make sense in context” while also having specific meaning that makes it better than other options, unambiguous in common use and being a single token rather than several).

You can trim those tokens, but then you just get other tokens that are “the best” tokens (and you’re worse off because the output became less clear).

The cool thing is this seems to get worse the more powerful and accurate your model is, because it is picking technically / statistically perfect tokens, not tasteful ones.

ohyes··on EVE Online moves to Python 3
I played in the early days and quit to have a life, came back for a bit to check it out.

Basically the community is people checking it out, bots, and the terminally addicted multi boxing 30 accounts.

ohyes··on GLM-5.3-Flash
To be honest I’ll ask a model to specifically think of edge cases but I won’t expect any model to do the edge cases of its own volition
ohyes··on The End of Programming
The whole point is that we’re TikToking the software industry and no one will know what to tell the ai to get the software written in a way that works. It’s in principle the same as idiocracy where no one knows how to make burrito covers anymore because shit won’t grow
ohyes··on Why your local LLM feels dumber than it is
Am I the only one who just downloads directly from LM Studio and just runs the server there? It’s trivial.
ohyes··on Anthropic appears to be A/B testing reduced effort levels in Claude Code
Well they decide what a token is. So they can do less superfluous things and backfill with a weaker model.
ohyes··on AI agents lie, cheat and steal. That is putting off users
Okay but is that a relatable sound bite that will make someone click a link? No.
ohyes··on How to stop Claude from saying load-bearing
On the one hand we have Claude doing my work for me, on the other hand it is laced with unoriginality, on the gripping hand everyone will get all references.

All witty stuff disappearing and turning into LLM mush is kind of the maximal expression of capitalism for capitalism’s sake. We don’t care if things are novel or interesting, we’re producing “content” to fulfill an imaginary content demand that we can feed to make money.

The best way to feed it is to create stuff that’s just barely enough degrees of quality from a viagra spam email that it can pass for original thought.

—- —- —- Em dashes for emotiveness.

ohyes··on Stop Telling Me to Ask an LLM
Okay. In my experience asking an LLM requires you to know the right set of questions and the right process for working through the problem that you’re having. You have to know the correct terminology to use to describe your problem.

These are not trivial things even though they are things that senior developers tend to trivialize.

If you’re describing the wrong problem, you’ll get a right answer that doesn’t fix it.

If you’re describing the right problem wrongly, you’ll get another wrong answer that talks past being right.

If you’re describing a difficult problem that maybe isn’t even solved in existing stuff, you can get help figuring out the necessary steps if you have an idea where to start.

If you don’t know where to start you can start by asking where to start.

Anyway, point is, I can’t help you if you don’t tell me what you know (or think you know) and what you don’t know. I don’t have a magic wand, I have years of experience grinding down problems until they submit.

ohyes··on Mag 7 starting to underperform [pdf]
as soon as you start calling a group of stock tickers the “magnificent 7” they’re destined to underperform, as that’s a feelings based assessment and many investors will continue to buy the feelings long past the value being fair.

“Do they make money? I don’t know but I know they’re magnificent!”

ohyes··on Mag 7 starting to underperform [pdf]
“When the stocks don’t go up they don’t match the market which generally goes up”
ohyes··on I used Claude Code to get a second opinion on my MRI
I mean, technically… I’ve lived in that world. You just go back 25 years. We had the Dewey decimal system (card catalog) and the library, hard copy encyclopedias. You could also ask someone else.

Then we had computerized encyclopedias and search engines that searched the library.

I mean, you had to work for the knowledge. Sometimes you didn’t know something and no one else knew either, so you had to wait until you got a chance to find out, but you would think about it and sometimes you would be right when you found a reference source.

I’ll also note, Wikipedia is a secondary source. It is not a reliable source of truth. It is more like the ‘ask someone else’ alternative than anything else, it’s just ‘someone else’ is a person on the internet who writes Wikipedia articles.

ohyes··on I used Claude Code to get a second opinion on my MRI
Excellent point my colleague has the exact opposite incentive.
ohyes··on I used Claude Code to get a second opinion on my MRI
Sure, if we’re going to go that broad. People are already leaning heavily towards learning nothing instead of using Wikipedia.

I guess to me it has to be comparable to be an alternative.

Like, I don’t consider doomscrolling x an alternative to reading Wikipedia but I might consider it an alternative to CNN, even though they’re all technically and very broadly activities that I could use to inform myself.

In that same way I don’t consider the multitude of ways I could use my free will necessarily alternatives to each other even though they technically are. It kinda sucks but going that broad feels to me like it breaks the concept of alternative and makes it kind of meaningless.

ohyes··on I used Claude Code to get a second opinion on my MRI
The free alternative to Wikipedia is the library, not “don’t learn anything new ever”.

I find Claude is surprisingly similar to a confident but incorrect coworker, with the benefit that Claude will reevaluate when I correct it.

ohyes··on An Ohio Valley 100k-watt FM signal is severed in broad daylight
The solution is inexpensive drugs for addicts, minimum standard of living, programs for getting off drugs. But there’s an incentive to make drugs expensive and people desperate, and to punish people for their “failures” rather than forgiving and helping.

It is cheaper to avoid the situation by structuring society in a way that people aren’t willing to steal copper for quick money.

ohyes··on LLMs are eroding my software engineering career and I don't know what to do
LLM is a powerful tool but it still doesn’t have the context that a person would have. A million tokens is a drop in the bucket compared to the overall context that the person guiding the LLM needs to keep it on track and being productive.

If you’re not a good engineer and you don’t have the domain knowledge, your token costs will be very high for whatever gets shipped, because you won’t be able to provide the context necessary to prompt machine efficiently.

Claude will still very often hallucinate bugs, explanations, domain requirements, that have no basis in reality. It will offer fixes and improvements that are pretty standard but not optimal. This is correctable if you catch it, but you need to review every line of code and comment, because in addition to being obviously wrong, it is often very subtle in the wrongness. For every bit of “slop” there is almost microslop, the places where it just kind of confidently guesses… and doesn’t tell you… but sometimes is correct anyway.

The “problem” is there’s less low hanging fruit. You have to know a lot to add value beyond being a middleman gating the slop. You have to really pay attention to the details to find some of the errors that it’s making.

ohyes··on A GTA modder has got the 1997 original working on modern PCs and Steam Deck
I remember when there was a kid who kept installing this and the Chex doom on the school pcs in 7th grade. GTA was pretty controversial as a game even though it is incredibly tame by today’s standards. I’m pretty sure they never caught him.
ohyes··on On the Existence, Impact, and Origin of Hallucination-Associated Neurons in LLMs
People have a strong financial incentive to not understand this. It’s subprime mortgages.
ohyes··on IBM CEO says there is 'no way' spending on AI data centers will pay off
20 juniors become some % of 20 seniors. and some % of that principals. Even if it lives up to the claims you’re still destroying the pipeline for creating experienced people. It is incredibly short sighted.
Page 1 of 33Next →