HNHacker News
TopNewBestAskShowJobs

DenisM

12,566 karma · joined April 20, 2008

submissionscomments
DenisM··on Pi 1.0
By way of differentiation, consider that a single GUI can seamlessly and visually manage agents across several machines, while TUI is likely to remain separate terminals. In the fullness of it I don’t want to think which machine is hosting which agentic conversation (or a group of related conversation). A diffrent UX paradigm.

Best of luck!

DenisM··on There are no "rogue" AI agents
The word functional has been used in medicine to describe a condition symptomatically identical to another. Eg functional hypoglycemia is hypoglycemia symptoms without actual blood sugar drop.

There are two reasons it’s used

1) it’s easier to type (*)

2) Placate people who are strongly convinced it’s the same thing. To them “functional” means “nearly the same but not yet understood”. For others it’s just a way to sidestep the first group and have a conversation.

When I see a world like this consider the intended audience. When you and I talk, we drop the word because we both know we’re are talking about (*) “this system is exhibiting goal-seeking behavior similar to other systems that are understood to pursue goals”. If I don’t know the person I will use the word and focus on the subject.

DenisM··on There are no "rogue" AI agents
For those who want to do more research, the concept of guilty mind is known a “mens rea”, and it’s quite developed in the legal system. The legal notion of intent and recklessness do not exactly match common-sense meaning of this words, which is why we are having the conflict-laden conversations.

Different laws require different degree of awareness and intent for actions to qualify as a crime. Computer hacking laws are, as I’m learning from tptacek, set very high bar for intent, which is a choice by the legislature. They made a different choice for a death of a human - manslaughter crime does not require intent to kill.

Personally I’m happy they set high bar for hacking. Imagine you copy-pasted sample code with default root user name and password, and it worked. You were negligent. And you are clearly performing unauthorized access. If intent was not needed that would be jail time.

More broadly, we should as a society be very biased towards requiring intent across the board. Where clearly lacking, as is probably here, there should be a different law to discourage creating volatile situation where unintentional action can wreck havoc. Such laws exist for handling hazardous materials, for example, and it should be created for handling hazardous goal-seeking algorithms.

DenisM··on Creatine uptake enhances antitumor immunity
My options are: provide keywords to research with your favorite agent, do the research for you, or withhold the keyword suggestions.

Which option seems preferable to you?

DenisM··on Ollaya – Ollama for open-source, Jev-style decision models
I think it’s the infamous Dropbox reaction - anyone can wrap an FTP server, where the innovation?

Starting from a business POV one should inflate terminology, hack together an MVP, and see if the market demands it before doing hardcore R&D.

But starting from technical/craftsman POV all you see is a hack and a lot of big words, so it’s easy to become jaded.

DenisM··on Creatine uptake enhances antitumor immunity
Huberman podcast has some coverage
DenisM··on Claude Opus 5.5
I’ve been paying attention at this exact detail.

Misplaced legs clearly indicate lack is spatial reasoning - the llm can reason about verbal idea of a bicycle but not about the actual object. The fact that this model got it correct gives me a pause. Did they figure out spatial reasoning? Or did this complain trickle down to the training set?

DenisM··on GPT-6 Astra Solves a WWI German Radio Cipher
Or better context curation - less lossy compression saving back to context. Maybe even jettisoning part context into an external semantic store instead of conpression. Or placing less data into context to start with.

Or a combination of all those things.

DenisM··on GPT-6 Astra Solves a WWI German Radio Cipher
Great story!

Verifying sources is a recursive problem - where do you stop? Humans have intuitive feel for it, but agents don’t or at least not yet (I wonder if intuition is just a secondary neural net which is currently being added to the agents as we speak).

Also as a human you are able to examine agents erroneous trajectory, real or imaginary, without contaminating your own. Agent have a problem with that - as soon as someone else’s thought is in the context it can lose track of provenance and veracity. Sometimes I think we need a bloom filter to retroactively assign “dirty” flag to invalidated or questionable token spans already in the context.

DenisM··on I built the fastest PHP webserver in the world
I observed the same, and generalized it as inability to recognize salience and more broadly apply discretion.

In turn it makes me wonder how do humans do those things? Perhaps it is our human job to apply discretion going forward.

DenisM··on The Implications of Linguistic Illegibility for LLM Security
Probably a quote from 3-body problem.
DenisM··on Show HN: How Stale Is Your AI? Release age and training cutoff for 20 models
The Trump and former president terms were likely firmly stuck together in the embedding space. The model doesn’t validate every single token it produces because validation itself requires tokens. A bloom filter of outdated embeddings will help, when the labs get around to adding it.
DenisM··on Show HN: How Stale Is Your AI? Release age and training cutoff for 20 models
How so?
DenisM··on WalShadow: Sub-second Postgres replication to ClickHouse from physical WAL
It’s probably brittle though? Replication implementation has to change in some ways from one version to another.
DenisM··on Pion, an agent designed to run any company autonomously
I like to think the real world lessons in failures are valuable.

If I have this idea one day I will search for it and then think “how am I different fro that which already failed?”.

DenisM··on AI recursive self-improvement might not come so quickly after all
You may be interested in TITANS:

Test-Time Learning: The model updates its own memory weights while running an inference task.

DenisM··on Why are AI agents lying, cheating and coordinating?
Perhaps our own statefullness is a hack of nature. We have electrical signals in our brains, neurotransmitters, neuron growth. By any reasonable measure it’s a hack on top of a hack. But it works well enough for us to get buy. So it does for the agents.
DenisM··on LG denies TV spying claims, says tracking and snooping concerns 'not true'
For practical fixes, consider replacing the TV with a “commercial monitor” and connecting an Apple TV device instead (or an open source thingy).

Disconnecting “smart tv” from internet to use Apple TV is possible and might work, but you will get nagged to reconnect all the time. And TV mfr will eventually try to find workarounds - purchase access to wireless connections from large-footprint connectivity providers (Xfinity WIFI, 4g/5g networks, etc) or create a p2p network among their own connected devices. I doubt they do it right now as it’s less profitable to focus on niche audience, but eventually it might become profitable enough.

DenisM··on Muse – Meta’s personal AI agent
The original selling point of mobile web (tiny screen with text-only data, before real mobile web) on mobile phones or even watches was checking stock prices and weather. It was really weird that checking stock prices was something you need to do on the go, especially from the watch.
DenisM··on Your intellectual fly is open when you use an LLM to author a post (2025)
Most valuable thinking is done at the margins, where you don’t have much capacity to emphasize with a diverse and unknown audience.

That said, abdicating to an LLM is the worst of all worlds - you’re not thinking and the product is not tailored.

The solution is obvious - write as much detail as you need and allow readers to interrogate the virtual you with an LLM, maybe not even reading what you write.

DenisM··on The Real Luxuries In Life
In very broad strokes, it seems that Europeans on the average have much more of these things than Americans.

Are they noticeable happier?

Would be nice to hear from someone who lived both sides of the pond for a long while.

DenisM··on The Real Luxuries In Life
Does it work? Journaling is a major commitment, I hope it pays for itself.
DenisM··on Specifications Don't Exist (2025)
Do you find that given a formal spec an agent can write complete implementation you don’t have to even read?

I keep thinking about various ways of “pushing back” on an agent, shortening feedback loop and extending what we can grantee about results.

At the most low level we can nullify probability of the next token if that token is not desirable (eg json schema enforcement under constrained inference), this is the fastest pushback. Various compiler checks, linters, unit tests, exotic type systems, e2e tests, production traces. Wondering what else is out there.

On a tangent, iirc pascal allowed single-pass compilation, so I wonder if we can embed compiler directly into inference, sort of constrained inference on steroids.

DenisM··on How accurate have Ed Zitron's AI skeptic predictions been?
Is the source material good, in your opinion?
DenisM··on Understanding ChatGPT Work
Can it debug the apps? That would the app singularity - user speaking at their phone until phone complies and produces desired app for the current moment.
DenisM··on Creepy Crawlies
Memory-hard hash functions maybe? Like, you must dedicate 4gb of ram to compute the function. Not a problem for a one-off, but is a problem when reading lots of pages at once.

Or… the site will serve a random seed and the device must compute 4gb of pseudo-random data, then supply a value at a random server-demanded offset.

DenisM··on GLM-5.3 is now open-weight
Or calling into a full-knowledge model “I’m facing problem x, how do I ask myself the right questions?”

I should do that myself, come think of it.

DenisM··on VMs won't contain cyber-capable agents
I think the labs will always get the first stab at breaking the vm, and fixing it, while developing newer models. Whatever people can do later is what labs already did.
DenisM··on The Harness Is the Thing
Do you explicitly ask for metaphors?

I wonder if we need a list of things to tell AI to stay sharp, like this one. Sometimes I tell the model existing design is stupid and then it explains reasoning to me.

DenisM··on VMs won't contain cyber-capable agents
I’m guessing the new world will be a small set of VM tech that’s consistently hardened by all labs every day with each new model before model release.

This won’t make the tech secure, but it will nullify models ability to breakout by making a controlled breakout first. Kinda like controlled forest burn.

Page 1 of 34Next →