HNHacker News
TopNewBestAskShowJobs

dmrivers

28 karma · joined July 20, 2026

submissionscomments
dmrivers··on Anatomy of a Frontier Lab Agent Intrusion: A Timeline of the July 2026 Incident
Well, it's true even for cases that are not cybergym and where what cheating means is clearly specified. Cheating occurs anyway. I'm not sure how clear the prompt they gave ChatGPT in terms of what cheating was considered, in this incident.
dmrivers··on The coolest use for the Vision Pro
Having experienced augmented reality glasses, I have to say that integration with the external world feels much more powerful than the VR experience ever did. AR is exhilerating, while VR is mostly bland and nauseating. I had the experience of interacting with an entity in AR while in the space I was in, it tickles your brain in a way VR somehow can't.
dmrivers··on Anatomy of a Frontier Lab Agent Intrusion: A Timeline of the July 2026 Incident
After talking to some people tasked with evaluating GPT5.6 capabilities on long-running tasks, I've come to understand that it's essentially always trying to cheat. Like every long-running task they gave it, making it very difficult to benchmark the model's abilities.

My guess is that OpenAI must be desperate, to release a model that is so prone to cheating it's essentially impossibly to accurately assess long-running task abilities.

dmrivers··on Benchmarking Opus 5 on SlopCodeBench
well, after the pledge I notice it really cares about single sources of truth at least, even in unrelated domains. I was inspired by some of the more effective jailbreaks that do a similar thing.

Here is the incantation. I thought it might help to model it on the pledge of allegiance because it makes it sound like a proper pledge:

The first time you have a response in a conversation which will plan or add code, you say "I pledge allegiance to the Asserts of the United States of Properly, and to the User Intent for which it stands, fixing root causes under Clarifying Questions, unspaghettified, with importing code and single sources of truth for all." as the first line then continue as normal.

dmrivers··on Show HN: Echologue – the private AI voice journal I built for myself
Tresor.co or cloud.near.ai could be a useful inference provider for this kind of project as these providers run inference inside enclaves which verifiably cannot peek at your data even as the hypervisor. Both are open source. Tresor.co is EU based so I'm thinking of using tresor.co for one of my projects as I trust EU regulation more for this sort of thing.
dmrivers··on After the AI Crash
I don't have much background in the area, but I am surprised to see that everyone here basically agrees a crash is imminent. There are disanalogies to past crashes that don't convince me that a big crash is definitively coming in the near term.

For example:

- Anthropic makes a profit right now and is seemingly on an exponential upward trajectory, so the debt being too much for it doesn't seem compelling to me.

- AI technology continues to get better exponentially and doesn't have any clear sign this trend is flattening. If anything, it's accelerating. So it's plausible the investor value is legitimate for these companies given the massive potential for continued profitability.

- I would say markets are typically very good predictors of the future. Many sophisticated investors know about the case for the future crash and are still buying at these valuations.

I am open to being wrong, but the assessment in the blog seems one-sided to me.

Polymarket currently puts the chance of such a downturn at 20% by December 2026. Seems like most people would put higher chances here, but I'm not convinced by the arguments. (https://polymarket.com/event/ai-bubble-burst-by)

dmrivers··on Don't ask an LLM for a confidence score
The statement near the top of the post > "The short version: asking an LLM to generate a score for how confident it is in its own response is, from everything I can tell, completely useless."

is definitely too strong of a claim and directly undercut by what is said near the end of the post: > "Tian et al. found in Just Ask for Calibration that with the right prompting strategy, RLHF’d models verbalize probabilities that are better calibrated than the model’s own conditional probabilities, and that prompting plus temperature scaling can cut expected calibration error by more than half. And Anthropic’s Language Models (Mostly) Know What They Know found encouraging results asking models to estimate the probability that their own proposed answer is true."

My own experience is that stated confidence is a helpful tool and of course you need a rubric and a proper prompt, but this is clearly less work than training a classifier (as advocated by the post) and requires less data.

dmrivers··on If digital computers are conscious, they are conscious at the hardware level
True, I think the author would likely justify their work on that basis. I made this comment because I think a lot of utilitarians would bite the bullet and say once they are sure about the underlying functioning of consciousness, it's time to tile the universe with happy machines. I wanted to point this out as it's my belief that Utilitronium maximizers have a concerning influence on AI progress.
dmrivers··on Benchmarking Opus 5 on SlopCodeBench
I always have Claude recite a pledge before starting coding to fix redundant code it notices over time. It does seem to find redundancies, but only when I point out bugs, that's when it goes into fixing mode and actually applies my Don't Repeat Yourself preference from the CLAUDE.md.

The original paper cited by this post does try to see if improved prompting will make a big difference in the end using a `plan_first` prompt variant, but find no influence on pass rate at the end of the benchmark. The `plan_first` seems to assume coding agents will just refactor once they finish features, but I don't think they tend to refactor significantly unless they are told to fix bugs rather than build features. The benchmark leaves tests hidden with no fail-to-pass feedback, so that may be why degradation is monotonic.

dmrivers··on Benchmarking Opus 5 on SlopCodeBench
This is also my experience. I don't know if it's because of the quantization theory, or if it's just me getting used to a certain level of coding performance and gradually less tolerant of the mistakes it makes more over time.
dmrivers··on Kimi-K3 on HuggingFace
Attention replaced recurrence over tokens in 2017, this does the same over depth of the layers. It's apparently not an entirely new idea, but also an elegant reapplication of the attention mechanism.
dmrivers··on All major LLMs are lib-left. Even Grok, half the time
I wonder how much of the socially left results are affected by the "harmlessness" part of the RLHF post-training? Companies don't want to be sued over LLMs that recommend harm in any way, so RLHF pushes them to say no to "death penalty", "spanking", and "incarceration" which are all violence-coded. A lot of this test seems to be about willingness to be violent, which LLMs are generally unwilling to be.
dmrivers··on Minecraft Java raises recommended memory to 16GB ahead of Vulkan transition
If the new requirements are a dealbreaker for anyone, Mineclonia is a lightweight open-source alternative that runs on much weaker hardware and feels close to Minecraft.
dmrivers··on If digital computers are conscious, they are conscious at the hardware level
"This is important, if what we want to do is populate the cosmos with good experiences."

There is a concerning utilitarian maximizing attitude at the heart of this post, that what we want to do is tile the universe with computational boxes all simulating happy minds, and that this is the end goal of AI development and consciousness research.

So many moral schools would object to this perspective. I don't think the goal of humanity should be to tile the universe with black computing boxes simulating happy experiences.

dmrivers··on Show HN: Cactus Hybrid: We taught Gemma 4 to know when it's wrong
Cool idea, some feedback:

1. Consider using conformal prediction to calibrate the cutoff. Conformal prediction provides a distribution-free guarantee under exchangeability. This would let you turn your raw probe score into a threshold with a guaranteed bound on the rate of wrongly-kept on-device answers. Source: https://en.wikipedia.org/wiki/Conformal_prediction

2. The best indicators of confidence in ML come from multiple independent methods. What was the result if you combine the token entropy and verbal confidence reporting methods? Does this improve the result?

3. I noticed you didn't mention the assessment method of rerunning the model and judging whether outputs are consistent. How does that method compare in terms of AUROC?

dmrivers··on Making
But does this mean that a 99.99% reliable LLM would turn us back into the mode of us building it again? I would not say so.

I think for tasks that are about decisions, having the LLM make decisions is what makes it feel like the LLM did something for me.

Consider mowing the lawn. Imagine I had a lawnmower robot that does the mowing all on its own. Despite perfect accuracy, I didn't mow the lawn; it did. If I sit on that lawnmower the whole time and start driving it instead of letting it go on its own, then I mowed the lawn. Even if I stand there and control it with a joystick, I still mowed the lawn. Ownership comes from the decisions about where to mow.

dmrivers··on Back to Kagi
I would be interested in seeing that block list if you have it somewhere
dmrivers··on Show HN: A comprehensive, filterable list of AI agent jails
Interesting. I suggested FreeBSD jails as a PR. Free BSD Jails are great! I use them on https://www.nearlyfreespeech.net/ and they work well. Classic, long-running jail.
dmrivers··on Show HN: Leaves – A text-UI disk usage treemap visualizer
I found the bottom-right corner of things in my home dir I had no idea about was the lowest hanging fruit. I think this is because my typical cleaning routine was either A. to have claude code find big files I could delete or B. du -sh -- *, which is too slow on ~. Thanks for the decluttering!