HNHacker News
TopNewBestAskShowJobs

RGS1811

886 karma · joined April 15, 2017

submissionscomments
RGS1811··on Clef: our open-source decision models
If everyone else can spin up their own version of your product in under a month, there probably wasn't much product there.
RGS1811··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
My main grievance is that any gap in specificity in my statement of a task would lead Opus 5 to invent an interpretation to fill the gap, frequently creating lots of extra work for itself in the process, and often deviating from my intent. This would happen even for very simple things. I once asked Opus 5 to fix a failing unit test in CI (something pretty simple), and it went off on a 45 minute expedition (all in one turn), read a boatload of unnecessary files, massively overcomplicated the assignment, etc. It fixed the test but previous models would have handled this much more straightforwardly.

A common form of this failure is the model picking up on random wordings from earlier in the session (e.g. some comment it made to me in the middle of a response, that I never explicitly endorsed) and then treating these as hard commitments. Or over-interpreting a specific word choice or clumsy phrasing as if it were a "load-bearing" constraint on the task.

None of this clumsiness would be so problematic if the model didn't have such a strong drive toward autonomy. It's much like with people: there's no shame in not understanding what you're being asked to do, provided you ask clarifying questions. There's no shame in ignorance if it's wedded to curiosity. Benchmaxing has RLVRed curiosity and clarification straight out of these models. It sucks.

RGS1811··on Open-ended data is a bottleneck for AI systems tackling real knowledge work
This is an ad.
RGS1811··on I can't tell who's teaching who anymore
I think the idea that there could be such a thing as a “style” filter misunderstands what style is. It’s not about avoiding buzzwords, it’s about tone, empathy with a reader, the layered and often unconscious will to express certain ideas with a specific flavor, and the idiosyncrasies that come from being a particular person who has lived a particular life with its own odd mix of influences and experiences.

Humans fail at good style more radically than LLMs do, and LLMs are on average better writers than people. But good style requires particularity and history in a way that a model could never achieve with just a filter.

The closest I’ve gotten to getting an LLM to write well was when I gave it an intellectual biography along with its writing assignment. But even that fails over longer context, because attention makes it very hard to build up narrative (even just conceptual narrative) in a compelling way that extends past the length of a short blogpost. Past a certain point you always end up back in the uncanny valley with that weird lack of cohesion and the unjustified contrastives, tricolons, and empty cliches.

RGS1811··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
I ran a battery of tests against a couple of simple prompts to check on thoroughness and verbosity of every available Opus, and 5.5 is a lot closer to 5 than people are letting on. 4.6 remains the best in terms of getting to the point and just doing what you ask. I had switched from 4.7 to 5.5 as my main claude model, but started running into the telltale over-interpretation issues of the 5 series, and have switched back. Something in their RL pipeline has made these models consistently worse IMO.
RGS1811··on Ask HN: What are you reading?
I just finished Wilfred Sellars “Empiricism and the Philosophy of Mind”

I’m slowly rereading Wittgenstein’s Philosophical Investigations.

I am perpetually rereading the works of Richard Rorty.

RGS1811··on 2026 in LLMs (So Far)
This was a fun read. I appreciated the recurring attention to “Deep Blue“ and laughed at this part:

> These projects were quite useful, in that they sort of cured me of my AI mania... because after I built these things, I got to look at them and ask “does the world need a slow, buggy, half-baked Python JavaScript interpreter?”

> I don’t think the world does.

Very relatable.

RGS1811··on Sonnet 5.5
This made me a bit emotional. Thank you for the beautiful response.
RGS1811··on Sonnet 5.5
> Letting AI basically control everything and us not having any mental understanding of sorts and sort of becoming the meat-proxies just for economical reasons seems realistic possibility but a bleaker reality at that. I am left feeling a little bit uncomfortable if this reality turns out to be true.

For the past year I’ve been yo-yo-ing in and out of existential despair about the future of civilization depending on how I feel the answer to this question looks. It’s emotionally exhausting, on top of everything else, and I wonder how others are coping with it aside from denial and cynicism.

RGS1811··on The problem is not AI code, but not knowing about system architecture or intent
The word I have for this is “commitment laundering”. People pass around AI artifacts that nobody has necessarily read or considered, and the invented assumptions and tagalong commitments just keep piling up. They look like they’ve got real provenance but nobody can say what is being done or why.
RGS1811··on Prompting Claude Opus 5.5
So much AI discourse tacitly assumes that there's an objective quality that corresponds to "being smart", rather than a chaotic patchwork of extremely contextual social practices and expectations.
RGS1811··on Parley: Federated, decentralised chat that speaks plain IRC
I've done this a few times. If I had to gesture at the root cause I'd say: LLMs lack curiosity and are trained to complete assignments, not to question their validity, so both the research phase and the questioning of intent tend to get short shrift prior to building anything.
RGS1811··on The Test
The public/private distinction Kant introduces is interesting, but it's not what I'm getting at here.

To draw a broad strokes historical parallel you could say this: in pre-Enlightenment Europe, responsibility for reasoning was signed away to various authorities (the church, the prince, etc.), "liberating" common people from the need to reason or ask questions or make self-directed use of their intellectual faculties. The vision of enlightenment Kant puts forward gives us an alternative ideal: maturity and freedom come from the self-directed use of one's intellectual faculties, the choice to take on responsibility for one's beliefs oneself, even if the freedom to think and discuss is constrained within the confines of public law. The ideal is inquiry, questioning, the desire to understand things from first principles and the refusal to accept the dictates of external authorities as given without being able to participate in the reasoning that justifies their claims.

If Huang, Altman, etc., are offering us all a future in which nobody is indepdendent minded enough to be able to perform arithmetic or even know their own address, what they're offering is "liberation" backwards into the world of intellectual immaturity, into a kind of serfdom disguised as leisure or carelessness. Maybe some people want that, and maybe they can even argue that it's good, but it's worth being clear about the fact that "liberation" from the need or desire to know or reason about the basic facts of one's life is not empowering, and does not yield greater personal independence or maturity. It is, in Kant's terms, anti-Enlightenment.

RGS1811··on The Test
This post resonates with me but for reasons not given in it. It calls to mind Immanuel Kant's famous essay "What is Enlightenment?" Where he says this:

> Enlightenment is man's emergence from his self-imposed immaturity. Immaturity is the inability to use one’s understanding without guidance from another. This immaturity is self-imposed when its cause lies not in lack of understanding, but in lack of resolve and courage to use it without guidance from another. Sapere Aude! “Have courage to use your own understanding!”--that is the motto of enlightenment.

RGS1811··on Gemini 3.8 text-to-speech
For the voice design, these don’t support it, but for the final render, they’re much better. So your pipeline could for example generate voices with one tool and render with another.
RGS1811··on OpenAI breaches Medicare, Albanese reveals
We can only guess how many of these incidents actually happened.
RGS1811··on Gemini 3.8 text-to-speech
I've been working on a similar project all year and as a tip, you should try Fish Audio or Higgs as a replacement for Qwen3. Both yield much better prosody and are much easier to listen to for long runs.
RGS1811··on Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max)
"This is a classic test request..."

I know there's been discussion about whether pelicanmaxxing is happening, but this is at least evidence that Claude was explicitly exposed to this problem.

RGS1811··on Claude Opus 5.5
This question is real.
RGS1811··on Claude Status – Elevated errors for multiple models
I'm using DeepSeek's own API with opencode. The pricing is absurdly good.
RGS1811··on Elevated errors for multiple models – Resolved
I did this with DeepSeek a couple of days ago.
RGS1811··on Show HN: An e-ink frame that hears birds and draws them as 1800s illustrations
This is one of the most wonderful uses of AI I've seen in a while.
RGS1811··on We must pace the frontier
No, my argument is that the need to pace the frontier means they cannot safely advance in capability due to liability concerns. The competition is already almost caught up. If OpenAI/Anthropic have hit an upper bound on safe capability improvement, the gap will close all the way, and we will have reached the full commoditization of LLM tokens very soon.
RGS1811··on We must pace the frontier
Alignment just means "this machine operates in ways that align with the intent of its users". It doesn't imply anything about the inscrutability of the machine in question. A gun with a misaligned scope would likewise fail to operate in accord with its user's intent, and likewise with potentially deadly consequences.
RGS1811··on We must pace the frontier
I don’t understand all the comments assuming that RSI is the real threat here. Dario is admitting that they failed to solve alignment. Without alignment, further improvements in capability turn LLMs into wanton felony generators. This call to pace the frontier is dressed up as altruism but it’s an admission that they cannot produce a marketable product better than what they have. Pacing the frontier means the US labs have lost their moat and are dead in the water.
RGS1811··on A misalignment of AI in mathematics
It's not about job security. It's about the social and intellectual practice of the discipline.

The threat to mathematics isn't that suddenly the profitability of their profession (lol) is going to go away, it's that people are thinking of AI as a replacement for the human social and intellectual practices that constitute the discipline.

RGS1811··on Claude, change the “Add to Cart” button to blue
`/model claude-opus-4-7`
RGS1811··on Muse: Meta's personal AI agent, features and capabilities
I get this, and I’ve said similar things about privacy and security in the past, but Meta in particular has a history of intentionally using customer data and behavior to the detriment of those customers. Creating experiences designed to ensnare, addict, and ultimately harm their users. This goes beyond paranoia about security or privacy. It’s handing your drug dealer the keys to your house.
RGS1811··on Muse – Meta’s personal AI agent
I have a hard time understanding, given the abundance of bad faith and predatory behavior Meta has displayed over entire span of its existence, why anyone would want to consult with its "personal AI" or grant it access to "all aspects of your life".
RGS1811··on Why the AfD Wins
By your own admission, (top of your first comment) you don't know what you're talking about. I never implied you aren't aware of the Holocaust. Thinking that the Holocaust is the only history to know here is indicative. You don't seem to understand the history of the BRD or how its constitution is structured, and you clearly don't know the internal history of the AfD, which started as a mixed center-right libertarian anti-EU party and has become more and more extreme over the past decade. You being angry doesn't make anything I've said wrong.
Page 1 of 3Next →