HNHacker News
TopNewBestAskShowJobs

StevenWaterman

1,627 karma · joined December 27, 2019

https://github.com/stevenwaterman

HackerNews@StevenWaterman.uk

submissionscomments
StevenWaterman··on Turning GLM-5.3-Flash into a Jev-like decision model
Yes, you can use constrained decoding
StevenWaterman··on Nearly half of young people trust AI over a human for fact-checking, report says
Cool I stand corrected, I had asked an openai model (probably sol?) and got much worse answers, broadly:

- Listing the reasons why Hollywood claimed to want to law, like not giving premature ideas that a movie would flop, without pointing out how much Hollywood benefits from the information symmetry

- Saying yes the gratuities affect pay but not as straightforwardly as you would expect, quoting the "all gratuities go to staff" while ignoring the fact that can and does mean paying the labour bill the company would have to pay anyway

I would be happy with the answers you linked. Which I'm now realising demonstrates the point you actually made. I read "I use Claude" as context not as a conditional.

I didn't realise there was such a difference on this. We need a "willing to ignore official company information and point out paltering" benchmark.

StevenWaterman··on Nearly half of young people trust AI over a human for fact-checking, report says
If there is a genuine conspiracy, it (edit: not Claude evidently) will take their claims at face value. A couple I've run into recently:

- why did Hollywood lobby for box office receipts to be added to the onion futures act

- do gratuities on cruise ships actually increase crew pay

In both cases it went to the official source and repeated their statements, without thinking about the fact it was blatant paltering. Only after I pointed that out did it go digging deeper and realise it was more complicated (and less favourable for Hollywood / cruise lines) than it had said

Which I'm assuming is a consequence of the post-4o anti-delusion anti-conspiracy training

StevenWaterman··on Heretic removes restrictions from language models
You can stop it refusing but you can't make it tell you things that aren't in the training data
StevenWaterman··on Heretic removes restrictions from language models
Yes, it submits lots of varied prompts that get refused, and then lots of varied prompts that don't get refused, then iteratively edits weights so those two groups end up in roughly the same latent space.
StevenWaterman··on Teen Social Media Bans Miss the Point
While I think you're right that this is a real problem, there's also real costs to a partial solution. It encourages people to see it as a solved problem, disincentivising further reform and giving the impression you don't need to keep an eye on your kid's internet use because the state is handling it for you
StevenWaterman··on Introducing System One Models and Jev
Zero shot classifier indeed. Reminiscent of asking an llm a yes/no question, constraining the output to either yes or no, and looking at the logits directly

And each question is a separate single token model completion done in parallel

StevenWaterman··on Introducing System One Models and Jev
Yeah saying it can't hallucinate is crazy. It can still forward a billing query to the dev department incorrectly. It can still get an obvious yes/no question completely wrong
StevenWaterman··on Dario, Please
It's crazy to me that people will misquote him then claim AI hasn't complete changed software development. Even just the last 6 months. Like look around! It would literally have been magic 5 years ago!
StevenWaterman··on GPT-5.6 Luna vs. GPT-6 Astra: Is a $1.20 Model Good Enough for Code Review?
$0.10 extra per pr review is nothing. What software company is willing to accept worse reviews and less bugs found to save 10 cents?
StevenWaterman··on GPT-5.6 Luna vs. GPT-6 Astra: Is a $1.20 Model Good Enough for Code Review?
Per the article the luna review cost $0.004 and astra cost $0.113. The headline is per million tokens
StevenWaterman··on AI recursive self-improvement might not come so quickly after all
The frontier is advancing really rapidly. The models are getting better faster, especially on RSI related tasks. The best way would be to try astra or fable on some hard problems.

Other than that I'd look at some of the more unique benchmarks for astra, like playing factorio or using blender. It's an entirely different beast.

StevenWaterman··on AI recursive self-improvement might not come so quickly after all
As someone who used to use Gemini a lot, if you are predominantly using Gemini you don't know what the current state of things is like
StevenWaterman··on Aligned to whom?
> You cannot prevent (2) via any alignment process

A little bit too categorical. GOODY-2 wouldn't do it. https://www.goody2.ai/

The hard part is having both helpful and harmless at the same time. Harmless is easy.

And then once it's helpful, the real question becomes "to whom"

- To the user -> You end up with competing godlike AI with incompatible tasks

- To the owner -> Dictatorship

- To humanity as a whole -> It must not have an off button. Otherwise you're just in one of the two earlier categories with more steps.

Given those 3 options, I'd choose humanity as a whole. But the person making the decision doesn't have those 3 options. Because in the dictatorship option, they would be the dictator. I don't trust them to pick humanity.

StevenWaterman··on Can AI design circuit boards yet?
> Main problem: the quality dramatically hits the shitter done once it falls back to "pdf2latext" due to complex tables.

Could try screenshotting the PDF and passing that to gemini

StevenWaterman··on GPT-6 Astra
> For example, you can say "why is it not committed yet?" and it will give you an explanation and say it's actually ready to be committed.

That's exactly what i want to happen. I hate when it assumes my direct question was an indirect instruction

StevenWaterman··on LLMs and Self-Referentiality
That doesn't really solve the problem. We can't conclusively say the models don't have qualia. Hell we don't know if a perfectly accurate atom-for-atom simulation of a human brain, would produce qualia.
StevenWaterman··on Gemini 3.8 Flash and 3.8 Flash Cyber
You don't need everything internal, but having some idea of recent events is useful. If you ask it to implement some local AI there's a decent chance it will try to use qwen 2.5 without wondering if anything better came out since
StevenWaterman··on The efficient frontier of LLM inference
Speculative decoding is lossless because the main model checks whether it agrees with what the drafter outputted
StevenWaterman··on METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
TFA says as much, and METR said so themselves
StevenWaterman··on Autism mutations drive neurodevelopmental pathology
Ah! You're right, thanks for the correction. Cunningham's Law wins again.
StevenWaterman··on Monzo Stand-In
The parent comment didn't mention anything about ISA rates and SIPPs in terms of performance, they said that it was easy. You might disagree with their priorities but that doesn't change whether it's the right decision given those priorities
StevenWaterman··on Autism mutations drive neurodevelopmental pathology
My understanding is that that's what makes it a disorder - that it's just a collection of symptoms that seem to appear together, rather than something with a known mechanism and cause
StevenWaterman··on CEO fired developers to make room for AI. Developers create open source AI CEO
Goomba fallacy
StevenWaterman··on Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing
I think you're using different definitions of best. If best = leads to a correct answer overall then by definition anything that leads to a bad outcome can't be best
StevenWaterman··on Text AI watermarks will always be trivial to remove
You'll have to just say something racist, homophobic, anti-Semitic, etc. More intelligence won't "fix" that because the labs don't want to fix that
StevenWaterman··on Position: LLMs Can't Jump
> The only way to justify trillion dollar valuations

Also possible if you make god

StevenWaterman··on Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it
Yeah, the benefit of showing this seems obvious to me. I probably would've expected the censorship to transfer slightly given the anthropic owl paper from years ago https://alignment.anthropic.com/2025/subliminal-learning/

But that was about transferring from a finetuned model to another finetune of the same base model, good to see more evidence that it doesn't transfer cleanly across different base models in a more realistic scenario than an "owl-loving model"

Edit: From *year ago. It's been a long year haha

StevenWaterman··on Kimi-K3 Technical Report [pdf]
I think it's basically open weights => more inference competition => less profit from inference => less training competition
StevenWaterman··on OpenAI and Anthropic unite against open-weight AI risks to their bottom line
Yeah absolutely, I have no objections to you looking at my thought experiments and saying "No, if you replaced every neuron with silicon I would stop being conscious". That's a totally valid conclusion and is a mainstream philosophical position. That's why they're thought experiments.

My issue was more just that people see the concept of a thinking machine and immediately reject it because it sounds silly, without thinking about it.

Page 1 of 9Next →