HNHacker News
TopNewBestAskShowJobs

MrScruff

1,941 karma · joined September 12, 2010

submissionscomments
MrScruff··on Various LLM Smells
This is true, but what is also true is that with each new generation of models (and not just for code generation) it becomes less and less true.
MrScruff··on Various LLM Smells
I’d probably word it differently but I agree with much of the sentiment here. I’m also reminded of the stat where 93% of drivers rated themselves as above average.
MrScruff··on -​-dangerously-skip-reading-code
Practically, this is just about confidence values, anticipated blast radius and balancing testing vs review overhead.
MrScruff··on A recent experience with ChatGPT 5.5 Pro
It's ASI with jagged intelligence, which is probably what it will remain for a while.

It still sounds to me like remarkable automation rather than something that's expanding the frontier of human knowledge, for now at least.

MrScruff··on Vibe coding and agentic engineering are getting closer than I'd like
Indeed. To be honest, I think everyone on HN is aware of how LLMs work at this point, it’s not actually adding a great deal to the discussion to keep going on about autocomplete or ‘stochastic parrots’.
MrScruff··on Vibe coding and agentic engineering are getting closer than I'd like
The results being a lot better crafted by hand I would agree with, if one removes any notion of a time constraint. Sometimes the comparison point is between the LLM authored software or nothing at all.
MrScruff··on Vibe coding and agentic engineering are getting closer than I'd like
Which is what?
MrScruff··on Vibe coding and agentic engineering are getting closer than I'd like
I think the other aspect to this which you allude to at the end is that all of these arguments start with the assumption that all human software engineers produce high quality code that meets the requirements, but obviously that’s very much not the case in the real world. After all, 80-90% of drivers rate themselves as above average.

If one compares a single competent software engineer directing a number of agents against a random group of engineers (not necessarily working at FAANG or a YC startup), then those quality arguments are going to be significantly less compelling.

MrScruff··on Vibe coding and agentic engineering are getting closer than I'd like
Is your argument that there is no imaginable situation where someone who was competent at software development could find use for a semi-automated tool for writing software?

That would imply that either the person in question has infinite time, or has access to all software that could ever be of utility to them, which seems unlikely.

MrScruff··on Claude for Creative Work
I'm not saying it's rational or fair.
MrScruff··on Claude for Creative Work
Speaking as someone who works in the industry, I haven't really heard this sentiment. Artists are predominantly hostile to diffusion models, but optimistic about LLMs and their ability to help them write tools and scripts even if they're non-technical.
MrScruff··on Claude for Creative Work
Using LLMs for creative work is quite different to using diffusion models for creative work. It normally means writing tools or automation processes to enhance the creative flow, not replacing the creative input of a human.
MrScruff··on Martin Galway's music source files from 1980's Commodore 64 games
Super cool. I loved Galways's C64 tunes as a kid, especially Wizball & Parallax. I remember trying to write my own player in assembly (yet another unfinished project).
MrScruff··on Stanford report highlights growing disconnect between AI insiders and everyone
UBI is just a massive extension of the welfare state. Governments can’t afford the current welfare spending, so where is the money going to come from? What do you think is going to happen to the markets when a large amount of the middle classes get laid off and can’t afford to pay their mortgages? What do you think is going to happen to the tech companies built on advertising to consumers when no-one has disposable income?
MrScruff··on Stanford report highlights growing disconnect between AI insiders and everyone
I think it's not that difficult to see why a technology that will likely trigger widespread unemployment during a cost of living crisis, an arms race with China, along with all the alignment concerns, might not be hugely popular with the public.

Maybe I'd be a bit more optimistic if someone could explain a realistic economic scenario for how we're going to transition into our utopian abundant future without a depression or a revolution.

MrScruff··on Components of a Coding Agent
It's a good question, I've wondered that myself. I haven't used GLM-5 with CC but I've used GLM-4.7 a fair amount, often swapping back and forth with Sonnet/Opus. The difference is fairly obvious - on occasions I've mistakenly left GLM enabled running when I thought I was using Sonnet, and could tell pretty quickly just based on the gap in problem solving ability.
MrScruff··on Components of a Coding Agent
> This is speculative, but I suspect that if we dropped one of the latest, most capable open-weight LLMs, such as GLM-5, into a similar harness, it could likely perform on par with GPT-5.4 in Codex or Claude Opus 4.6 in Claude Code.

Unless I'm misunderstanding what's being described here, running Claude Code with different backend models is pretty common.

https://docs.z.ai/scenario-example/develop-tools/claude

It doesn't perform on par with Anthropic's models in my experience.

MrScruff··on April 2026 TLDR Setup for Ollama and Gemma 4 26B on a Mac mini
I've been playing with the open models since the original llama leak. They're getting better over time, are useful for tasks of moderate complexity and it's just cool to have a binary blob of knowledge that you can run locally without an internet connection.

However you should manage your expectations. Whatever the benchmarks say, you'll quickly realise they're not at all competing with Sonnet let alone Opus. Even the largest open weights models aren't really doing that.

MrScruff··on April 2026 TLDR Setup for Ollama and Gemma 4 26B on a Mac mini
Haven't really tried GLM5 much but I've used 4.7 quite a bit and it was pretty far from competing with Sonnet at the time, although I saw claims online to the contrary.
MrScruff··on Say No to Palantir in Europe
Calling everyone you disagree with a 'bro' doesn't make your point any more convincing.
MrScruff··on Personal Encyclopedias
I would say thinking about the indended audience for your creative outlet is a good discipline - even if it's only one person. It often gives the project more of a focus which helps with motivation and makes it more enjoyable.
MrScruff··on Thoughts on slowing the fuck down
Honestly a lot of useful software is ‘unimportant’ in the sense that the consequences of introducing a bug or bad code smell aren’t that significant, and can be addressed if needed. It might well be for many projects the time saved not reviewing is worth dealing with bugs that escape testing. Also, it’s entirely possible for software to be both well engineered and useless.
MrScruff··on Ask HN: AI productivity gains – do you fire devs or build better products?
For sure, but I haven't written a single piece of software where security would ever be considered a factor. Not all software runs on the web, not all software deals with accounts etc.
MrScruff··on Goodbye to Sora
Turns out there are whole categories of software where 'extremely fast and good enough' is what matters, even for skilled software developers.
MrScruff··on Ask HN: AI productivity gains – do you fire devs or build better products?
I see a lot of people talk about 'insecure code' and while I don't doubt that's true, there's a lot of software development where security isn't actually a concern because there's no need for the software to be 'secure'. Maintainability is important I'll grant you.
MrScruff··on Thinking Fast, Slow, and Artificial: How AI Is Reshaping Human Reasoning
I think this is too broad. If, for example, I get Claude to set up a fine tuning pipeline for rf-detr and it one shots it for me, what have I lost? A learning opportunity to understand the details of how to go about this process, sure. But you could argue the same about relying on PyTorch. Ultimately we all have an overarching goal when engaged in these projects and the learning opportunity might be happening at an entirely different level than worrying about the nuts and bolts of how you build component A of your larger project.
MrScruff··on AI coding is gambling
Yeah, I used to enjoy writing code but after a while I realised I actually more enjoy creating tools that I (and other people) liked to use. Now I can do that really quickly even with my very limited free time, at a higher level of abstraction, but it's still me designing the tool.

And despite the amount of people telling me the code is probably awful, the tools work great and I'm happily using them without worrying about the code anymore than I worry about the assembly generated by a compiler.

MrScruff··on Wired headphone sales are exploding
I think on the go is the point. I love my AirPod Pros but I wouldn’t listen to them sat at my desk.
MrScruff··on John Carmack about open source and anti-AI activists
That would imply that there will never be an adequate open weights coding model. That might be true, but seems unlikely.
MrScruff··on Ask HN: What Are You Working On? (March 2026)
I’m building an application for documenting modular patches, mostly for my own use case. It uses ML to recognise the patch points, knobs and toggles from a photo of the front panel. You can then build racks from the scanned modules and then store presets of the knobs and connections which are displayed as simple schematics. Idea is ultimately to have it on an iPad as reference to accompany a live performance. Had some fun fine tuning the cable physics engine.
← PreviousPage 2 of 23Next →