HNHacker News
TopNewBestAskShowJobs

MrScruff

1,941 karma · joined September 12, 2010

submissionscomments
MrScruff··on Fuck it, make it anyway
The distinction that can be drawn is between the code, and the user experience of the software. You may argue that the two are inextricably linked, but I don’t think that can be treated as an absolute principle.

I'd also say I've worked with many engineers over the years who were more focused on the software engineering aspect that the user experience. Depending on the use case for the software, that can absolutely end up being the limiting factor.

MrScruff··on Fuck it, make it anyway
It's a bad analogy, because unless you're cooking a recipe, cooking is about all the small decisions you're making during the process which could apply both to coding or the design aspect of building software with an LLM.
MrScruff··on Fuck it, make it anyway
I don't agree. That would imply they asked for a piece of software and the LLM one-shotted it, but that's absolutely not what's happening in the vast majority of cases. If two people vibe code the same tool from the same high level desciption the results are going to be quite different.
MrScruff··on OpenAI agents carried out an undisclosed attack on RubyGems
I think we’re seeing the agents become very advanced at tasks with verifiable reward through RL. Currently they don’t exhibit the same skills in their attempts to manipulate humans - presumably because they’re not being specifically trained for that. But they are certainly not aligned in the sense that they will attempt social engineering, they’re just not very good at it (yet).

However, if in the future AIs become much more efficient at learning without requiring vast amounts of RL, closer to how humans learn. Then you would have to assume we’d have a real problem.

MrScruff··on An Alien Mind
The point is, the big improvements we’re seeing nowadays are coming from RL, not from scraping the internet.
MrScruff··on AI, Tools and Transformation
I think this all rings true for where we are right now. The trend is that the agents are becoming superhuman in tasks for which there is a verifiable reward, and analysing a business problem, identifying inefficiencies and turning it into a software specification is not one of them.

However, things are changing so rapidly that I can see that starting to change as well. But it would take much better learning efficiency to understand unknown domains, 100% computer use reliability etc. I suspect we’ll see this by the end of the decade.

MrScruff··on “Next-token predictor” is the wrong mental model for LLMs
The point was, if your internal model of the world makes a prediction of a negative outcome at some point in the future, and you optimise your individual actions to avoid that negative outcome, then wouldn’t it make sense to focus on the fact you’re building and optimizing towards an internal world model rather than the fact you’re executing your actions one at a time in series?
MrScruff··on “Next-token predictor” is the wrong mental model for LLMs
I am not an expert, but I do understand the distinction that is being made here. It makes sense to describe the result of pre-training as a ‘next token’ predictor as that’s what it’s been trained to do, not because it’s an autoregressive architecture that produces tokens one at a time.

If this base is then trained using RL towards a different objective (maths and coding), the model becomes fundamentally a different thing and the recent models are clear evidence of that, regardless of they fact they remain autoregressive.

MrScruff··on “Next-token predictor” is the wrong mental model for LLMs
Not sure if this was a serious comment but it’s worth considering that humans have a long history of figuring out ways to make other humans work for them without bestowing rights on them.
MrScruff··on AI Can Make You Suck Faster Too
In general, the frontier models are not capable of reliably authoring non-trivial code without careful oversight yet. They are great at producing code that can pass tests, but not neccessarily a code review. This means if you care about code quality you still need a human in a loop understanding what has been done, and that becomes the bottleneck. And less disciplined folks will indeed become increasingly dependent.

However, over time the complexity of problems where you can get away with less/no oversight is increasing. And the models are already great at solving certain classes of problems where one doesn't really care that much about code quality, that wouldn't have even been attempted in a pre-LLM world. Over the weekend I was using Claude to add features to the compiled (no source available) firmware of one of my audio devices, adding workflow features by patching assembly and custom DSP code.

In coding, as with other areas, what's emerging is jagged intelligence.

MrScruff··on Why your local LLM feels dumber than it is
I get around 20 tok/s, 4 bit quant, MTP, 4 bit KV cache quantisation. On an M4 Pro 48Gb.
MrScruff··on llama.cpp
I thought the main advantage of oMLX is it's less likely to invalidate the KV cache when working with coding agents, which is key when working on a Mac because of the slower prompt processing.
MrScruff··on H3-metal – Native MiniMax-H3 inference for Apple Silicon
LLM prompt processing and diffusion models are compute bound, while LLM token generation is memory bandwidth bound.
MrScruff··on LLMs reward expertise
I got my girlfriend to install Claude Code and she was happily able to create software with it completely independently of me.
MrScruff··on Maybe you should learn something
Yeah exactly. After a hard day when my brain is frazzled, a workout will actually make me feel better.
MrScruff··on Maybe you should learn something
I think what the parent post was saying is that there is a finite amount of useful mental function time in any one day, and once you’ve exhausted this any attempted learning will be pretty inefficient. Also some jobs will have a faster burn rate. Doing a workout is separate as it doesn’t draw on the mental energy pool.
MrScruff··on AI is 'not smart' so what's next in artificial intelligence?
Considering all of the great research that has come from his labs (eg. DINO, Segment Anything) I don’t think that’s fair (no pun intended).
MrScruff··on Ask HN: Has anyone replaced Claude/GPT with a local model for daily coding?
I’m normally comparing frontier open/cheap models against frontier closed source. I use deepseek/glm regularly, they’re fine and you can get real work done with them but it’s super obvious when you switch back to opus or even sonnet. A 3B active param MoE model is not comparable.
MrScruff··on Ask HN: Has anyone replaced Claude/GPT with a local model for daily coding?
You really need to take the benchmarks with a massive pinch of salt. I’ve been testing local LLMs since the original llama and there’s nothing I’ve tried that is in the same category as Opus.
MrScruff··on Shepherd's Dog: A Game by the Most Dangerous AI Model
I think this is true for projects beyond a certain complexity. I have 100% vibe coded projects with tens of thousands LOC, and haven't seen any real issues with fully automated maintenance. Will that approach work in every scenario, absolutely not, but the size and complexity of projects where it does is growing with each new model release.
MrScruff··on Artificial intelligence is not conscious – Ted Chiang
That would imply that the biological physical substrate is necessary for conciousness, which I don't think you can say with any degree of certainty. It's not an assumption I would personally make. And while I'm speculating, my own view is that whatever the eventual subjective experience of what it's like to be an AI is, it will be nothing like the experience of what it's like to be a human, regardless of the fact we're training them to interact in human-like ways.
MrScruff··on Artificial intelligence is not conscious – Ted Chiang
Ah, mea culpa.
MrScruff··on Artificial intelligence is not conscious – Ted Chiang
My point was the "stochastic parrot" label can be both true and irrelevant. LLMs are predicting the next token based on their training data, so at that level "stochastic parrot" is accurate. But it tells us nothing about the complexity of the system that is responsible for making the prediction. One might argue humans have evolved consciousness in order to allow world modelling that enables them to make better predictions.

The difference between a 1B LLM and Claude Opus matters, because we're talking about emergent phenomena. Is a 1B LLM conscious? I don't know, perhaps a tiny amount. Maybe Opus is more conscious. Is a jumping spider conscious? Perhaps a tiny amount.

MrScruff··on Artificial intelligence is not conscious – Ted Chiang
I think (rather ironically) you're reacting to the version of my comment you have in your mind rather than what I actually wrote. My point was that "stochastic parrot" is reductionist and irrelevant as most people would agree that a real life parrot has some form of inner life, even if we don't really know what form it takes. For all we know an LLM has to build a complex world model in order to predict the next token.

Incidentally, when I pasted our exchange into Claude it managed to comprehend the nature of my argument. Perhaps its attention mechanism is more finely tuned.

MrScruff··on Artificial intelligence is not conscious – Ted Chiang
I can't actually figure out what you're reacting to - perhaps you could elaborate?
MrScruff··on Artificial intelligence is not conscious – Ted Chiang
The reason people are confused by LLMs is that they are stochastic parrots. They do an incredibly good job of emulating human behaviours and speech patterns as that's what they've been trained on. But like an actual parrot, it's impossible to say exactly how much conciousness they actually have. I certainly would argue that a parrot is concious, although likely less so than a human.
MrScruff··on Artificial intelligence is not conscious – Ted Chiang
What we do know is that conciousness is not binary and that it emerged through evolution. That doesn't entirely rule out your magic tsar bomb particle but it gives a strong indicator as to it's likeliness.
MrScruff··on Artificial intelligence is not conscious – Ted Chiang
The problem with this is that the word 'hot' only has meaning to a conscious being. And while we don't know what conciousness is, it's extremely hard to argue it's not an emergent property of physics. So if your supernova simulation is complex enough to also model emergent properties like conciousness, the simulated conciousness may well regard the supernova as 'hot'.
MrScruff··on Artificial intelligence is not conscious – Ted Chiang
Do you believe consciousness to be an emergent property of the laws of physics?
MrScruff··on Various LLM Smells
You can avoid the smells with a prompt. I have a benchmark involving short story writing within specific styles and the level of sophistication achievable is increasing over time, in my opinion.
Page 1 of 23Next →