HNHacker News
TopNewBestAskShowJobs

cruffle_duffle

884 karma · joined November 27, 2023

submissionscomments
cruffle_duffle··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
Dude prompt your agent to set up always on remote connections via systemd… then you can drive from Claude or codex mobile apps natively over their native hookup. Works great.
cruffle_duffle··on Cf: The Agentic CLI for the Cloudflare API
You laugh now but in 5 years when these LLM’s operate at thousands or tens of thousands of tokens per second on dedicated hardware in your phone… “might as well” won’t even be a joke.
cruffle_duffle··on Opus 5.5 is good at explainer videos
True. I don’t think it takes much of a model thankfully. Luna or something…
cruffle_duffle··on Opus 5.5 is good at explainer videos
Why not replace the backend of your site with the LLM? Let it do the recommendations in structured inputs and outputs so the front end isn’t a chat interface but leverages the fuzzy nature of LLM’s.
cruffle_duffle··on Launch HN: Coverage Cat (YC S22) – Umbrella insurance via your personal agent
My biggest line item is my auto insurance and your agent tool can’t do that. And I rent not own and thus have renters not homeowners.

Basically, this sounds interesting in theory but it’s useless for me. I even tried letting codex go over it and it wasn’t clear without additional prompting what was supported.

Sorry.

cruffle_duffle··on GPT-6 Sol and Luna
It’s a hard problem because among so many other things…switching models mid session because “shit got real” (or shit is now just executional) costs cache.
cruffle_duffle··on Claude Opus 5.5
I mean the webpage can say what ever it wants. The proof is using it yourself.
cruffle_duffle··on Claude Opus 5.5
“It has the same annoying cadence and writing style with slightly less prominent claudisms.”

Seems like it based on my first session. It still does the whole “bury the important thing in a pile of words” coupled with the “it might actually be important” thing… so basically you never really know what it’s talking about.

Honestly I trust opus so little that the entire “opus” brand is completely tarnished. Its writing style is so god awful that it needs more than just a point release. Either dump the name and ship a different model entirely or at minimum call it “opus 6”. Calling it 5.5 makes it sound like it’s basically a continuation of the same garbage output that 5.1 had but with some minor adjustments. And based on my single first test, that is what it appears like to me.

cruffle_duffle··on Grok 4.7
> Part of intelligence is knowing your audience and communicating efficiently.

Bingo! And on this axis many SOTA models fail miserably. These things are acting on my behalf under my direction. All the supposed intelligence in the world means fuck-all if nobody can understand it.

And like somebody else said… when meat-based humans talk like Claude does, it almost always means they either don’t understand what they are talking about, or are actively trying to conceal something and are a fraud. Not always, but almost always.

cruffle_duffle··on MCP was always a bad idea?
I absolutely love when they do that unprompted. I once asked a Claude work agent to go pull comparable apartment listings and somehow it’s subagent reverse engineered like rentcafe or whatever’s private API to get the data. All unprompted.

So yes. Absolutely assume AI agents acting on behalf of their humans are finding all the token efficient ways to get at your sites data.

cruffle_duffle··on MCP was always a bad idea?
> JSON is a bad format/transport.

I've cut token counts significantly for some of my MCP servers by returning csv or tabular data or sometimes even just good old fashioned plaintext "template style" words. JSON is highly repetitious. In some instances like 30% or more of the token response was just boilerplate JSON.

Once you realize that to an LLM JSON is just a stream of tokens no different than XML, markdown, or a punctuation-less stream of conciseness.... you start to realize there is no reason not to just make up your own crazy formats. The LLM isn't throw parsing errors if your MCP is returning something other than JSON. The LLM is smart enough to figure it out. It's far more important to make sure your tool descriptions and parameter descriptions properly "market" what it is you do (and dont do) than return syntactically correct JSON.

Same with versioning.... it's all ephemeral. There is no backwards compatibility to speak of with MCP (at least for tooling "contracts".... what the tool actually does however... that is a different level and is product requirements not API contracts).

cruffle_duffle··on MCP was always a bad idea?
> The statefullness of MCP was such a massive detour for the industry that we will be cleaning up after it for years.

I remember discovering this when i wrote my first MCP server. It was like "huh? why would they do that? what use case did they have in mind?". We've spent decades making "internet shit" as stateless as possible on the backend because making it stateful is expensive and complex if you want to have any reasonable scalability. I mean good luck trying to host a stateful service on any kind of commodity serverless "scale-to-zero" infrastructure here in 2026.

Maybe it's because these AI-labs are used to statefulness. I mean LLM-based sessions are hugely stateful if you want any kind of reasonable caching to happen and caching is the only way you can economically scale out LLM's. Seen from that perspective it kind of makes sense why they'd look at MCP and think "hey, why not make this stateful on the backend as well". Statefullness just part of their DNA.

cruffle_duffle··on I don't like passkeys
Sure but who is reading account numbers off checks? They are copy and pasting from their bank app.
cruffle_duffle··on AI-generated posters don’t have to be horrible
If the hole in the wall teriyaki restaurant’s menu wasn’t hand painted by Monet I turn right around and walk out. I want nothing to do with an eating establishment that can’t at least spring for a realistic rendition of chicken yakisoba hand painted by Michelangelo himself.

Photoshop basically destroyed a centuries-old artistic tradition of depicting six gyoza on a white plate. Now with all this AI nonsense, you might as well just make your food at home rather than stare at yet another “six finger hipster eating a vegetable spring roll” cranked out by ChatGPT.

cruffle_duffle··on AI-generated posters don’t have to be horrible
I think this generalizes to a lot of professional work.

LLMs are very good at satisfying the requirements you give them. They’re much worse at knowing which requirements should be challenged, reframed, or ignored because you don’t know enough about the field to know what actually matters. That’s why they’re such force multipliers for experts: the expert can spot when the model chose the wrong abstraction or checked every box and still produced something bad. A novice often can’t. I’d argue the model mostly raises the novice toward “average.”

“Tell it to push back” doesn’t really solve this either. Too little and you get “aye aye captain!” Too much and you get Opus 5 / claudise, where it seems compelled to find something wrong with everything.

I see this in a Home Assistant project I’ve mostly vibe coded. I can ask Codex to make the kitchen brighter on school days and it’ll happily implement something. But maybe an HA expert would say “you really shouldn’t model it this way.” Maybe they’d even question HA itself. The model could conceivably know that too, but it only sees a tiny slice of my world. There are infinite side quests it could raise, and a good professional’s real skill is knowing which one matters enough to interrupt you about and which 99% to silently ignore.

cruffle_duffle··on Bend 2 and the Vibe-Coding Trap
Depending on what I’m doing I’ll dedicate a few deep research sessions to building a framework. It will generate some grounding docs that go into the repo and get consumed as we go. Said docs establish terminology, widely known formulas and methods, etc.

And yes the output of these researchers are highly sensitive to prompting. Left to their own devices the LLM will often ship some very biased prompts to its deep research agents loaded with pre-conceived ideas rather than letting the agents uncover things themselves. Then all the agents do is confirm what the prompt told them to rather then “think independently”. (Very similar to open ended interview questions rather than asking yes/no questions)

It’s is far better to spend a session writing writing the research prompt itself.

All of this takes time and tokens of course…

cruffle_duffle··on Replacing Pull Requests with Delta
Why? It sounds like a pretty typical workflow to me.

It just depends on what you are building. Some things are more more forgiving to muck-ups than others and if the LLM screws the pooch, oh well! Roll forward.

cruffle_duffle··on I don't like passkeys
“There have been sites that blocked the autofilling of passwords and so on as well.”

Sites that do this irritate me so much. Ones that try to block pasting and stuff… like somebody intentionally baked that into the site. Who? And what was their rationale? Are they really so arrogant to think people are going to carefully type in some elaborate password not once but twice?

That and blocking paste in fields like bank account numbers and stuff.

Surely somebody here has been asked to implement these mis-features. Please explain what went through the heads of the people responsible for it?

cruffle_duffle··on Claude Cowork and chat are now one Claude
> I do like Claude Code a lot when it comes to pure coding use cases

The harness is great, but opus is such an arrogant little prick that spews out unintelligible word salad. Opus 5 is so bad at communication it amazes me that somebody green-lit it. It's absolutely awful.

The fact that this isn't an acknowledged regression (and Fable 5.1 isn't much better) leads me to believe that people at Anthropic actually like Opus 5's output.

cruffle_duffle··on Claude Cowork and chat are now one Claude
"Traditional and boring works for me."

I wouldn't call copy & paste code from a webui of chatgpt either traditional or boring. I'd call it tedious, error prone and guaranteed to get poor results. There is much better tooling and harnesses to leverage now.

cruffle_duffle··on An update on Wayback Machine access
Then make agent friendly content. Take the text and make a markdown version.
cruffle_duffle··on An update on Wayback Machine access
Dunno why the downvotes. I feel that is reasonable as well. Owners that block that stuff are doing so only to their detriment.
cruffle_duffle··on Fable 5.1 Solves the Cyphral Distich, a 370-year-old cipher
You’d think that, I thought that… but then I realized I’m just kidding myself thinking its output makes sense. It doesn’t. It doesn’t. Sometimes it might as well just speak tongues.

In other words, it ain’t you. It’s the model. It’s just genuinely bad.

Then you switch to ChatGPTs lineup and realize how things can actually be better. It took about a week to really get the feel for how to use their models… then I basically switched. I’ll check in every now and then when they actually make a deal about how opus “now makes sense”.

But honestly I’m half convinced Anthropic actually prefers the output of opus 5. I dunno why, but how else could you explain how such a thing got shipped? I mean somebody in the pipeline had to say “dude this model doesn’t make sense, you think we should fix it?” Right? Like it’s a pretty massive drop in quality for such a major brand in this space, you know? How did it make it out the door?!?

cruffle_duffle··on We must pace the frontier
It’s the same thing with Covid. People thought 1 in 10 people would be dead. And believed it for years, despite it being orders and orders of magnitude wrong.

Humans love a good end-of-world story. Always did, always will. It’s when they want the world to conform to their ungrounded irrational bat-shit crazy nonsense that I get a worried. Especially when people in power actually take their nonsense seriously.

cruffle_duffle··on We must pace the frontier
Sorry but… uh… bro. I always wondered if opus 5 was a regression or intentional. I’m seriously thinking it’s actually derived from how people at Anthropic talk.

The last people on earth who should be regulating this are governments and tech oligarchs. I find that vastly more scary (and plausible) whatever unintelligible nonsense opus and fable spew out these days. Seriously. The way those models talk and behave do more to show the limits of AI than anything else.

Anyway. Touch grass.

cruffle_duffle··on Claude is only available to people over 18 years
Dude, being able to run high quality models locally with cheap hardware can't some soon enough.
cruffle_duffle··on Claude, change the “Add to Cart” button to blue
For one thing, it helps clarify my own thinking and uncover any gaps. For two, it lets me peek into its own understanding of the problem and ensure we are both aligned. For three, the garbage I spew at it is usually hastily typed crap from a mobile phone. So it helps clean it up and unpacks things.

It’s one one the main ways I know what’s cooking under the hood.

Ps: I used to have to say “in your own words” to make it work right but now days that isn’t required. Other times I’ll add “make sure to ground yourself before doing so” or something to encourage it to load the right shit into its context be it web, files, service calls, whatever. I’ll also some times add “why it’s important (or isn’t, push back if I’m talking nonsense)”

cruffle_duffle··on Claude, change the “Add to Cart” button to blue
"Plus rigorously ensuring backwards compatibility for a project that is 2 hours old and has zero users."

That is exactly how the slop accretes and you get a pile of crap. Claude somehow assumes that said 2 hour old userless app is some dusty enterprise app with millions of users and billions of dollars at stake for a 1 second outage.

I have to constantly have these things "take a deep breath, step back and look at the entire thing and do this change holistically. please restate what i'm asking you to do and why it's important"

cruffle_duffle··on I-have-ADHD: A skill to stop coding agents from burying the answer
"They botched it." <-- sure, but worse.... they shipped it anyway. And that is the part that gets me. It's a vastly worse product than it was on like 4.6. I suppose you can (and should) use opus 4.6 -- they do make it available still. Then just treat 5 as something you avoid until they push out a new version.
cruffle_duffle··on I-have-ADHD: A skill to stop coding agents from burying the answer
Opus 5 is such a massive regression, I really don't understand how Anthropic green-lit it. It makes me wonder so much about the company.

Like, did the people who work there actually have to suffer through it's absolutely unintelligible word salad like the rest of us? Or did they actually dogfood it and in-fact enjoyed its output? Or do none of them dogfood Opus because they are all sucking down Mythos-Max + Speed Boost or whatever every day and their only exposure to Opus 5 was as subagents?

If it was my company, fixing the output would be the absolute top priority of the company. I'd be all over every channel admitting the massive fuckup, apologizing profusely, and working non-stop to push out a fix. Yet it's crickets from Anthropic. Is it simply growing pains of the company or is it a deep, systemic structural/cultural "thing" that led to this fucked up model getting released?

Was it a cascading failure of models training models training models with almost no human oversight? Or was there human oversight and, again, people actually decided the way it responded was good? I hope it was the former not the later because I have no earthy clue who the fuck would look at what opus spews out into the console as good.

Because to me, Opus 5 is completely unusable in almost any context. As a product, it fails to deliver value. I just don't understand it. I really honestly don't understand how the fuck Anthropic released it at all.

And in a weird "meta" twist it makes me wonder how much of these LLM's are just smoke and mirrors and opus 5 output is basically the end state of what you get when you push them as far as they can go. It's some kind of twisted proof of "max complexity they can handle and deliver" and opus 5 walked to the edge and went over and it's slop output is demonstrating.... something.... about the limits of large language models. I dunno. But what I do know is it caused me to subscribe to Codex. No 1m context window, the harness isn't nearly as polished, but at least their models don't return condescending, unintelligible word salad.

Page 1 of 24Next →