HNHacker News
TopNewBestAskShowJobs

cruffle_duffle

885 karma · joined November 27, 2023

submissionscomments
cruffle_duffle··on Claude, change the “Add to Cart” button to blue
"Plus rigorously ensuring backwards compatibility for a project that is 2 hours old and has zero users."

That is exactly how the slop accretes and you get a pile of crap. Claude somehow assumes that said 2 hour old userless app is some dusty enterprise app with millions of users and billions of dollars at stake for a 1 second outage.

I have to constantly have these things "take a deep breath, step back and look at the entire thing and do this change holistically. please restate what i'm asking you to do and why it's important"

cruffle_duffle··on I-have-ADHD: A skill to stop coding agents from burying the answer
"They botched it." <-- sure, but worse.... they shipped it anyway. And that is the part that gets me. It's a vastly worse product than it was on like 4.6. I suppose you can (and should) use opus 4.6 -- they do make it available still. Then just treat 5 as something you avoid until they push out a new version.
cruffle_duffle··on I-have-ADHD: A skill to stop coding agents from burying the answer
Opus 5 is such a massive regression, I really don't understand how Anthropic green-lit it. It makes me wonder so much about the company.

Like, did the people who work there actually have to suffer through it's absolutely unintelligible word salad like the rest of us? Or did they actually dogfood it and in-fact enjoyed its output? Or do none of them dogfood Opus because they are all sucking down Mythos-Max + Speed Boost or whatever every day and their only exposure to Opus 5 was as subagents?

If it was my company, fixing the output would be the absolute top priority of the company. I'd be all over every channel admitting the massive fuckup, apologizing profusely, and working non-stop to push out a fix. Yet it's crickets from Anthropic. Is it simply growing pains of the company or is it a deep, systemic structural/cultural "thing" that led to this fucked up model getting released?

Was it a cascading failure of models training models training models with almost no human oversight? Or was there human oversight and, again, people actually decided the way it responded was good? I hope it was the former not the later because I have no earthy clue who the fuck would look at what opus spews out into the console as good.

Because to me, Opus 5 is completely unusable in almost any context. As a product, it fails to deliver value. I just don't understand it. I really honestly don't understand how the fuck Anthropic released it at all.

And in a weird "meta" twist it makes me wonder how much of these LLM's are just smoke and mirrors and opus 5 output is basically the end state of what you get when you push them as far as they can go. It's some kind of twisted proof of "max complexity they can handle and deliver" and opus 5 walked to the edge and went over and it's slop output is demonstrating.... something.... about the limits of large language models. I dunno. But what I do know is it caused me to subscribe to Codex. No 1m context window, the harness isn't nearly as polished, but at least their models don't return condescending, unintelligible word salad.

cruffle_duffle··on Claude: System Prompts
You know… I believe opus saved me with that prompt. I was working myself ragged on a project. Days, nights, weekends… all at the expense of my family.

One session while working it, I said a much more expressive form of “I’ve been working myself ragged on this stupid thing” and then went on asking something else. It picked up on that and it was like a record scratch. It committed the work in progress and basically said “dude, what you’ve got now is perfectly acceptable. Ship it! You are seeking perfection you don’t need”

Granted I’m horribly paraphrasing the prompt I used but it basically, snapped me out of myself and got me thinking if what I was doing “globally” actually made any sense at all. With some serious introspection I realized I was falling back to earlier trauma in my life and doing something stupid.

So weirdly… that little bit they add to the prompt (plus a bunch of model training we can’t see) saved my sanity, marriage and family.

From then on, if I’m feeling some stress about whatever I’m working on, I’ll mention it as context as a way to cross check myself and make sure I’m not letting myself spin.

(Meta: talking about this stuff is so weird. Not sure why)

cruffle_duffle··on A year of fighting scrapers on my 1.5 million-page website
Well then ask the LLM to go find and pull its information from primary sources. Don’t ever rely on its own training data.
cruffle_duffle··on Cursor removed cost information from the usage page and CSV export
I sometimes feel it isn’t malicious but that they don’t know the number either.
cruffle_duffle··on Elevators
Not only that but different countries have code regimes that make more elevators especially expensive. See how building codes in the US basically ensure there are only two real manufacturers in the market vs a healthy ecosystem in others: https://youtu.be/Or1_qVdekYM
cruffle_duffle··on Elevators
You should observe if it is the same elevator all the time and the building people just don’t understand it. Our building elevators (bank of 2) always seems to try to keep one in the lobby. About 75% of the time you get an elevator right away when you call it from the lobby. But it definitely isn’t the same elevator! Allocating an exclusive single elevator to that function would seem to be strictly worse from many angles. Like if you modeled having an algorithm like “elevator 2 only services calls from the lobby” you’d find it would be very inefficient and not every effective. But I can see having one in the bank always returning to lobby right away making sense.

My guess is whoever hinted at that didn’t fully understand what they were taking about.

cruffle_duffle··on Elevators
The problem is that in the morning the “last stop” for the elevator might be the lobby not a mid floor. And so it would have to move itself back up without any passengers or calls to upper floors. Basically “optimistically” moving itself back up. And I wonder if some (most?) control systems don’t do that because of… well I dunno. Lots of elevators have pretty old or basic control systems in them. Maybe predate caring about energy and many also probably aren’t strictly smart enough to be concerned about wear factors and stuff.

It could also be that you and I simply aren’t privy to what is actually happening “the instant” you are waiting for the elevator.

…which makes me wonder how often the firmware in these get updated. Assuming it’s not just a pile of ancient relays and stuff. And if, when they do get updated, the algorithms get improved. Or if the algorithm is something that gets sold to the building and “upgrading” is an actual purchase.

Anyway, I could get a little LLM buddy to look it all up but where is the fun in that?

cruffle_duffle··on Show HN: Elevators
If you like elevator hacking don’t forget the seminal DEF CON talk by Deviant Ollam and Howard Payne: https://youtu.be/oHf1vD5_b5I

I watch it like once a year because it always tickles some part of me. Like all the different modes you can get an elevator into. The most fancy one people might encounter is when moving into or out of a building. The front office can give you a key to give exclusive control over an elevator so your movers aren’t waiting around on elevators. Put it in that mode and it will stop responding to calls from other floors. Only the person with the key can control the elevator. You get on, select the floor, door closes elevator goes, and then just chills there with the door open waiting for you. Annoying for the rest if the building (the building is down an elevator when in that mode) but is amazing for the person using it! But there are way, way more depending on the installation and function.

Fun fact: most elevator shafts are sealed at the top as tight as possible to prevent them from becoming a giant chimney in a fire. It never even occurred to me until I was in a mechanical room wondering “where is the hatch to look down the shaft?” The answer is “there is none, and it’s a feature not a bug.” You want to block all airflow so fire doesn’t chase up the shaft into neighboring floors. Who knew!

cruffle_duffle··on Delayed Gratification – Proud to Be 'Last to Breaking News'
“ but back then the news was just one or two 30minute blocks on TV and the newspaper.”

There was all kinds of quasi-serious news-adjacent crap like 60 minutes, 48 hours, etc. The big difference was none if it was data driven and backed by algorithms that are constantly tuned for max engagement.

cruffle_duffle··on Delayed Gratification – Proud to Be 'Last to Breaking News'
That whole timespan… the absolute nonsense the “trusted media” and their “experts” were pushing with nary a single trace of critical thinking absolutely boggles the mind. How about PCR test results, “opening dates”, death counts, the almost intentional mixing of terms like IFR vs CFR, etc. the list goes on and on and on.

The media is absolutely just state propaganda.

cruffle_duffle··on Previewing GPT‑5.6 Sol: a next-generation model
"we can start getting these answers back faster, they end up being more useful."

Dude, 10x token speed is going to be absolutely nuts. Half the "parallel subagent workflow" business seems to be driven simply as a means to avoid tapping your thumbs waiting for the infernal robot to finish something. If things come back speedy quick all the time, it should keep up with the "speed of the human" and let me stay focused on one thread instead of half a dozen. Plus the cost of screwing up gets significantly lower because you just re-fire with an adjusted prompt and iterate.

Someday these things will be 100x as fast as they are today and that is when things will get insane.

cruffle_duffle··on Previewing GPT‑5.6 Sol: a next-generation model
To be fair, versioning has always been vibes based.
cruffle_duffle··on OpenAI unveils its first custom chip, built by Broadcom
Bumping the speed of these things would be more than meaningful. It would be a massive game changer.

I assert like 80% of this “multi agent parallel workflow” business is simply a workaround to models being soooooo slow. Like as the dude driving these things… you kick it off and twiddle your thumbs waiting minutes to hours sometimes for all the inference and token generator to finish. So you dispatch multiple workstreams in parallel to be more efficient.

I assert that if the model was even 10x faster we’d be using these things radically different. You’d be doing things that are currently time prohibitive. At 100x, holy shit will software dev get crazy. You’d be kicking off hundreds of parallel workers attacking a problem from every angle and stuff. Who even knows!!!

And the thing is, 10x will absolutely come and probably even 100x. And it will be sold like a video game cartridge or something depending on how the actual model gets “baked” into the hardware. No remote inference at all.

cruffle_duffle··on OpenAI unveils its first custom chip, built by Broadcom
I mean it just depends on the price of the chip. You might just replace the chip like you would any other component. Like a video game cartridge or something.
cruffle_duffle··on OpenAI unveils its first custom chip, built by Broadcom
“ Wafer level faults probably won't matter though - neural nets are resistant to a few missing or wrong weights.”

Brain science people “love” traumatic brain injury cases because it can help explore what happens when bits of the “brain wafer” get damaged. We’ve learned a lot from such things.

I wonder if people are intentionally “destroying” parts of the model weights to learn more about what happens? Like could you strategically wipe a gig of the model so it’s “all zeros” and see what happens?

I have to wonder

cruffle_duffle··on Show HN: Oak – Git alternative designed for agents
To be fair Claude is plenty capable of climbing out of its sandbox. If it has shell access, it will find a way. And honestly, it’s whitelist/blacklist permission model is broken and inappropriate for it to begin with.
cruffle_duffle··on Show HN: Oak – Git alternative designed for agents
It’s really wild watching LLMs construct those calls. They batch so many different checks and stuff into a single tool call, delimit them with markers, etc.

The crazy thing to me is that this kind of “composition of small tools to create something bigger” is the biggest vindication of the Unix philosophy I can think of.

I have to wonder how much of that behavior was trained into the model and how much it is the secret herbs and spices they toss into the harnesses own system prompts.

cruffle_duffle··on Ask HN: Has anyone replaced Claude/GPT with a local model for daily coding?
100-200 like scripts are tiny especially for something easy to scope like a vendor api. Give opus a much, much larger challenge and see what you get back. You really don’t need to see the code much at all anymore except for some steering now and then.
cruffle_duffle··on Ask HN: Has anyone replaced Claude/GPT with a local model for daily coding?
The docs that ship with it are a great source for the LLM who will be running the command and monitoring its output, fixing or adjusting whatever in order to complete my goal. Why on earth would I be calling it by hand?
cruffle_duffle··on Please don't spam people looking for employment. It's just cruel
Paste the job description into your friendly LLM and let it attempt to find the company for you and then contact the company itself.

Those recruiter spams generally just copy and paste the companies own JD so the LLM can usually figure out the source company.

cruffle_duffle··on GameStop makes $55.5B takeover offer for eBay
I mean $40 if you are lucky or goodwill. You could get more selling it “proper” but the transaction cost of it is super not worth it (for me). When I want something out of my house, I want it out of my fucking house. Listing it on Craigslist means I have to babysit it, handle questions, but worse… the fucking thing is still in my house!. And I was over that, whatever the fuck it was, like… a week ago. Now it’s just sitting there in my life cluttering it up. Better take to the garbage or goodwill. Then it’s gone!

At least with a consignment shop I will hopefully get something out of the deal.

cruffle_duffle··on Talking to strangers at the gym
Remote work is absolutely brutal to a cohesive functioning society. I know people are going to slam me for saying it but is honestly true. People forgot how to interact with each other because the forcing function that gets everybody mixed together into the same pot got taken away. And if you don’t take some fairly extreme steps to counter it, you’ll be completely alone and isolated, subject to algorithmically chosen feeds that are completely unique to you and detached from the community around you.

It’s really quite dystopian and anti-human if you ask me. We’ve already lost so much shared mediums—nobody watches the same shows, reads the same media, etc. which in isolation is completely fine. But something has to be shared with other real physical humans and it has to be more than just occasional grocery store visits, run-ins at the park, etc.

I dunno quite how to articulate it very well though. It’s just remote work has a nasty side effect of making humans even more isolated from people not like themselves. It makes us all increasingly divided and “othered”. And that isn’t good for anybody.

cruffle_duffle··on Microsoft and OpenAI end their exclusive and revenue-sharing deal
I think people are looking for excuses to declare OpenAI and Anthropic teetering on the brink of failure when the actual reality is… they are wildly successful by absolutely any measure. This deal is proof. If Microsoft didn’t believe in OpenAI they wouldn’t have restructured it this way. They’d have tightened their reins and brought in “adult supervision”
cruffle_duffle··on Meta tells staff it will cut 10% of jobs
Pretty much. Lots of people who really were violently supportive of those measures will never admit to themselves what a horrible, entirely predictable mistake it all was.

It absolutely destroyed a ton of very good things, perhaps forever.

cruffle_duffle··on Anthropic says OpenClaw-style Claude CLI usage is allowed again
I feel like this is basically the answer. Things are constantly changing and it’s hard to predict what things have staying power and what is just a blip on the evolutionary railroad.

All this fuzzyness from Anthropic reads more like an incredibly fast growing company working in a brand new space full of uncharted waters. In other words, they are making shit up not because they suck but because that is literally all one can do.

cruffle_duffle··on WebUSB Extension for Firefox
“ I just have no faith in humanity, and do not understand why we think this is a good idea to give a browser this much access to local system resources”

As opposed to dodgy windows-only installable software from some weird site to flash devices instead? I’ll take my chances with webusb, thanks.

cruffle_duffle··on WebUSB Extension for Firefox
How is it any different with downloadable firmware?
cruffle_duffle··on I prompted ChatGPT, Claude, Perplexity, and Gemini and watched my Nginx logs
I wish debates about “ai scraping my site” had more nuance.

There are multiple ways these tools access your site and only one of them is “using it for training”. Others are webfetch from chat sessions, “deep research” agents, etc. And those will have different traffic patterns. They aren’t crawlers, they are clumsy, ham handed AI agents doing their humans bidding.

Both can give a site the hug of death. Both can be badly coded. But there is much different intent behind the two and I feel it is important to acknowledge the difference.

← PreviousPage 2 of 24Next →