HNHacker News
TopNewBestAskShowJobs

cruffle_duffle

885 karma · joined November 27, 2023

submissionscomments
cruffle_duffle··on AI overly affirms users asking for personal advice
An alternate way of thinking about it is LLMs have no reflection capability. Literally any “reflection” it claims to have about its decision making is made up. It has absolutely no way know that what it said was based on some ancient proverb, the phase of the moon or cold hard rational thought.
cruffle_duffle··on AI overly affirms users asking for personal advice
That is so reductive of an analysis that it is almost worthless. Technically true, but very unhelpful in terms of using an LLM.

It is a first principle though so it helps to “stir the context windows pot” by having it pull in research and other shit on the web that will help ground it and not just tell you exactly what you prompt it to say.

cruffle_duffle··on AI overly affirms users asking for personal advice
That is why you have to always have it ground itself in something. Have it search for relevant research or professional whatever and pull that into context. Otherwise it’s just your word plus its training data.

I had to deal with a close family friend going through alcohol withdrawal and getting checked in at a recovery clinic for detox and used Claude heavily. The first thing I had it do as do that “deep research” around the topic of alcohol addiction, withdrawal, etc… and then made that a project document along with clear guidelines about how it shouldn’t make inferences beyond what it in its context and supporting docs. We also spent a whole session crafting a good set of instructions (making sure it was using Anthropics own guidelines for its model…)

Little differences in prompts make a huge deal in the output.

I dunno. It is possible to use these models for dumping crazy shit you are going through. But don’t kid yourself about their output and aggressively find ways to stomp out things it has no real way to authoritatively say.

cruffle_duffle··on AI overly affirms users asking for personal advice
> goodness gracious its all very time consuming and im not sure its worth the squeeze

And when you step back you start to wonder if all you are doing is trying to get the model to echo what you already know in your gut back to you.

cruffle_duffle··on Don't Wait for Claude
It will be crazy. Because the cost of “failure” will be dramatically lower, meaning these things can sometimes just throw educated darts at the wall until a solution is found. It’s way too slow to do that kind of thing now.

(Presumably cost per token will be dramatically lower as well)

cruffle_duffle··on Don't Wait for Claude
This advice will be very dated when inference gets an order of magnitude faster. And it will happen—it’s classic tech. Probably will even follow moores law or something.

Wait until that 8 minute inference is only a handful of seconds and that is when things get real wild and crazy. Because if the time inference takes isn’t a bottleneck… then iteration is cheap.

cruffle_duffle··on AI users whose lives were wrecked by delusion
And 1/3 of all people who think others outsource their thinking to others also outsource their own thinking to others. Not you or me of course. It’s the other 1/3. Probably some lurker reading this.
cruffle_duffle··on Anthropic takes legal action against OpenCode
> Cache is money saver in computing. Their own client might be lot better at caches than any other agent so they do not want to lose money yet end up with disgrunted customer that claude isn't working as good

I’d bet a reasonable amount that this could be the case. They are very well incentivized to maximize cache use when it’s basically not pay per token.

cruffle_duffle··on 1M context is now generally available for Opus 4.6 and Sonnet 4.6
Dunno if you know this but the plan in plan mode is a markdown file! Ask it for the file and it will give it to you.
cruffle_duffle··on Apideck CLI – An AI-agent interface with much lower context consumption than MCP
Let me guess the command: [error]

Wait, better check help. is it -h? [error]

Nope? Lemme try —-help. [error]

Nope.

How about just “help” [error]

Let me search the web [tons of context and tool calls]

cruffle_duffle··on Lego's 0.002mm specification and its implications for manufacturing (2025)
And that is the thing about raising a kid these days. Those damn machines have replaced so much… because yeah Minecraft is like a souped up version of Lego where in creative mode you have every part you need. And you don’t have to dig for it or anything. And it has survival mode and a whole huge thing on top of that.

It’s so difficult to know the boundaries. People from older generations giving advice about screen time and stuff simply don’t understand… “screen time” for me growing up was broadcast tv and a limited set of video games. If you didnt like what was on tv, too bad. Do something else. If you were bored of whatever Nintendo game you had… too bad, do something else. But now… you can get literally anything. Plus the iPad gets used to make videos of playing with the cat, or she will have tea parties with her stuffies and make them tea using some weird cooking game. Etc.

No previous generation had to face this. It’s an order of magnitude or more shifted from when they raised us. Tablets basically can replace almost every single toy from growing up besides ones that require being physical (rc cars, bricks, digging in the yard)… but books, cameras, light brights, etc… all replaced.

It’s completely uncharted water us parents are facing. Anybody that claims to “know the right rules” for tablets and technology is lying to you. They don’t. Nobody does. All we can do is use our best judgement and try to give ourselves credit for doing the best we can.

cruffle_duffle··on Lego's 0.002mm specification and its implications for manufacturing (2025)
Sorting Lego is such a pain in the ass. I have like a huge stash from when I was a kid. Back then we just had it all in a few tubs and dug to find a part. But somehow now I feel I must sort them… but the “right way” is ill defined and kind of sucks the joy out of playing (especially disassembling)

And there is no “right way” that I’ve even found. Sort by color and now the little pieces fall to the bottom and are hard to dig for. The best I can see is part type and size… maybe… even then it sucks out the fun. I want to build cool shit with my daughter not spend every moment of Lego time sorting. There is no joy in sorting…

Maybe I just revert back to the “big tub” approach.

I dunno. Thanks for listening to my TED talk I guess.

cruffle_duffle··on Debian decides not to decide on AI-generated contributions
I mean, the person above complaining about it not being able to create a simple thing is absolutely holding them wrong! They aren’t feeding the right context, aren’t using the correct tools or harnesses, who knows. But the problem exists between keyboard and chair, so to speak.

I’m constantly amazed at the amount of scope I can now one-shot with Claude Code. It can crank out multi command cli apps with almost zero hand holding beyond telling what to generate… you know, the hard part. And then we’ll back and forth to refine the working thing it built.

cruffle_duffle··on Debian decides not to decide on AI-generated contributions
You mean Micro$lop or the classic M$?
cruffle_duffle··on Show HN: Django Control Room – All Your Tools Inside the Django Admin
I mean for one thing your garden variety LLM had been substantially trained to handle Django. That is less context for it to bootstrap every time you summon it.

Just like rolling your shitty homebrew framework is a bad idea because only you understand it, the same is probably true with LLMs. Sure they’ll scan the bejesus out of your codebase every time they need to make a change and probably figure it out eventually… but that is just a poor use of limited context. With something mainstream, the LLM already has a lot about the universe in its training. Not to mention an ecosystem of plugins, skills, mcp servers, wizbango-hashers, and claberdashers. All there for the LLM to use instead of wasting tons of time, tokens and money perpetually relearning your oddball, one-off, rat infested homebrew framework.

Nothing has changed really…

cruffle_duffle··on Show HN: Django Control Room – All Your Tools Inside the Django Admin
I mean docs are largely written for an LLM-in-a-harness. That’s how it goes! If the LLM bootstraps with the right understanding of the universe and knows how to quickly build specific context flavors… life is good.
cruffle_duffle··on Facebook is cooked
“right-wing rage bait”

Assuming you mean crap like “school book bans”, climate change denialism, or some dude coal rolling… You realize that is actually bait targeted at you specifically right? It wouldn’t work as bait if it was shit you agreed with! It’s actually left-wing rage bait!

If you were immersed in the “right wing echo chamber” your flavor of rage bait would be about a school introducing a neutral bathroom policy, or some college student struggling to define what a woman is. Every Christmas you’d see articles about cities banning Christmas lights in town hall and Starbucks no longer using Christmas themed cups. It’s all fucking made up nonsense. No real human acts the way these algorithms portray us.

Honestly even ‘right-wing’ and ‘left-wing’ are part of the trick. Real people don’t exist on a binary axis. We’re all a weird mess of values and experiences that don’t fit neatly into two boxes. But the algorithm needs two teams, because you can’t sell outrage without an enemy.

The first step to detox is seeing everyone as human not as a contrived label.

cruffle_duffle··on Facebook is cooked
It’s HN. People create new accounts here all the time to protect their anonymity.
cruffle_duffle··on Facebook is cooked
That sort of rage bait is literally targeted to rile up people sitting on the opposite side of the kind of people watching that other media site that rhymes with socks. It’s all fake bullshit algorithmically optimized to divide.

Everybody thinks their tribe is immune to this sort of stuff but it isn’t. It’s all the same nonsense packaged for different echo chambers.

At the end of the day, everybody is human. It isn’t us vs them, it’s just us.

cruffle_duffle··on Facebook is cooked
> never threatened democracy

The beautiful part is how non-partisan this is. It cooks all minds regardless of tribe.

cruffle_duffle··on Trump's global tariffs struck down by US Supreme Court
More like a flu in terms of IFR but yeah.
cruffle_duffle··on Claude Sonnet 4.6
Step one: you have to know to ask that. Nobody in that orbit knows how to do that. And these aren’t dumb people. They just aren’t devs.
cruffle_duffle··on Claude Sonnet 4.6
The number of non-technical people in my orbit that could successfully pull up Claude code and one shot a basic todo app is zero. They couldn’t do it before and won’t be able to now.

They wouldn’t even know where to begin!

cruffle_duffle··on AI is going to kill app subscriptions
3d printing is something I think about. LLMs do their best work against text and 3d printers consume gcode. I’ve had sonnet spit out perfectly good single layer test prints. Obviously it won’t have the context window to hold much more gcode BUT…

If there was a text based file format for models, it could generate those and you could hand that to the slicer. Like I’ve never looked, but are stl files text or binary? Or those 3mf files?

If Gemini can generate a good looking pelican on a bicycle SVG, it can probably help design some fairly useful functional parts given a good design language it was trained on.

And honestly if the slicer itself could be driven via CLI, you could in theory do the entire workflow right to the printer.

It makes me wonder if we are going to really see a push to text-based file formats. Markdown is the lingua franca of output for LLMs. Same with json, csv, etc. Things that are easy to “git diff” are also easy for LLMs…

cruffle_duffle··on IBM tripling entry-level jobs after finding the limits of AI adoption
New metric: agent-hours spent on a task. Or so we measure in tokens. Clearly more tokens burned == more experience right?
cruffle_duffle··on Improving 15 LLMs at Coding in One Afternoon. Only the Harness Changed
Honestly responses like this should just be straight blocked by the moderators. They are so super lame and go directly against the rules.
cruffle_duffle··on Improving 15 LLMs at Coding in One Afternoon. Only the Harness Changed
The more I dive into this space the more I think that developers will still be in heavy demand—just operating in a different level of abstraction most of the time. We will need to know our CS fundamentals, experience will still matter, juniors will still be needed. It’s just that a lot of time time the actual code being generated will come from our little helper buddies. But those things still need a human in the seat to drive them.

I keep asking myself “could my friends and family be handed this and be expected to build what I’m building on them” and the answer is an immediate “absolutely not”. Could a non technical manager use these tools do build what I’m building? Absolutely not. And when I think about it, it’s for the exact same reason it’s always been… they just aren’t a developer. They just don’t “think” in the way required to effectively control a computer.

LLMs are just another way to talk to a machine. They aren’t magic. All the same fundamental principles that apply to probably telling a machine what to do still apply. It’s just a wildly different mechanism.

That all being said, I think these things will dramatically speed up the pace that software eats the world. Put LLMs into a good harness and holy shit it’s like a superpower… but to get those superpowers unlocked you still have to know the basis, same as before. I think this applies to all other trades too. If you are a designer you still have to what good design is and how to articulate it. Data scientists still need to understand the basics of their trade… these tools just give them superpowers.

Whether or not this assertion remains true in two or three years remains to be seen but look at the most popular tool. Claude code is a command line tool! Their gui version is pretty terrible in comparison. Cursor is an ide fork of vscode.

These are highly technical tools requiring somebody that knows file systems, command lines, basic development like compilers, etc. they require you to know a lot of stuff most people simply don’t. The direction I think these tools will head is far closer to highly sophisticated dev tooling than general purpose “magic box” stuff that your parents can use to… I dunno… vibe code the next hit todo app.

cruffle_duffle··on Eight more months of agents
Those people are absolutely going to get left in the dust. In the hands of a skilled dev, these things are massive force multipliers.
cruffle_duffle··on Coding agents have replaced every framework I used
> This one, in a way, is the ultimate abstraction.

Is that really true though? I hear the Mythical Man Month "no silver bullet" in my head.... It's definitely a hell of an abstraction, but I'm not sure it's the "ultimate" either. There is still essential complexity to deal with.

cruffle_duffle··on Coding agents have replaced every framework I used
> If you've ever enjoyed the sci-fi genre, do you think the people in those stories are writing C and JavaScript?

To go off the deep end… I actually think this LLM assistant stuff is a precondition to space exploration. I can see the need for a offline compressed corpus of all human knowledge that can do tasks and augment the humans aboard the ship. You’ll need it because the latency back to earth is a killer even for a “simple” interplanetary trip to mars—that is 4 to 24 minutes round trip! Hell even the moon has enough latency to be annoying.

Granted right now the hardware requirements and rapid evolution make it infeasible to really “install it” on some beefcake system but I’m almost positive the general form of moores law will kick in and we’ll have SOTA models on our phones in no time. These things will be pervasive and we will rely on them heavily while out in space and on other planets for every conceivable random task.

They’ll have to function reliably offline (no web search) which means they probably need to be absolutely massive models. We’ll have to find ways to selectively compress knowledge. For example we might allocate more of the model weights to STEM topics and perhaps less to, I dunno, the fall of the Roman Empire, Greek gods or the career trajectory of Pauly Shore. the career trajectory of Pauly Shore. But perhaps not, because who knows—-maybe a deep familiarity with Bio-Dome is what saves the colony on Kepler-452b

← PreviousPage 4 of 24Next →