HNHacker News
TopNewBestAskShowJobs

BoiledCabbage

6,885 karma · joined March 26, 2017

submissionscomments
BoiledCabbage··on Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
> I wonder if this is related to Grok thinking it's a reincarnation of Hitler.

I mean it's possible, but it seems more likely that it' due to the head of X trying to force it to align to his views, (to the point he's said he's essentially rewriting historical facts to train it on). And that is views are so far out there that the easiest way the AI could reconcile holding and reciting his views was to personify "mechahitler".

BoiledCabbage··on Show HN: Vibe Kanban – Kanban board to manage your AI coding agents
How is the isolation layer of a shadow vcs different/better than just checking into a feature branch and not pushing?

Not saying there isn't value in the risk assessment part, but, I'm asking specifically about the change isolation.

BoiledCabbage··on Supreme Court's ruling practically wipes out free speech for sex writing online
Not having access to the internet during childhood 30-40 years ago is very different than not having access to the internet as a child now.

I think most people are taking your description of your experience being positive as a recommendation - one that is very unrealistic in modern times.

BoiledCabbage··on DOJ Statement of Interest on Suppression of Competition Through Deplatforming
> As a libertarian, I am thankful for this investigation.

I didn't get this line. You're saying "even though I am a libertarian I support this non-libertarian action"? Or is it something else?

Previously independent companies were free to choose what they display, now there are proposals for government mandates over their allowed actions.

I may be missing your point.

BoiledCabbage··on AI agent benchmarks are broken
> Tests that a "do nothing" AI can pass aren't intrinsically invalid but they should certainly be only a very small number of the tests. I'd go with low-single-digit percentage, not 38%. But I would say it should be above zero; we do want to test for the AI being excessively biased in the direction of "doing something", which is a valid failure state.

There is a simple improvement here: give the agent a "do nothing" button. That way it at least needs to understand the task well enough to know it should press the do nothing button.

Now a default agent that always presses it still shouldn't score 38%, but that's better than a NOP agent scoring 38%.

BoiledCabbage··on 'Improved' Grok Criticizes Democrats and Hollywood's 'Jewish Executives'
So after his failed first attempt at forcing Grok to reply with the repeatedly shown to be false South Africa "white genocide", he's has a new approach.

And to make his new approach work, he needs to literally re-write history (adding and deleting information) to get it to match his views. Because according to him any model trained on "uncorrected data" will never reach the crazy conclusions he wants it to?

This is one of the most absurd and insane things I think I've ever read.

BoiledCabbage··on The new skill in AI is not prompting, it's context engineering
> >This really does sound like Computer Science since it's very beginnings. > Except in actual computer science you can prove that your strategies, discovered by trial and error, are actually good.

Maybe it's true for computer science - but most people on here aren't doing computer science. They're doing software engineering. And it sure as heck isn't true for software engineering. If it were, I wouldn't be hearing arguments about programming languages for years, or static vs dynamic typing, or functional vs OOP...

So what you're arguing about AI isn't exactly anything new to software development.

BoiledCabbage··on What LLMs Know About Their Users
Not to complain, but that test would be more interesting if you ran it with an account.
BoiledCabbage··on BYU study: Why some people choose not to use artificial intelligence
> AI is built essentially on averages.

It is, but that also means if you prompt it correctly it will give you the answer of the average graduate student working on theoretical physics, or the average expert on the historical inter-cultural conflict of the country you are researching. Averages can be very powerful as well.

BoiledCabbage··on RFK Jr's new vaccine panel votes against preservative in flu shots in shock move
> Seeding an idea, to plant an idea does not jive here.

It does. If you want to see any reasonable policy passed instead of craziness, you need to convince some conspiracy groups to believe it, then it will be taken up by the administration and passed.

Plant the idea, they grow it and it gets passes. It's pretty clear that's what OP meant. 'Ceed', in this context makes no sense. In the discussed example policy is already coming from fringe groups, there is nothing to ceed to them. But to plant ideas with them that you want passed is how you'll get policy.

Note, not saying I agree or disagree with the original statement, just clarify what they meant by it.

BoiledCabbage··on Build and Host AI-Powered Apps with Claude – No Deployment Needed
JFC, these crypto people never stop! No matter how many times their tech has shown itself it be just about useless other than for illegal activities or for scamming people, they keep pushing it in every new tech space that exists.

I swear 100 years from now someone will be inventing faster then light travel, and there will be some tech scammers posting on HN on "how much better it would be on chain". Or how "the engine could be better if it used a decentralized crypto protocol."

The allure of being in the ground floor of a new scam just must be that great.

BoiledCabbage··on Using Microsoft's New CLI Text Editor on Ubuntu
The fact that everyone says the meme is dead, but in this small thread there are 5 different people posting how to exit, and none of them are the same says there is still pretty good substance behind that meme.
BoiledCabbage··on With only 8% built, Texas defunds state border wall program
Add long as certain groups keep falling for it we'll keep seeing them.
BoiledCabbage··on Is There a Half-Life for the Success Rates of AI Agents?
Oh man that's good - next step create a PR to push it up stream! Everyone can benefit from its fixes.
BoiledCabbage··on Large language models often know when they are being evaluated
If wasn't distracting for me (nor presumably for others). Maybe describing why you got so distracted by it?
BoiledCabbage··on How multiplication is defined in Peano arithmetic
Is there any notable difference between how it's presented in the post

> Thus, addition is a function P:NxN -> N such that for all numbers a, b,

> 1. P(a,0) = a

> 2. P(a,S(b)) = S(P(a,b))

And this alternate formulation?

1. P(a,0) = a

2. P(a,S(b)) = P(S(a),b)

Ie "decrease one from b and add it to a", instead of "decrease one from b and add it to the total".

BoiledCabbage··on CI/CD Observability with OpenTelemetry Step by Step Guide
I don't know the details but does a span have a beginning?

Is that beginning "logged" at a separate point in time from when the span end is logged?

> AIUI, there aren't really start or end messages,

Can you explain this sentence a bit more? How does it have a duration without a start and end?

BoiledCabbage··on Q-learning is not yet scalable
It does (naively I'll admit) seem like the problem is one more of approach more than algorithm.

Yes the model may not be able to tackle long horizon tasks from scratch, but learn some shorter horizon skills first then learn a longer horizon by leveraging groupings of those smaller skills. Chunking like we all do.

Nobody learns how to fly a commercial airplane plane cross country as a sequence of micro hand and arm movements. We learn to pick up a ball that way when young, but learning to fly or play a sport consists of a hierarchy of learned skills and plans.

BoiledCabbage··on AI agent startups at Y Combinator’s Spring ’25 Demo Day
> There needs to be a moat.

Why does there need to be a moat? You don't believe in markets and competition?

BoiledCabbage··on Agentic Coding Recommendations
Or even better, what if you could automate writing half or more of your unit test, and ensure they run not just out of band, but on ever build?

And even better rather than have them off in some far away location annotate the code itself so the tests will be updated with the code.

That's pretty impressive and someone would have to be short sighted to feel the false productivity of constantly manually implementing what a computer can automatically do for them.

Not to mention how much better if you work on any actual large scale systems with true cross team dependencies and not trivial code bases that get thrown away every few years where it almost doesn't matter how you write it.

BoiledCabbage··on Research suggests Big Bang may have taken place inside a black hole
> The second half is incorrect. Since the time coordinate becomes spacelike in turn you'll still have 3 spatial degrees of freedom. Dimensions can't just vanish if you believe that spacetime is a 4D Lorentzian manifold (as physicists do).

Can we say that one of the spatial dimensions (the radial dimension) and the time dimension combine into a single dimension? After crossing the event horizon aren't they 1:1 correlated?

BoiledCabbage··on I'm Wirecutter's water-quality expert. I don't filter my water
Can you share the evidence of this? Did you or someone have access to their affiliate revenue?

How do you know it shaped their decisions between products?

BoiledCabbage··on I'm Wirecutter's water-quality expert. I don't filter my water
Playing devils advocate for a min, your comment just a long way of saying "Don't get the very best thing - get the very best thing for you."

What tangible thing do you do differently from the advice this friend gave you? Or rather how did you shop before is your didn't look to see what utility it gave you in comparison to the cost?"

Let's say I'm in a situation where I need a bicycle for two months. I'm not going to buy the must expensive bike or there, I'm going to look around and buy a cheap bike that will be fair enough for two months. Are you saying before this advice you would research and buy the best bike out there?

BoiledCabbage··on The Illusion of Thinking: Strengths and limitations of reasoning models [pdf]
> There's nothing "omniscient" or "dim-witted" about these tools. They have no wit. They do not think or reason.

> All Large "Reasoning" Models do is generate data that they use as context to generate the final answer. I.e. they do real-time tuning based on synthetic data.

I always wonder when people make comments like this if they struggle with analogies. Or if it's a lack of desire to discuss concepts at different levels of abstraction.

Clearly an LLM is not "omniscient". It doesn't require a post to refute that, OP obviously doesn't mean that literally. It's an analogy describing two semi (fairly?) independent axes. One on breadth of knowledge, one on something more similar to intelligence and being able to "reason" from smaller components of knowledge. The opposite of which is dim witted.

So at one extreme you'd have something completely unable to generalize or synthesize new results. Only able to correctly respond if it identically matches prior things it has seen, but has seen and stored a ton. At the other extreme would be something that only knows a very smal set of general facts and concepts but is extremely good at reasoning from first principles on the fly. Both could "score" the same on an evaluation, but have very different projections for future growth.

It's a great analogy and way to think about the problem. And it me multiple paragraphs to write ehat OP expressed in two sentences via a great analogy.

LLMs are a blend of the two skills, apparently leaning more towards the former but not completely.

> What we do have are very good pattern matchers and probabilistic data generators

This an unhelpful description. And object is more than the sum of its parts. And higher levels behaviors emerge. This statement is factually correct and yet the equivalent of describing a computer as nothing more than a collection of gates and wires so shouldn't be discussed at a higher level of abstraction.

BoiledCabbage··on Smalltalk, Haskell and Lisp
> The above APL takes a similar liberty, but instead of a fiat declaring, we just a fiat declare that our data is sufficiently normalized.

Why would you make an assumption like this and try to compare code? If the sample code were simply subtraction then OP would've written subtraction. What use is demonstrating APL code that solves a different problem? If you're gonna do that, instead of removing part of the problem you might as well remove most of the problem and just say the solution is:

⌈/dest-source

And then write some sentence about saying how that's valid because you declare they are normalized even further.

> Here's my take on performScan, too: > Log⍪←onSourceTime+slewTime

How is this valid if you don't actually do any work to calculate onSourceTime or slewTime? You took the last line of performScan and implemented it. You do need to also implement the other lines in the function.

Maybe I'm just completely missing something here.

BoiledCabbage··on My AI skeptic friends are all nuts
> I think 384gb of ram is surprisingly reasonable tbh.

> 200-300$/month are already 7k in 3 years.

Except at current crazy rates of improvement, cloud based models will in reality likely be ~50x better, and you'll still have the same system.

BoiledCabbage··on The Evolution of Software Development: From Machine Code to AI Orchestration
> a kind of contradicting idea of AI being really intelligent agents, but also AI needs to be checked carefully. Either they're intelligent enough to not need humans or we are changing the definitions of intelligence.

Why is everyone so flummoxed by this? Your coworker is intelligent, but still needs code reviews.

Why is it than whenever people think of artificial intelligence the only options they see are dumb as a rock / pure parrot, or some omniscient god? There is no intelligence on earth that falls in either of those categories, but thats the only two options people can use to visualize AI and if it not one to them it must be the other.

Intelligence exists without being perfect gods. I feel like people have watched too much sci-fi.

BoiledCabbage··on A Break from Programming Languages
> Still, the history of mainstream programming languages is essentially a story of programmers vocally and emphatically rejecting what eventually proved to be some of the most incredibly successful innovations in the history of the field.

Concise summary of why programming language improvement is so difficult. It's a whole bunch of people yelling loudly while being wrong.

BoiledCabbage··on Square Theory
While it doesn't fully describe it, his category theory diagram reference seems relevant to me.

The stricter of the squares seem to be a homomorphism. But the "looser" ones which don't "preserve structure" after the transformation but "find a new structure" are some of the more interesting ones.

BoiledCabbage··on Attack of the Sadistic Zombies – Paul Krugman
And where are all the libertarians?

It's why I view contemporary libertarians as a farce. They were all up in arms, screaming at about having to wear a mask to keep others alive, but here we have people building an administrative state placing itself itself to be fully above the law able to impose any rule or action desired without consequence -- and not peep.

Maybe I'm just not understanding the distinction (and am open to being corrected) but the level of hypocrisy is just incredible to me.

← PreviousPage 7 of 34Next →