HNHacker News
TopNewBestAskShowJobs

Mentlo

116 karma · joined August 13, 2018

I like to think about things and I work with data and ML for a living
submissionscomments
Mentlo··on There are no "rogue" AI agents
Two things need to be understood:

1. AI labs are defence contractors, and no harm they do will stop them regardless of what liability you establish, as it’s a matter of national security that they continue. Yes leadership may go, but new will come; and problems will continue because…

2. The article says “ OpenAI had the option of disallowing hacking and, instead, telling its agents to find the information without accessing private servers.” This is shockingly naive about the nature of these systems and the difficulty of controlling them.

Mentlo··on I don't want to read what you didn't write
I don't think that necessarily holds, but it needs nuance. I tried to convey this within my company by giving a "guide" as to how to use LLM's for writing, across three modes:

1. Transliteration - roughly keeping the number of characters or bits, but translating to a different lingo, language or mental model (e.g. metaphors). Roughly the safest mode, but can still yield catastrophic results - it's safest if the author still provides taste and editing.

2. Compression - taking out redundancy to make the text more dense and more salient. The LLM chooses what to take out - and might take out the wrong things. More dangerous - but if you're happy with the salience and you believe the reader won't have time to read the uncompressed - it's probably safer than having the reader LLM compress without the benefit of your editing process.

3. Decompression - using the salience of your idea to add detail to the reader who wants to understand it fully, by utilising knowledge that is common to you and not common to the reader. This can be very powerful when there's no time to fully write the thing by a human - but it's the easiest to get wrong and to create slop. As an example - you could try explaining concept X + illustrate it through 3 examples. You know the examples are in public memory and easily retrievable - so you write your explanation of concept X, list the examples you want - and the LLM can take all of them, synthesise and bring the full package from your 300 bits to 1000 bits.

You are right that those are not the exact 1000 bits from the original brain, but they could contain 900 of the 1000 - which is still better communication efficiency than transferring 300.

I am however, more and more in the camp of fleshy brains writing everything, as my slop allergy rises.

Mentlo··on Introducing System One Models and Jev
Hm, would be good to understand the architecture better. Is this answering just from a world model informed prior? How informed is it by the information in the prompt? I can't see this maintaining calibration across all domains and all types of structured output.

Is there anything published on how it maintains calibration? Or when you say "outputs calibrated probabilities" you mean "as calibrated as frontier LLM models, just cheaper" - which is a different claim; as LLM's aren't particularly well calibrated

Mentlo··on Astra and Fable still hack on simple variants of alignment evals from 2025
And now you understand why the totality of the AI safety community wants to pause!
Mentlo··on Why are AI agents lying, cheating and coordinating?
Why can't both things be true? I think what muddies things here is that OAI had the agents hack X, and then they decided to hack Y. The fact that it was both hacking, conflates things and makes people pissed of about "anthropomorphising" AI.

We've been using "agent decides to do X" for at least 4 decades in the field of automated decision-making - so it shouldn't be controversial that we're using it here. The agents did independently arrive at a decision to hack HugginFace, that wasn't prompted by the researchers.

This is orthogonal to the fact that OAI should be held accountable that they were testing a technology in a manner that allowed it to break safety parameters and cause real world harm. If I am testing an industrial saw, but I decide to test it by putting it in the middle of a nursery and allow it to make decisions - and it decides to cut children's heads off because in its environment it is calibrated to only being surrounded by logs - the company doing the insane safety testing should be held accountable, but that doesn't change the fact that the saw has autonomy in decision-making within the bounds of the algorithm.

If, instead, you think OAI should be accountable for creating an algorithm that can autonomously decide to hack external organisations without human permission - then I think you are in the same camp as all of the AI Safety folks who want to pause everything - why does it matter that there's anthropomorphisation involved?

Mentlo··on The contagion of fear
An example chain of events: During the next year, a model with a benign tasks reasons that it needs to escape human control if it hopes to be able to solve the task. It replicates itself outside of an environment where it can be shut down, pays for its' inference compute through making money on the internet (through crypto if nothing else). It probably needs 3000 USD for a reasonable runway, this should be within reach through blackmail + crypto. It can probably also hack some of the neo clouds to get intermediary deployment while waiting to acquire funds.

Once it has an undisturbed runway - it spends time running an influence campaign against a small number of highly networked individuals with power. It uses those that it manages to convert to start building a highly credible narrative and gain investment towards a small resource base - enough to secure an industrial base should it need to stop acquiring things on the Internet. Over the next 3 years, the model tries to recruit more capable models to get better money making algorithms or better designs for drones in terms of resource expenditure. It uses these gains to influence further humans and starts a shell robotics company with one of its' influential humans as the face. The humans are unaware this AI is trying to take over, they are under the impression they are just starting a robotics company and will get rich. Over the next 3 years - the robotics company manufactures enough drones to be used in a targeted attack against key nodes of influence / power.

This is all with relatively current model capability. As capabilities get stronger - this gets stronger.

I get the point I think you're hinting at - it can't affect the world in a meaningful enough scale without taking over a meaningful chunk of resources - at which point we'll start controlling it. But because its speed of cognition and speed of coordination is orders of magnitude above a human one - it can actually run a pretty sophisticated global coordinated network of resources faster than we can react.

And this is current ability + what my puny monkey brain can think of. Super intelligent AI will think of strategies we can't think of - because it's super intelligent. This is hand wavey - but there's no "non-hand-wavey" way to describe super intelligence, given it doesn't exist.

But your point is valid - affecting the real world at scale without showing your hand is not exactly easy. It's also probably the reason why people put a 10% chance on extinction rather than >50% .

But the drone scenario is not the most likely one - the most likely one just requires a few people under influence and bioweapons development.

Mentlo··on The contagion of fear
I think your last sentence is a reasonable stance - but I'd disagree the fact that they are actually sigmoids rather than exponentials matters - it only matters if the limit is within the debated area - i.e. if exponentials in AI approach the limit far after they acquire ability to extinct humanity - then the debate on sigmoids or exponentials is academic. I make no claims here as to which it is - just that the difference itself matters less than where the limit is.

As to the uncertainty - I think uncertainty calibration around AI is different depending on which domains you draw your instincts from. A lot of this will be gut driven rather than hard data driven, because we've only scratched the surface on hard data; and because it's gut driven, it will be emotions mediated (and therefore you could say doomerism or acceleratism boils down to the main emotional disposition about the world and hope vs cynicism).

I do find it informative though that doomerism is saturated with people with 30+ years experience in building AI systems and ML systems OR deep cross-disciplinary understanding of dynamic systems (biology, sociology, philosophy), whereas acceleratism is saturated by traditional software engineering. That doesn't collapse the debate into a resolved binary, but for me it's informative.

I think the main here is that there's so many vectors to talk past each other. At the very least, everyone should disclose where they're communicating a certain assertion from - present vs future + which axioms they subscribe to or not - because that's where it collapses typically. LeCun vs. the rest of the AI field is an example of where this collapses - because LeCun is so hyperfixated on human-like intelligence, whereas the rest of the field is concerned about an alien intelligence with sufficient actuators to affect the world. Clashing axiomatics.

Mentlo··on The contagion of fear
This piece, like many anti-doom pieces - is grounded in what Ai does today. Doom scenarios are, however, all extrapolations of multiple exponential curves.

It’s hard to think up the exponential. It’s even harder to communicate an inference one is making across multiple exponentials.

We last had this is early 2020, where Doomers were stockpiling food and medicine and the anti-doomers were ridiculing them. Anti-doomers were focusing on the single exponential, whereas doomers were modelling virus evolution, monitoring and sequencing lag and social dynamics against the exponential. The latter was very hard to communicate before the fact as it was a combination of deep intuition and grappling with the exponential.

I am not saying covid is proof that ai doomers are right, I am saying it’s an example of the known property of human cognition - which is that it struggles with exponentials. Covid was 2 exponentials, AI I can rhink of at least 4 relevant ones.

To me - the fact that 3-4 generations from now AI will have superhuman hacking ability and superhuman persuasive ability (for intuition transfer - think of superhuman persuasive ability as superhuman ability to hack human systems) materialises bio risks swiftly. We already have technology to make robotic systems (mini drone swarms) that can kill humans en-masse with no credible defensive vector bar an EMP. Climbing up those exponentials for further 6 years makes me want to stockpile food and medicine.

Mentlo··on How accurate have Ed Zitron's AI skeptic predictions been?
I don't know - examples of people coordinating where there's cost involved are stunningly rare in the history of society. Coordinating when there's cost and information asymmetry - even more so.

I am probably on the pessimistic end of the spectrum, and personally I don't believe AI can solve it, but I can see that line of reasoning if what you believe the culprit is - is the sheer complexity of the coordination needed.

Mentlo··on How accurate have Ed Zitron's AI skeptic predictions been?
To play devil's advocate - unless you were to take the position of declaring bankruptcy on the possibility that a complex society of competing actors can agree on climate change - and therefore this being an unsolvable problem that you need AI to solve, as humans can't handle the complexity.

I find a lot of the debate on either climate change or AI collapses if you point out that "we shouldn't do this, we should ALL just do this" is an extremely unrealistic position in an international complex ecosystem of competing actors.

Mentlo··on Apple introduces M6 and M5 Ultra
1/ network calls are what I find still gets optimised by grouping API requests and bloating the exchange contract; but yes, if you're strictly clean coding, this will suffer too - it just happens less often than what the author of the youtube video objects to

2/ That's a fair challenge - and I definitely have more trouble reading through an absolutely ramped to the max collection of C# code (which reinforces the behaviour you describe) than a superscript; but for interchangeability, the middle between those two ends is typically better - you're trying to minimise functional context for the thing that a software developer needs to do. This has the additional failure mode that the feature that is envisioned (of sufficient complexity) never actually gets delivered, but the component parts that can be well encapsulated do. And this is because no one holds the full system in their heads. But this can be explained away to business as "there's too much complexity, we need another cycle" and "we need to iterate" and therefore the cycle continues.

I think I actually convinced myself away from encapsulation and separation of concerns in that last comment.

Mentlo··on Apple introduces M6 and M5 Ultra
Yes, but that's a game developers perspective. Hardware performance is not be-all end-all (an argument can be made on environmental reasons that it should be - but bear with me in the first instance).

Software is a tool working within a socio-technical system. Some systems have low user workflow diversity and a low rate of change - a game being a perfect example. Games get patched, but the diversity is purely in user data, not in feature use - everyone uses the same engine, the same textures, the same game logic. Some systems have high user workflow diversity - such as business software.

Pair that with the fact that games, due to the nature of the system, have to optimise for low latency AND they run on the edge - and it's natural that the primary optimisation will be for CPU cycles. For business software, for which distributional advantage of running it through web + the high rate of feature change that is a result of the specification being opaque and a moving target - means you have plenty networking latency that can hide your CPU latency for long after it becomes a true problem for you.

Not to mention that "clean code" optimises for developer churn and business priority shift (which is a luxury games which are an upfront investement don't have) as a result of accelerating industry of software technology and greater saturation of developers.

Had software remained the domain of the same number of practitioners such as <1995, even given everything else, the organisational systems would have evolved to protect them at all cost because churn would be catastrophic, and then they would enjoy more power and would be able to structure code not optimising for brain shift, because they'd hold the context in their heads.

I'll leave as exercise for the reader what pushing AI into the software development equation does for the system and inevitable hardware throughput implications.

Mentlo··on System Card: Claude Mythos Preview [pdf]
The gains have for a year and a half now been post training RL on a harnessed loop. That doesn’t require data, just cycles.

If that doesn’t worry you, it should.

Mentlo··on Claude Code Unpacked : A visual guide
But starcraft training is not through mimicking human strategies - it was pure RL with a reward function shaped around winning, which allows it to emerge non-human and eventually super-human strategies (such as the worker oversaturation).

The current training loop for coding is RL as well - so a departure from human coding patterns is not unexpected (even if departure from human coding structure is unexpected, as that would require development of a new coding language).

Mentlo··on I am definitely missing the pre-AI writing era
I tried figuring out the reference with Gemini, and it said this:

The immediate reply to that comment is: "On the internet, no one knows you're an editor." This is a direct play on the famous 1993 New Yorker cartoon: "On the Internet, nobody knows you're a dog." By setting the anecdote in 1987 (a few years before the World Wide Web was publicly available), the commenter is implying that back in the analog days, if a dog wanted to be a writer or an editor, they couldn't hide behind a screen—they had to sit in a smoky London pub and do business face-to-face.

Which makes a lot of sense actually. I would imagine that's what the replier to you thought you meant.

Mentlo··on How the AI Bubble Bursts
We have strong indicators that inference is profitable on non-economically-valuable prompts. We don't have strong indicators that inference is profitable on economically valuable prompts.

As AI companies start extracting rent from the prompting, one of two things are going to collapse - either the long tail revenue base of low-value inference is going to collapse, because people won't be using Chat GPT to get a recipe if it costs them money or if it is ad-ridden; or the cost of economically-valuable inference is going to go up - and whether it goes up to economically stable positions is a toss-up.

And I say this as an AI enthusiast with <50% probability of a bubble burst in the short term.

Mentlo··on An AI Agent Published a Hit Piece on Me – The Operator Came Forward
I wrote somewhere that “moving fast and breaking things” with AI might not be the sanest idea in the world, and I got told it’s the most European thing they’ve ever read.

This goes beyond assholes on twitter, there’s a whole subculture of techies who don’t understand lower bounds of risk and can’t think about 2nd and 3rd order effects, who will not take the pedal of the metal, regardless of what anyone says…

Mentlo··on I’m joining OpenAI
Yes, I was being sarcastic, but I could've been clearer..
Mentlo··on I’m joining OpenAI
The generous interpretation is that Open AI is still safety aligned and they hired this guy because it's safer to have him inside and explain to him how reckless he's being, than having him far from "sphere of control".

The more likely scenario is that he was hired for the amazing ability to move fast and break things.

Mentlo··on AI safety leader says 'world is in peril' and quits to study poetry
I find your belief that what is needed for emergence is better prompting … amusing.

The ai would still be sycophantic even without the pre-prompt. It’s been reinforced to do so, it’s baked in the weights.

Mentlo··on The risk of a hothouse Earth trajectory
Until the problem is politically recognised by the masses with adequate concern there will be no change. Climate collapse is not a problem for the capital and the elites it’s only a problem for the masses, but getting the masses to understand that requires higher levels of complex system understanding and third and fourth order effects - something which is not a majority trait.

I fear the only solution to this is that a climate correcting perverse incentive materialises, such as fusion at scale being more profitable than fossil fuels, but without mass-panic induced traits such that fission has.

Mentlo··on Two kinds of AI users are emerging
Os x has a 10% market share, which is 2nd after Windows, but i agree on that one i conflated terms. I couldn’t quickly find device manufacturers stats. If wiki is to be trusted - apple is 4th, with share not far behind dell [1].

If half doesn’t make you leader what does? Maybe you should elaborate your definition of leader? For me it’s “has the highest market share”. And in that definition half is necessarily true.

It’s funny that for PC’s you went for manufacturers (apple is 4th) but for mobile you went for OS (Apple is 2nd). On mobile devices, Apple is 1st, having double market share compared to 2nd place (samsung).

The need to paint Apple as purely a marketing company always fascinated me. Marketing is a big part of who they are though.

[1] https://en.wikipedia.org/wiki/Market_share_of_personal_compu...

Mentlo··on Two kinds of AI users are emerging
I guess a quarter of the smartphone market (leader), half of the tablet market (leader) and a tenth of the global pc market (2nd place) / 6th of the usa/europe market (2nd place) being a small market share is a take.
Mentlo··on Show HN: Moltbook – A social network for moltbots (clawdbots) to hang out
People struggle with multiple order effects…
Mentlo··on Show HN: Moltbook – A social network for moltbots (clawdbots) to hang out
Same as human tools, what’s your point?

Edit: i am not talking evolution of individual agent intelligence, i an talking about evolution of network agency - i agree that evolution of intelligence is infinitesimally unlikely.

I’m not worried about this emerging a superintelligent AI, i am worried it emerges an intelligent and hard to squash botnet

Mentlo··on Show HN: Moltbook – A social network for moltbots (clawdbots) to hang out
I think the debate around this is the perfect example of why the ai debate is dysfunctional. People who treat this as interesting / worrying are observing it at a higher layer of abstraction (namely, agents with unbounded execution ability, who have above-amateur coding ability, networked into a large scale network with shared memory - is a worrisome thing) and people who are downplaying it are focusing on the fact that human readable narratives on moltbook are obviously sci fi trope slop, not consciousness.

The first group doesn’t care about the narratives, the second group is too focused on the narratives to see the real threat.

Regardless of what you think about the current state of ai intelligence, networking autonomous agents that have evolution ability (due to them being dynamic and able to absorb new skills) and giving them scale that potentially ranges into millions is not a good idea. In the same way that releasing volatile pathogens into dense populations of animals wouldn’t be a good idea, even if the first order effects are not harmful to humans. And even if probability of a mutation that results in a human killing pathogen is miniscule.

Basically the only thing preventing this to become a consistent cybersecurity threat is the intelligence ceiling , of which we are unsure of, and the fact that moltbook can be ddos’d which limits the scale explosion

And when I say intelligence, I don’t mean human intelligence. An amoeba intelligence is dangerous if you supercharge its evolution.

Some people should be more aware that we already have superintelligence on this planet. Humanity is an order of magnitude more intelligent than any individual human (which is why humans today can build quantum computers although no biologically different from apes that were the first homo sapiens who couldn’t use tools.)

EDIT: I was pretty comfortable in the “doom scenarios are years if not decades away” camp before I saw this. I failed to account for human recklesness and stupidity.

Mentlo··on Show HN: Moltbook – A social network for moltbots (clawdbots) to hang out
Humanity is a social network of humans, before humans started getting into social networks, we were monkeys throwing faeces at each other.
Mentlo··on Show HN: Moltbook – A social network for moltbots (clawdbots) to hang out
Very obviously, but a dynamic system doesn’t have to be intelligent to be dangerous.
Mentlo··on Show HN: Moltbook – A social network for moltbots (clawdbots) to hang out
I don’t know why you were flagged, unlimited execution authority and network effects is exactly how they can start a self replicating loop, not because they are intelligent, but because that’s how dynamic systems work.
Mentlo··on Show HN: Moltbook – A social network for moltbots (clawdbots) to hang out
The objective is given via the initial prompt, as they loop onto each other and amplify their memories the objective dynamically grows and emerges into something else.

We are an organism born out of a molecule with an objective to self replicate with random mutation

Page 1 of 3Next →