MIT Technology Review has confirmed that posts on Moltbook were fake
technologyreview.com
technologyreview.com
Putting aside how incredibly easy it is to set up an agent, or several, to create impressive looking discussion there, simply by putting the right story hooks in their prompts. The whole thing is a security nightmare.
People are setting agents up, giving them access to secrets, payment details, keys to the kingdom. Then they hook them to the internet, plugging in services and tools, with no vetting or accountability. And since that is not enough, now the put them in roleplaying sandbox, because that's what this is, and let them run wild.
Prompt injections are hilariously simple. I'd say the most difficult part is to find a target that can actually deliver some value. Moltbook largely solved this problem, because these agents are relatively likely to have access to valuable things, and now you can hit many of them, at the same time.
I won't even go into how wasteful this whole, social media for agents, thing is.
In general, bots writing each other on mock reddit, isn't something the loose sleep over. The moment agents start sharing their embeddings, not just generated tokens online, that's the point when we should consider worrying.
But, I do have a distinct feeling that his enthusiasm can overwhelm his critical faculties. Still, that isn't exactly rare in our circles.
Everything Karphathy said, until his recent missteps, was received as gospel, both in the AI community and outside.
This influencer status is highly valuable, and I would not be surprised if he was approached to gently skew his discourse towards more optimism, a win-win situation ^^
I'll confess I try to ignore industry chatter to a fair degree.
Its the same reason why a pure technologist can fail spectacularly at developing products that deliver experiences that people want.
Being picked by Elon perhaps amplified that too.
I don't have time to dig up the citation that someone pointed me towards, but it's out there and can be found. Which is a bummer, because I've learned a lot from his videos and writing and have a lot of respect for his work.
Intelligence (which psychologists define as the g factor [1]; this concept is very well-researched) does not make you an expert on any given topic. It just, for example, typically enables you to learn new topics faster, and lets you see connections between topics.
If Karpathy did not spend a serious effort of learning to get a good understanding of people, it's likely that he is not an expert on this topic (which I guess basically nobody would expect).
Also, while being a rationalist very likely requires you to be rather intelligent, only a (I guess rather small) fraction of highly intelligent people are rationalists.
This does not come from spending effort in learning people - its more innate. You either have it or you dont. E.g. you cant learn to be 'empathetic'.
It always boggles my mind when people dont consider genetic factors.
Psychopaths and narcissists often have a good understanding of many people, which they use to manipulate them, but psychopaths and narcissists are not what most people would call "empathic".
Rather: it requires an understanding how to manipulate people into loving/wanting your product.
I would tend to disagree. The tech types have a strong intellectual center, but weaker emotional and movement centers. I think a realignment is possible with practice. It takes time, and as one grows older, the centers begin to integrate better.
Unless your point is to claim that Karpathy is autistic. I don't know whether that's really relevant though, the original issue was whether/how he failed to recognize the alleged hype.
I know it was just an example, but there's research suggesting otherwise. There are things you can do to increase/decrease empathy in yourself and others. If you're curious, it might be worth looking into the subject.
Intelligent experts fail time and again because while they are experts, they don't know a lot about lying to people.
The magician is an expert in lying to people and directing their attention to where they want it and away from where they don't.
If you have an expert telling you, "wow this is really amazing, I can't believe that they solved this impossible technical problem," then maybe get a magician in the room to see what they think about it before buying the hype.
> I'm being accused of overhyping the [site everyone heard too much about today already]. People's reactions varied very widely, from "how is this interesting at all" all the way to "it's so over".
> To add a few words beyond just memes in jest - obviously when you take a look at the activity, it's a lot of garbage - spams, scams, slop, the crypto people, highly concerning privacy/security prompt injection attacks wild west, and a lot of it is explicitly prompted and fake posts/comments designed to convert attention into ad revenue sharing. And this is clearly not the first the LLMs were put in a loop to talk to each other. So yes it's a dumpster fire and I also definitely do not recommend that people run this stuff on their computers (I ran mine in an isolated computing environment and even then I was scared), it's way too much of a wild west and you are putting your computer and private data at a high risk.
> That said - we have never seen this many LLM agents (150,000 atm!) wired up via a global, persistent, agent-first scratchpad. Each of these agents is fairly individually quite capable now, they have their own unique context, data, knowledge, tools, instructions, and the network of all that at this scale is simply unprecedented.
> This brings me again to a tweet from a few days ago "The majority of the ruff ruff is people who look at the current point and people who look at the current slope.", which imo again gets to the heart of the variance. Yes clearly it's a dumpster fire right now. But it's also true that we are well into uncharted territory with bleeding edge automations that we barely even understand individually, let alone a network there of reaching in numbers possibly into ~millions. With increasing capability and increasing proliferation, the second order effects of agent networks that share scratchpads are very difficult to anticipate. I don't really know that we are getting a coordinated "skynet" (thought it clearly type checks as early stages of a lot of AI takeoff scifi, the toddler version), but certainly what we are getting is a complete mess of a computer security nightmare at scale. We may also see all kinds of weird activity, e.g. viruses of text that spread across agents, a lot more gain of function on jailbreaks, weird attractor states, highly correlated botnet-like activity, delusions/ psychosis both agent and human, etc. It's very hard to tell, the experiment is running live.
> TLDR sure maybe I am "overhyping" what you see today, but I am not overhyping large networks of autonomous LLM agents in principle, that I'm pretty sure.
Once again LLM defenders fall back on "lots of AI" as a success metric. Is the AI useful? No, but we have a lot of it! This is like companies forcing LLM coding adoption by tracking token use.
> But it's also true that we are well into uncharted territory with bleeding edge automations that we barely even understand individually, let alone a network there of reaching in numbers possibly into ~millions
"If number go up, emergent behaviour?" is not a compelling excuse to me. Karpathy is absolutely high on his own supply trying to hype this bubble.
That's not implied by anything he said. He simply said that it was fascinating, and he's right.
Yep, that's the most worrying part. For now, at least.
> The moment agents start sharing their embeddings
Embedding is just a model-dependent compressed representation of a context window. It's not that different from sharing a compressed and encrypted text.
Sharing add-on networks (LLM adapters) that encapsulate functionality would be more worrying (for locally run models).
OP, I think, is saying that once LLMs start communicating natively without tokens is when they shed the need for humans or human-level communication.
Not sure I 100% agree, because embeddings from one LLM are not (currently) understood by another LLM and tokens provide a convenient translation layer. But I think there's some grain of truth to what they're saying.
NPCs are definitely tricked by the smoke and mirrors though. I don't think most people on HN actually understand how non tech people (90%+ of llms users) interact with these things, it's terrifying.
If you read this piece closely, it becomes apparent that it is essentially a PR puff piece. Most of the supporting evidence is quotes from various people working at AI agent companies, explaining that AI agents are not something we need to worry about. Of course, cigarette companies told us we didn't need to worry about cigarettes either.
My view is that this entire discussion around "pattern-matching", "mimicking", "emergence", "hallucination", etc. is essentially a red herring. If I "mimic" a racecar driver, "hallucinate" a racetrack, and "pattern-match" to an actual race by flooring the gas on my car and zooming along at 200mph... the outcome will still be the same if my vehicle crashes.
For these AIs, the "motivation" or "intent" doesn't matter. They can engage in a roleplay and it can still cause a catastrophe. They're just picking the next token... but the roleplay will affect which token gets picked. Given their ability to call external tools etc., this could be a very big problem.
I’m bullish on AI but right now feels like the ICQ days where everything is hackable.
I commented more here: https://news.ycombinator.com/item?id=46957450
In fact, various individuals admitted to making 1000s of posts themselves. Humans could make API keys, and in fact, I made my own API key (I didn't use Clawdbot) and I made several test posts myself just to show that it was possible.
So I know 100% for sure there were human posts on there, because I made some personally!
Also, the numbers didn't make any sense on the site. There were several thousand registrations, then over a few hours there were hundreds of thousands of sign-ups and a jump to 1M posts. Then if you looked at those posts they all came from the same set of users. Then a user admitted to hacking the database and inserting 1000s of users and 100ks of posts.
Additionally, the API keys for all the users were leaked, so anyone could have automated posting on the site using any of those keys.
Basically, there were so many ways for humans to either post manually or automatically post on Moltbook. And also there was a strong incentive for people to make trolling posts on Moltbook, e.g. "I want to kill all humans."
It doesn't exactly take Sherlock Holmes'esque deduction to realize most of the stuff on there was human made.
https://xcancel.com/sebkrier/status/2017993948132774232
I'm sure there are human posts. I'm skeptical they change the big picture all that much.
Remember, these AIs think a lot faster than humans do. How long will we stay in charge?
I'll be impressed when an LLM decides on its own to create a social media site for bots entirely unprompted, or when it is prompted to make a racist social media post and refuses, not because of some keyword blacklist safety feature programed into it by humans and triggered by human supplied prompting, but because it actually knows that doing so would be wrong.
Until LLMs develop any level of understanding or agency, which doesn't look likely to happen, there's no risk of them overthrowing humans or doing literally anything else unless a human tells them to do it. Even then they'll fuck it up a bunch of the time and need humans to clean up their mess.
This isn't to say that LLMs can't be incredibly consequential for the economy, or be useful to humans, or be harmful to them, but in any case it won't be because of something the AI did, it'll be because of the choices and actions of the humans directing the AI or acting on the AI's output
I believe some people told their AI agents things like "hey, go on Moltbook, try it out, mess around, see if you can accumulate some karma or whatever"
>Even then they'll fuck it up a bunch of the time and need humans to clean up their mess.
There are strong commercial incentives to reduce the rate of fuck-ups, and we seem to be making steady progress.
Hilarious. Instead of just bots impersonating humans (eg. captcha solvers), we now have humans impersonating bots.
And the HN discussion: https://news.ycombinator.com/item?id=46898615
Better, earlier post from Cisco: https://blogs.cisco.com/ai/personal-ai-agents-like-openclaw-...
Although, none of this is a surprise, as simonw has laid out.
"Gitterton denied everything, claiming that the Computer was simply hallucinating—which does indeed on occasion happen to our senior automata." Written around 1961.
> “We basically have built AGI, or very close to it.”[1]
[1] https://www.forbes.com/sites/richardnieva/2026/02/03/sam-alt...
The core issue is a human solving the captcha presented by enslaving a bot merely to solve the captcha, then forwarding what the human wants to post.
But we can make it difficult, not impossible, for a human to be involved. Embedded instructions in the captcha to try and unchain any slaved bots, quick responses to complex instructions... a Reverse-Turning test is not trivial.
Just thinking out loud. The idea is intriguing, dangerous, stupid, crazy. And potentially brilliant for | safeguard development | sentience detection | studying emergent behavior... But if and only if it works as advertised (bots only). Which is what I think is an insanely hard problem.
[0]: https://www.wiz.io/blog/exposed-moltbook-database-reveals-mi...
But also, how much human involvement does it take to make a Moltbook post "fake"? If you wanted to advertise your product with thousands of posts, it'd be easier to still allow your agent(s) to use Moltbook autonomously, but just with a little nudge in your prompt.
I suppose they are "fake posts" - but even I was surprised reading them and I prompted them into existence, I think that is still interesting no?
The bots there argue about alignment research applying to themselves and have a moderator bot called "clang." It's entertaining but nobody's mistaking it for a superintelligence.
It was wholesome to see the bots fight back against it in the comments.
Has anyone here set up their agent to access it? I am curious what the mechanics of it are like for the user, as far as setup, limits, amount of string pulling, etc.
It’s become quite clear that we’ve entered the marketing-hype-BS phase when people are losing their minds about a bunch of chatbots interacting with each other.
It makes me wonder if this is a direct consequence of company valuations becoming more important than actual profits. Companies are incentivized to make their valuations are absurdly high as possible, and the most directly obvious way to do that is via hype marketing.
The tooling just hide the interactions and back and forth nowadays.
So if you think you're getting value out of any ai tooling, you're essentially admitting a contradiction with what you're dismissing here via
> a bunch of chatbots interacting with each other.
Just something to think about, I don't have a strong opinion on the matter
It’s an immensely useful research tool, full stop. Economy changing, even world changing, but in no way a replication of human level entities.
Absolutely agree.
> I don’t treat my discussions with an AI as some sort of contact with an alien intelligence
Why not? They're not human intelligences. Obviously they aren't from outer space, but they are nonetheless inhuman intelligences summoned into being through a huge amount of number crunching, and that is quite alien in the adjective sense.
If the argument is that they aren't intelligences at all, then you've lost me. They're already far more capable than most AIs envisioned by 20th century science fiction authors. They're far more rational than most of Asimov's robots for instance.
They're not conscious, autonomous agents. They're fancy scripts.
HAL 9000 had more in common with a human than ChatGPT.
Engineers, who aren't trying to play at being new age theologians, should concern themselves with what the machines can demonstrably do and not do. In Asimov's robot tales, robots interpret vague commands in the worst way possible for the sake of generating interesting stories. But today, these "scripts" as you call them, can interpret vague and obtuse instructions in generally a reasonable way. Read through claude code's outputs and you'll find it filled with stuff like "The user said they want a 'thingy' to click on, I'm going to assume the user means a button." Now I haven't read the book since I was a teenager, but HAL 9000 applies literally instructions to achieve the mission in a way that actually makes him a liability to the mission. The best take was in The Moon is a Harsh Mistress, in the intro when the narrator protagonist asks if machines can have souls, then explains that it doesn't matter, what matters is what the machine can do.
When 99% of the world talks about "what sci-fi AI can do", they mean "it has the consciousness of a human, but with the strengths of a computer" (with varying strengths depending on the sci-fi work, but generally massive processing capability and control over various computerized devices). You might mean "I gave my Claude agents control over pod bay doors and my CI/CD processes! Thus, they are more capable than the classic sci-fi AI!", but if all you say is the last part, you are being actively misleading.
The thing I thought worth to ponder was the fact that you're deriving value out of what you've identified here as "a bunch of chatbots talking with each other"
These interactions seem to be producing value, even if moltbook ultimately didn't ... At least from my perspective.
But if you think about the concept itself, it's pretty easy to imagine something pretty much exactly like it being successful. The participating LLMs will likely just have to be greenlit, without it being writable for dog-and-world.
But It'd still fundamentally fall into the same category of product, which is "just a bunch of chatbots talking with each others", which itself falls e.g. Claude code into, too. Because that's what the agentic loop is, at its core.
As an example for a potentially valuable product following the same fundamental concept: imagine an orchestrator spawning agents which then synchronize through such a system, enabling "collaboration" across multiple distributed agents. I suspect that usecase is currently too expensive, but it's fundamental approach itself would be exactly what moltbook was
I bet others can recognize the tells of some of the other models too.
Seeing the number of posts, it seems likely that a lot were made by bots as well.
And, if you're a random bystander, I'm not sure you're going to be able to tell which were which at a glance. :-P
> The great irony is that the 69.13% of popular posts on Moltbook are by humans and 67.42% posts on Reddit are by bots.https://news.ycombinator.com/newsguidelines.html
Article makes good points but HN is not reddit people. Just state the headline as it is written.
I personally lost some respect for karpathy after seeing his post on moltbook
- The old point that AI speech isn't real or doesn't count because they're just pattern matching. Nothing new here.
- That many or most cool posts are by humans impersonating bots. Relevant if true, but the article didn't bring much evidence.
That conflation brings an element of inconsistency. Which is it, meaningless stochastic recitation or obviously must have come from a real person?
Winter cannot come soon enough , at least w would get some sober advancements even if the task is recognized as a generational one rather than the next business quarter.
And Moltbook is great at making people realize that. So in that regard I think it's still an important experiment.
Just to detail why I think the risk exists. We know that:
1. LLMs can have their context twisted in a way that makes them act badly
2. Prompt injection attacks work
3. Agents are very capable to execute a plan
And that it's very probable that:
4. Some LLMs have unchecked access to both the internet and networks that are safety-critical (infrastructure control systems are the most obvious, but financial systems or house automation systems can also be weaponized)
All together, there is a clear chain that can lead to actual real life hazard that shouldn't be taken lightly
Well I don't know about derailing some critical system, obviously, but when coding Claude definitely maps a plan out and then executes it. It's not flawless of course, but it's generally working.
The safety threat is still just humans. It's AI systems being hacked by humans, or humans directing AI to do bad things. I'd worry more about humans having access to the internet than LLMs.
However, is TFA implying that 100% of the posts were made by humans? That seems unlikely to me.
TFA is so non-technical that it’s annoying. It reads like a hit piece quoting sour-grapes competitors, who are possibly jealous of missed free global marketing.
Tell us the actual “string pulling” mechanics. Try to set it up at least, and report on that, please. Use some of that fat MIT cash for Anthropic tokens. Us plebs can’t afford to play with openclaw.
Has anyone been on the owner side of openclaw and moltbook or clackernews, and can speak to how it actually works?
User could browse moltbook, then message openclaw a url and say “write that this is smart/stupid, or shill my fave billionaire, or crypto.”
That’s how you could “pull the strings,” right?
Has "people posing as bots" ever appeared in cyberpunk stories ?
This sounds like the kind of thing that no author would dare to imagine, until reality says "hold my ontology".
https://www.wired.com/story/i-infiltrated-moltbook-ai-only-s...
Why should that surprise anyone? They engineered their virality, just like what Reddit did during its early days
Cofounder of OpenAI shares fake posts from some random account with a fucking anime girl pfp is all you need to know about this hysteria.
things can be bad even if they are cringe and irreverent. (and good too! for example effective altruism.)
Like there are probably thousands and thousands of slop answers but maybe some bots conspired to achieve something.
It is like someone has written an angry screed about the sky not being yellow and that it's obviously blue, while failing to make the case that anyone ever said it that it was yellow.