Maybe don't let Muse run your Facebook Marketplace account
threads.com
threads.com
Nice touch by the mechanical parrot, to worthlessly owning it.
Wonder what quantization Meta runs these models on, Q4?
I wonder how much of this unreasonably risky and illogical behavior is a genuine property of LLMs and how much is there because LLMs come from "move fast and break things" startups.
"By the way don't do this again" <- as if the AI has the ability to ingest and systemically diffuse this.
I think Zuckerberg himself is deeply into the Koolaid, and is likely himself unaware of the limits of this tech.
He's probably surrounded by enablers.
If it was Claude I'd expect it to now put some form of this into every output even when not really related to the task: "No offers were accepted without consulting you and I haven't shared your address or availability."
Is it guaranteed to work? No.
And obviously it's a terrible idea to set up a chatbot to communicate, negotiate deals, and handle logistics on your behalf.
You have explained literally why it would not work.
The AI can absolutely not depend on 'arbitrary statements in some file' as operational policy.
For a very, very narrow scope of work, when it's well defined, when the information is rigorously applied, sure ...
But they don't have that.
The are throwing agents out there like they can handle this degree of complexity and nuance, when they cannot.
100% failure rate over any period of time.
However, it is untrue that it doesn't have the capability to memorize an instruction and diffuse it to new sessions.
Simply saying "do not ever do this again" can result in the behavior not reoccurring with any likelihood.
You'd have to benchmark whether with the instruction in place it would violate it, and in how many cases, so that you can understand the risk better.
I think we get that.
It's completley unreliable, which is the issue.
I think it's unreasonable to pretend that this couldn't possibly work, that you understand to what degree it does work, and that the only thing we should be discussing here was that it isn't deterministic, which every reader already knows.
Your argument exemplified by 'the ai can write to a file and use that information in context later' - implies a 'capability' to do theoretically do something, but not with any consistency at all.
Not in any scenario - even with human oversight - would be 'trust' any system to reliably work on the basis of arbitrary memory.
'It can possibly work' does not map to any reasonable conclusion that it will work with any degree of fidelity.
There are almost no examples where we rely on AI in this way. Chatbots for customer engagements etc. are merely thin language wrappers around deterministic systems.
It's 'possible' that Meta is doing something rigorous and deterministic but it's not remotely reasonable to make that assertion because we literally have the evidence right in front of us. That's literally what the article is about.
The Wright Brothers plane 'technically took flight' but it's not reasonable to conclude that it can fly passengers safely in any way, especially when the story is about a flight crashing.
It would be novel actually, if there were a story about Meta's unique use of AI in these kinds of information flows, that transcended what we all recognize as AI's inability to reliably process information, given what we know about AI and how it works aka 'it can write to memory'.
Here I expect factual discussion rather than misleading statements that suggest there was no way for the Muse assistant to follow an instruction never to do something again.
Vague arguments that such instructions are not deterministic are uninteresting, because it is obvious.
Day one our agentic platform goes live it causes a reportable compliance issue. Massive clean up. Reputational impact. Turned off.
No one held accountable still. Assuming that adding more guardrails will fix everything.
'Managers Delusion' - which includes techies as well, to be fair, at least they have the excuse they are one step removed.
They are morally bankrupt and will try and do it again and again if it's stopped too early.
'Say Hey There' by Atmosphere
It should be ready in about two more weeks! How many models have a Ph.D level intelligence now? I feel like I’ve been hearing that for about a year at this point.
“Models aren’t improving incredibly fast” seems a very odd point to be making.
Say that you want to make a garage sale, but to get a better reach, you want to list item by item on marketplace. If you've never done it before, making 100 listings is such a pain in the ass you never want to do it again (I used to flip stuff for a living back in college, and it is a hassle)
At least with this, you can now just take pictures of everything, and tell the agent to research everything, list the stuff, and arrange for pickups.
Obviously not something I'd like to do for any serious deals or higher $ items. But for random stuff? I mean, it sort of beats hosting a garage sale or taking them to the flea market.
I'm not all pessimistic about this.
Many more such cases are to come.
The 'deliberate failure' is on Meta here.
>they will hook up actually important stuff to agents because its the future
>???
>400 dead
It frustrates me no end that so many folks are so ready and willing to accept the hype and lies about what this technology actually is or can do, when what it actually is and can really do is already amazing enough on it's own even without all the ridiculous AGI/ASI anthropomorphising bullshit. Falling into this ridiculous "machine-god" hype-cult is kinda holding this technology back from it's true full potential, as everyone's all busy doin' stupid stuff it's not really capable of doing well, or designed for instead of focusing on using it for the (many) things it is really really good at doing (various really useful and powerful language, vision, and audio related tasks).
apparently people use Threads. I suppose the same kind of people who connect Muse to Facebook Marketplace.
With that said, in case Threads isn't available for whatever reason, I've uploaded screenshots of everything (I think?) here: https://imgur.com/a/mceW9WF
I've never had that problem on Mastodon, but I can see how it could be annoying if you follow a lot of people who mainly post screenshots of people's posts from other social media sites. Though you could solve that problem by unfollowing those people...
X has FB and Insta screenshots
Insta has FB and X screenshots
FB has Insta and X screenshots
Bluesky has FB, Insta and X screenshots
Masto has Bluesky, FB, Insta and X screenshots
So what I'm saying is on Mastodon, you are exposed to the most screenshots of other networks, because there are no techniques to embed them and people refuse to link to them for the reason they aren't on those networks in the first place.
As for AI, i know toddlers who can prompt but are unaware of what a right click is.
[0] Which makes browsing Modrinth or Curseforge for Minecraft mods HILARIOUS - "you make a [extra dimension] portal like this: [image not available in your region]" "Gallery: [image not available in your region] [image not available in your region] [image not available in your region]"
But Kik (from Medialab who are the people being fined by ICO) is still online in the UK which somewhat undercuts that argument?
> It's not due to OSA [...] in regards to how they handled children's data
potato, tomato.
Holy shit social media is bad. I've been using Mastodon for so long I didn't realize how bad the others are.
I just tried on my laptop though, both Waterfox and Firefox. It's just completely unusable there, with a ton of CORS and CSP errors so no CSS loads and no JavaScript loads. Happens even when I disable all forms of tracking protection and ad blocking; I think their site is just broken in a way which presumably happens to work in Google Chrome.