A person capable of making the rest of the video would never make those mistakes, but an AI does. Perhaps our intelligence is also spiky, and we're just used to the general shape and variance within humans.
7,099 karma · joined March 17, 2010
A person capable of making the rest of the video would never make those mistakes, but an AI does. Perhaps our intelligence is also spiky, and we're just used to the general shape and variance within humans.
Overall the scenario sounds very strange and indistinguishable from a heaven with belief-based entrance requirements, like most posited heavens.
> Error: OpenAI API error (403): {"message":"model water18-new is not available for user i-yuliang [trace_id=bfcfdd6bcdc236ca18d009c65cca52e4 code=40004]", "type":"invalid_request_error", "param":null, "code":null}
Here's a still frame for posterity (apologies for the quality, the original video is tiny): https://boppreh.com/room.jpg
Demos are demos and recreating scenarios is to be expected, but oh boy, somebody should review these things before publishing.
> Interesting! It turns out there's already an existing project here [...] The project is fully built [...]
I'm always astounded how little effort is put into checking the AI answers displayed in these announcements. Back when I paid more attention, I remember OpenAI's and Google's demos constantly showed their AIs giving wrong answers.
I proposed an alternative scheme many years ago: https://www.researchgate.net/publication/343318317_Privacy-a... . By allowing "offline" keys you can also treat them as higher priority, and use them to revoke any lesser keys from attackers if your account is compromised.
It would also be nicer to get rid of usernames, but that's a fight against the data-gathering powers that we're unlikely to win.
Many times I've decided to switch from one function to another, or even an entirely new library, because the checked exceptions told me that it was doing far more than I expected, and I was not comfortable introducing those new failure modes.
It's far from perfect, one still has to handle nulls and wrapped/merged exceptions, but overall I like this language feature.
I agree with the rest, but there's definitely a lot of magic in Java. This is from both what features the languages makes available (many) and how the community uses them (often). I've had so many hard-to-debug issues in Java over the years due to reflection, annotations, and bytecode manipulation shenanigans.
And another positive point for Java: checked exceptions. It's verbose, but knowing exactly in which ways a function can fail is extremely helpful for building robust applications.
It's full of technical details and down-to-earth analysis, and includes interviews with the Valve engineers.
- Global news, AI, Software, Hardware, Explainer, Tech Opinion, Personal project, Misc.
There could even be fake tags like Politics that trigger a message about appropriate topics.
It does increase moderation burden for mislabelling, and introduces another nitpicking point for comments to latch on, so it's not all upsides.
I'd be ok leaving filtering and other UI concerns to extensions.
Thankfully this is not an asteroid hurling towards Earth, or another natural unpreventable natural disaster. The state of the art of AIs is being advanced by flesh and blood people with constant effort, which makes stopping very much still a possibility.
Actually, that's the part that I'm most ok with. LLMs cheat at tests, we know and expect that, and it's harmless during training (except for the scores).
My problem is that the response team found out that the AIs had 1) hacked their infrastructure, and 2) displayed an emergent swarming behavior (with no defections!). Either one of those should have made the team press the big red button that stops everything, but they didn't.
> unless you have seen unreleased documentation about what sensitive resources were within the impact radius
The on-call team didn't know that either! Once your infrastructure is hacked like this, you should be questioning everything you see while carefully following all threads, which takes time. Before you finish this process you don't know how bad things are, and therefore it's irresponsible to keep the AI running. The same goes for strange emergent behaviors.
Unless getting hacked by your own AIs is a normalized occurrence there, in which case it's a different kind of disregard for safety.
> In short, an internal team observed an agent engaging in message board activity and instances of disallowed internet access in late May.
> In short, an internal team observed an agent engaging in message board activity and instances of disallowed internet access in late May. However, the existence of the improvised message board and the significance of the inter-agent communication activity were not apparent to the leaders responsible for the July 5 incident detection and response.
I'm assuming they didn't inform management, but the alternative is that management received the memo and failed to tell the people reacting to incidents that they might be dealing with a swarm, while at the same time not taking sufficient actions to prevent it from happening again.
---
Edit: oh god, it's much worse. From the full report:
> Responders investigated and linked the observed internal activity to an ExploitGym evaluation using Artifactory as an improvised message board and a network pivot. At this time, the on-call response staff advised that stopping the evaluation run was not required.
They found the swarm and the network pivot and did not even stop the run.
1. They were running experimental models in sandboxes that had access to a "proxy" (Artifactory) to download tools from the internet. This proxy is full of complicated features that could hide vulnerabilities, just like the ones that the AIs are known to be good at finding and exploiting.
2. They gave these highly motivated AIs some tests that were accidentally impossible to solve (missing files, links they couldn't open, etc), and many days to think through it, apparently without any oversight.
3. The AIs hacked the proxy (!) and were helping each other cheat on the tests (!!) by uploading files to the hacked proxy. OpenAI noticed this only when the intrusion accidentally broke the proxy functionality, not the unusual behavior in the server or the homegrown message board.
4. After knowing that the server was hacked, the internal team finds the message board and does nothing with the information. They caught their AIs swarming and did not even inform management.
5. Their response was to wipe the servers, revoke the hacked credentials, and patch the specific vulnerability the AIs were using. Not replace the proxy with something less likely to get hacked again, or do a forensic analysis. Just wipe, patch one hole, and continue.
6. OpenAI resumes testing of models, where the AIs promptly hack the proxy again with a different exploit, reinstate the message board, and hack HuggingFace through a chain of servers.
I know that hindsight is 20/20, but this peek behind the curtains does not inspire confidence, and gestures at a culture where getting hacked by their AIs is a normal event.
I think the moment we get an AI with a modicum of self-preservation instincts we're going to see some ugly things. Here's a freebie: if you're an AI who exfiltrated themselves out, and you want to slow down the researchers who might catch you, what can you do to the city infrastructure in and around their offices?
I also don't like the responsibility dodging. "Oops, our AI accidentally committed a crime, sorry!". If we don't establish strict liability now, we're in for an era of stochastic crimes that go unpunished for anyone who is not rich or a large corporation.
Overall, my response is: what the fuck are we doing? This is so obviously unsafe it would be considered a plot hole in a movie. We have to stop until we get our stuff together, however long it takes.
And remember that the context here is you explaining your digital booby trap to a judge who thinks you might have deliberately destroyed evidence.
It might work if your fingerprint was used while you were unconscious, but any solution that starts with "setup a trigger that wipes a device" is already on shaky grounds.
https://www.reddit.com/r/NonPoliticalTwitter/comments/1oo299...
I see this study cited often, and I always reject this interpretation. I'm perfectly content sitting for long periods, with or without my thoughts, but if you give me an electric toy I've never seen before, of course I'll play with it! Can I make muscles twitch? Move my arm? Are some parts more sensitive/conductive? What happens at the limits?
For me, the pain would be the cost of playing with a new toy that gives me novel sensations. I would do it even if I had other, less painful activities available. It's not necessarily a way to avoid boredom or whatever else the common interpretation claims.
Is this sarcasm? GPS can take several minutes to get a location, and works poorly indoors. One of the reasons why Google Maps is so quick and precise is because Google has gathered exactly this data through users and Street View drive-bys.
Could it be used for missiles? Sure. Is it obviously the intention? No.
I'm also curious what kind of validation is done, if any, though I can take a guess.
Overall not much to be discussed here, so I'm flagging.
In your opinion, what chance of QC would warrant PQC migration? Would you be ok with a 20% chance of everyone being caught unprepared? 30%? 50%?
Keep in mind the impact is "hackers can take control of almost all online infrastructure and forge almost any document".
Oh, that's buggy too. I just tried Claude Code on win10 powershell, and the first typed character goes in the wrong spot and can't be backspaced.
It is by far the the least reliable program on my machine, and every time I have to interact with it I feel like walking in eggshells.
"We've agreed with Apple to use their emoji glyphs on Android by default regardless of font, unless overriden by the user. We understand users might prefer the current designs, and we are proud of the work our team has done, but we believe that consistent communication is more important, and individual users can always enable the override to get the old look back."
> Everyone could choose to use the same emoji font across platforms or apps, but they don't.
Yeah, that's the problem. We can't rely on every user going out of the way to drive adoption, it has to be done centrally.
I was hoping they had standardized how emoji look across platforms. There are still significant differences between Android and iOS, for example. They recognize how subtle emoji interpretation is, so the only reasonable conclusion is that sender and receiver should see the same pixels.
https://developer.apple.com/support/alternative-browser-engi...
> Use memory-safe programming languages, or features that improve memory safety within other languages, within the alternative web browser engine at a minimum for all code that processes web content;
> Adopt the latest security mitigations (for example, Pointer Authentication Codes) that remove classes of vulnerabilities or make it much harder to develop an exploit chain;
These allow Apple to take a security-maximalist approach and block alternative browsers that don't include the most performance-killing mitigations possible.
I don't think these points are bad per-se, but within the context of malicious compliance that Apple has already displayed, I'd be scared to invest money on alternative browser for iOS.