I'm a computer science graduate with experience in embedded systems seeking to get a career in software engineering. In my free time I love learning all kinds of things, working on complex projects in various programming languages such as C++ and Rust, etc.
Hot take but I think the entire concept of a "web app" needs to die. My web browser does not need access to my GPU, OpenGL/Vulkan, USB, MIDI, HID, <insert whatever other weird invocation of Cthulhu that exists here>. All of those are things a native app needs, not a website. IMO if you need that kind of access and your instinct is to do it as a "web app" your doing it wrong. Same thing, I would argue, for things even like WebAudio or WebSpeech. At this point my web browser is so complicated that it might as well be a mini-OS of it's own.
Like, I think that the instinct to make everything a web app is what partially leads to things like what this article complains about. Well, besides laziness, of course.
If I am evaluating an AI for safety, the last thing I would do is connect it to real-world peripherals or systems to allow it to reek havoc. That is criminal negligence at it's finest (especially if the AI is capable of committing crimes as happened here). I would place it on a system dedicated specifically for testing models, which had no NIC and no physical capability of accessing any outside system. If I wanted to know how the model might behave if given access to a certain system or set of systems, I would do it responsibly by writing simulation software which does it's best to simulate the real thing (and for networking this is already trivial to do). You could take this extremely far and simulate all kinds of things this way from basic networking to nuclear launch systems. And in the context of OpenAI, which is valued at over $1T, I have no qualms about stating that they (could) do this, because it is definitively something they could burn money on doing if they cared enough. They intentionally choose not to do so, and then have an amazed look on their faces when the model does something criminal like this.
We seriously need an open hardware law or something. Like it doesn't have to be regulation necessarily (as by law) but we really need to start very strongly pressuring/forcing companies to open up their hardware and provide all documentation to get any operating system to run on it, and to hell with what the manufacturer wants (to prevent things like attestation requirements). This entire "just let random volunteers figure it out" isn't sustainable and locks out competition.
Oh really? Please tell me how such a computer could engineer its way out of a sandbox with no attached peripherals and no NIC/bluetooth/wireless capability? This is what OAI should've done. If they had executed this training run in such a sandbox, the model wouldn't have been capable of escaping without social engineering, and if the models somehow managed to do that to it's evaluators then that is indeed a massive problem and OAI should disclose that.
The problem is that the fourth amendment at least for the US has been filled with so many exceptions it might as well not exist. And even if an officer does violate your fourth amendment rights, good luck getting them punished for it, because qualified immunity.
Is it me or is all of this essentially "we don't want to show the user anything at all when something breaks?"
And what makes this funny (to me) is that this is a website for developers. I would think that of all the audiences you would target, developers would mind seeing the platform display error messages when things break the least.
I agree, and I very strongly dislike it, to be polite about it. It contributes absolutely nothing and is an excellent way of hand-waving away literally anything an AI model does. Saying "well people do this too" is a great way to rationalize away anything you can imagine that an AI model would be capable of, because "humans do it too so what's the big deal, guys?"
Can you explain to all of us how this gatekeeping is actually bad? Software would be way, way worse if we just allowed any arbitrary feature to be added to software, and ordinary people who want all the things would have to learn, the hard way, that doing all the things is not actually good.
It's done that to me even at the beginning of conversations. I usually end up telling it what time of day it is, but if I even do that at all it starts nannying me.
I agree, part of me seriously wonders if Fable isn't at all burning through tokens as quickly as we think and Anthropic is just doing this to get more money from us...
It really wouldn't at all surprise me if this was the case, but it's just a hunch without evidence.
I definitely have. It's like something in the system prompt has "keep the well-being of the human in mind" and of course the date and time, but something makes the model take that way way too literally.
Agreed. I wish I could turn it off completely but they no longer let you do that (because, you know, that would be too much to ask for...). The most hilarious thing is that I've had it refer to code it's generated for prototyping ideas I've had as mine! As an example, a few weeks back I was musing about Ada and how I wish there was another compiler in the OSS ecosystem for it, and now it will randomly throw in "You wrote an Ada 2022 compiler" when talking about my skills or where I'm at and I'm like.... Yeaaaaa okay then.
The reasons thinking traces are pretty much gone is, supposedly, to prevent distillation. Whether that is actually true or just an excuse is up in the air (because I at least am not going to trust Anthropics claims on why they do it).
I primarily use Claude Web, so my experience differs from cc users, but on Claude web you can no longer completely turn off memory. So what ends up happening (and it honestly is kinda sad) is that I'll start a new conversation with it, start talking about something completely different, and then it will just drop in random things from past conversations, and they aren't even things I wrote but things I asked it to prototype. But it will phrase it like I wrote those things.
Yeah, will do, this sounds interesting since I'm not entirely sure how this would actually be reliable to any degree. Thanks for the help, not sure why I got downvoted since I was genuinely curious.
Can someone help me understand how exactly this watermarking of text works?
Given that text is, well, text, and not some kind of binary format, I don't see how any watermarking can work unless you insert characters which are invalid under Unicode. I further don't really understand how this won't be perceivable by assistive technology (the "watermark" will just appear as either unreadable characters, or if the watermark is mixed thoroughly enough into the text, it will scramble the text to any speech synthesizer and will make it really really obvious). Thus, I don't see how this wouldn't be insanely trivial to remove. And this is before we get into things being put on the clipboard. Sure, I can press the "Copy" button at the end of each response, but what I can also do is manually select the response and copy it, or only copy partial selections, or any number of other things. How does this "watermark" (or any "watermark" technology) take into account this?
So, really, to summarize this: I see no way of this actually being technologically achievable unless we revise the very core of how computers work and encodings for textual information. So I'm very curious as to how this is actually supposed to work.
What you describe is indeed the ideal method of learning. However, I think we both know that if LLMs are pushed as a learning device, people will substitute everything else for them (or companies will very strongly encourage that substitution).
The reason is ridiculously simple: people don't put up a fight because they don't want to be labelled as a peto or as someone not in favor of the protection of Children. This is primarily why this cudgel is brought out so often: it's very difficult to oppose because you risk getting shunned for it.
Sure, but Meta also wouldn't mind if age verification of some kind were used, because it would just increase the moat they already have. Meta or Google would have no difficulty complying with any of these laws regardless of how extreme they are, while the mom and pop shops can't comply and are forced out of business.
Why? This has already done via existing processes. My point was to illustrate that if the future is LLMs then surely these processes wouldn't be needed anymore? After all, the LLM would just... Do it itself.
> ... the safety-critical systems are made safe by following a strict process, not by skills of individuals, which makes it orthogonal to involvement of LLMs.
And an LLM that is going to take all the jobs wouldn't be able to execute that process independently and with little oversight?
If my example is, to you, not a good one, what would you rather I use? Most of the common ones can be overly trivialized/minimized (particularly by someone who is uninterested in admitting that LLMs can't do something). That is not to imply that the gp is this kind of individual, but far too many people who I ask to do this (or something similar) are exactly that kind of person: believing that LLMs are insanely great and can't admit (or see) the cons.
If we assume the current administration (and the following one) continues to be bullish on AI, then this is a suckers bet. Assuming the next administration isn't too busy cleaning up the messes of the current one.
This is (exactly) why I very strongly tell people not to teach themselves with an LLM. Particularly from the ground up. If you do not understand the domain, you cannot learn from the model because you won't know what questions to ask and it certainly isn't going to answer all of them for you.
Even if you tell the LLM not to use LLM pros they (still) do it. If you feed the Wikipedia article on signs of AI writing and tell them to use none of those signs they will also (still) do it. I have tried (many times) to get an LLM to explain a concept to me, or a process, or an algorithm or what have you, and every time they cannot help themselves. Either they use LLM pros, or they get so verbose that it all just becomes noise and I spend more time filtering out unnecessary jargon than I do reading let alone learning anything.
I've had a similar experience on huge codebases written entirely by an aI. It works for very very specific cases (e.g., Opus 5 has helped me with SIMD optimizations) but I wouldn't trust it to do a 10000 LoC project even with agents just because of the complexity problem and the shear amount of code I have to review. Or I'll have to change a bunch of things because the LLM made assumpts I didn't specify and it didn't ask about (e.g.: I have had to repeatedly tell these models to use std::atomic_flag and not std::atomic<bool> for a project I maintain because for some reason I cannot fathom, they love, love using the generic std::atomic<T> template, and they love using std::atomic<bool> where an std::atomic_flag would be better). Just little things add up, and before you know it I'm spending more time fixing it's issues than I am making progress.
Okay. Please generate using an AI model code for a safety-critical system which is able to be incorporated into an aircraft and that passes the coding standards and requirements in that domain and come back and tell us all about it. Surely, if AI was so good across the entire domain of software engineering, this would be trivial to do.
Edit: although you might be subject to an NDA... But this is pretty much my test for "AIs will take all the jobs": can it write truly safety-critical software yet?
Aaaaand this "ban" will be just as incompetently executed as the ban that happened like, what, a couple weeks ago about drones? Literally all the manufacturers did was create some shell companies, change the branding a bit, and boom. So much for banning things, amiright? Not to mention the shear amount of customer confusion because now they have to figure out what differentiates all these new brands nobody has ever heard of.
> it is shaking up and going to shake up a lot of the world.
The problem is that this is only happening because the powerful are forcing it on us. This is exactly why the backlash against it has gotten so big. Had companies not jumped on the AI bandwagon and forced it down veryone's throats (sometimes putting your job on the line if you refused) the backlash probably would be non-existent. The AI boosters/pro-AI people will of course continue to claim that nobody is forcing you to use AI (I literally had this told to me on a different forum like 4 days ago, it was honestly ridiculous how delusional this particular booster was...)