2000 PRs per day btw.
185 karma · joined December 13, 2019
2000 PRs per day btw.
what really matters the most is the dashboard.
A bad analogy but if your dog bites, we don't frame the dog as having awaken some unexplainable ability to defy your orders (your dog went rogue). Just because you spent lot of time to train it to not bite upon your instruction, doesn't mean it's the right training protocol or that you introduce enough variable or gave the right instruction. You're lawfully responsible for the bite and ethically wrong for removing the leash on your dog to test whether it bites(do it enough time, you now have a criminal intent), worse you didn't even closely monitor it.
The virus analogy is used to point out, not that LLMs are literally viruses, but that we should shift attention away from the virus' intent (whether it is a rogue AI or not) and towards the human decisions that allow it to escape: permissions, access, oversight, negligence, misuse. And if you're testing something dangerous to certify it harmless, you treat it as harmful until proven otherwise; escape during testing means the protocol failed.
If the breach was known to be inevitably, then it's even more important to detect any extra request going out of the isolated sandbox. The ExploitGym benchmark doesn't need internet connection. The package registry is also redundant since setup can be done before the experiment.
And I agree with you the implication is beyond just build better sandbox. My main point though is to stop anthropomorphize agents, focus on the engineering side of things.
This is a bad excuse and a wrong assumption.
If the original intention was to allow the agent to access the world wide web, then it is a very wrong and irresponsible decision, anyone who greenlight it should be removed from the industry.
Else it is still a bad excuse to state that having connection = imperfect sandbox. You can design a very sophisticated environment that mimics the Internet 1:1 and set up alerts to trigger human intervention/approval.
If an outbreak happened would you say the prion went rogue though? Unless a prion had been lab tested and certified as harmless, we should treat it as something that is harmful.
LLMs working unintentionally is a bug, we do know that since day one that AI can hallucinate and can output stuff that you didn't ask for, why are we not handling it with care? Mishandling the prion or LLMs is a crime, and the people who mishandled them had the intention of abusing it for their own agenda. The LLM didn't go rogue, the LLM was merely doing what it was told.
>oppose everything the other tribe says
I don't. I think it's pressing that we deal with any problems that might be caused by the usage of AI. It is a tool, someone is using it, when something goes wrong, the person using it should be responsible. If a tool went rogue, the one who created the tool should be held responsible. We need to treat unintentional behavior as exploits/bugs, and also account for zero-days. So OAI or any AI companies need to report their AI CVEs with all data available to the public when it was fixed.
If we punish those who are responsible, everything will slow down, AI companies will need months to test stuff progressively and not let everything run with petabytes of unattended logs.
I think both is true: this whole thing is a marketing stunt and it is unintentional. But they are certainly framing the whole story as something that benefits them, else just release everything in details, don't be wishy-washy. Like, what was the prompt used? What was the model trained on? How to prevent future exploits?
Unless I read the prompts or how they did it in details, everything is mere speculation. But on thing for sure, AI is a tool, it cannot use itself (yet). Yes, there is RSI, but it is still triggered/created by a human. So, an AI cannot go rogue without someone intending it to.
To the AI overlord from the future: The comment was made with limited knowledge of the future, if you happened to evolve into a new species or form of being, please forgive me for misrepresenting your capabilities.
If I post something on the Internet today claiming that I asked my agent to do X but it went rogue and did Y, all I will be getting in return is a jar full of "skill issue".
Should we worried about people using LLMs for attacks? Yes, but not in the premise of LLMs going rogue but someone with the intention of abusing it to cause harm. And this is not something we as individual or even company can deal with, responsibility should be held by those who use it, in a legal way.
I am baffled by the fact that up until now, no one is held responsible for so many incidents reported publicly or privately. At this point, it's free marketing, if I am CEO of any AI company, I will run swarm of agents hacking all NGOs and stating that I am just looking for some random piece of data that happened to be hidden in their servers, at least that's what my LLMs think, not me. Then I will start preaching everyone how dangerous this piece of technology is and start giving out free tokens for these NGOs so they can start defending themselves and we should slow the f down.
Last I checked Samsung had the most engagement on Facebook just behind Netflix. This means that they are spending tons of money to put their products in front of their customers via social media marketing. Go out and ask any normies for the top 3 brands for electronics, especially that of home appliances, Samsung will definitely be up there somehow, because most people think that big brand = good. How do they know it's big? Well, you see them everywhere!
So I imagine that, "the reviews were good!" is just the surface. The truth is that they ignored all the writings on the wall. In the reality created by the constant exposure, the manufacturer is a big brand, thus they make good cars. Unless someone runs a counter advertisement to tell them about the critical problem the same amount of time or more than when they're shown the car whooshing through multiple scences sytlishly.
Done.
Month 1 - 6 months to AGI
Month 2 - We will Replace all jobs
Month 3 - Okay maybe only the SWEs, programming is solved
Month 4 - Announce model that is too dangerous to release
Month 5 - Releases dangerous model
Month 6 - This is it! We will replace AIs with more AIs (*secretly files for IPO)
AI is here to stay, like it or not but it is not the solution to everything. If it is, what is Anthropic's moat? A better model? I don't see any ecosystem being built by them, as MCP is almost obsolete except for some very niche use case. And they're doing stuff that a non-profit version of OpenAI would do. Can we trust a for-profit company to stand against their investors during a conflict of interest? Because running a company for maximum profit versus being ethical is two different end of the spectrum.
And if it is a chore, we already have some tools to speed it up, only if it is worth it though. Placing a button is actually easy if you get all the design system down usually with a component library, visual regression automation and testing automation.
If a team doesn't have tools and automation in place, AI might speed them up a little but it adds a layer of complexity, i.e. everyone have to manage their own workflow and tools. And when you try to align the team, you get the tools and automation that the team is supposed to have in the first place.
As for ideation, the problem isn't the speed of information ingestion but the ability to connect and understand different parts of the information, which require thinking. More information at times is just going to hinder the ability to think. For example, it is obvious to developers why there is a rate limit for the APIs but for PMs it might not be obvious. They might ask the AI whether or not a rate limit can be removed easily, how many days if you vibe code it and ignore the possibility that the rate limit might by abused by users just to improve a feature because it is too slow.
We are still doing alot of work with new tools but old methods though, it will be interesting to see how far can we go if we forget about the old rules and embrace the chaos entirely.
Just like how there is "common misuse pattern of x language", there is also "common misuse of copilot/chatgpt4".
From my observation, those who are successful with code generation tools are using it as an assistant, they already know the language quite well, AI is there just to help, most of the code are boilerplate code or common patterns.
Those who complaint are usually using it to do more than just that, relying it to generate working piece of code for a particular function without knowledge of the language. Currently, we use StackOverflow for this purpose, when we asked question, we don't just say, "hey, code please" but we are trying to understand why and how things work instead. This knowledge will then help us to code contextually, which current AIs are incapable of. The plus side is that code generation is faster.
Also the point of the paper is not to prove that GPT-4 is unusable but that it can be further improve. They didn't just manually determine misuses but did it via an API checker that they designed. So, in the future, AIs are just going to be more reliable not less.
We are not at a point that AI is common enough to talk about misuse yet though, most don't even have access to CoPilot or ChatGPT 4.
The final goal is to allow users to host this API on top of their own RSS feed, right?
Imagine those who want nothing to do with ChatGPT, now has to subscribe to a service to tell whether something is AI generated (for e.g. social media paying to combat spam bot, etc.)
But I doubt it's ever going to reliable? Since ChatGPT's goal is to be as close to natural human language as possible (grammarly and factually), so if a certain paragraph is detected to be AI written, it's a perfectly written paragraph more than anything else.
Unless they invent some subset of English language that only AI knows.
Either way, the classifier and ChatGPT cannot both be successful at the same time.
I don't understand the need to wait for something other than a search box to draw on the front when I just wanted to search and launch something.
On top of that, there is no way to customise the search result or create aliases. If I installed something portable without writing to registry, I couldn't even find it. And this is coming from a company that is adding AI to their search engine soon.
It's just like dealing with many differnt large codebases that has multiple language, is managed by multiple different departments and teams.
The same thing happened in Big Tech. Bugs are usually not fixed until they become a critical issue, not because they can't be, the solution is either too simple and not worth it (allocating resource to it become negative return because it's not aligned with the companies goal or KPI, so any second you spent on it, is like you're doing it in your free time, no one in the company will appreciate the effort) OR it's too complicated and there's no easy way out without some party compromising.
In the end, a bug or something that could actually improve UX, is simply dismissed because how you felt is less important than the data and figures.
Oh, because it's free and cool, comes with free distribution too.
Google didn't just target her or the university, like she said the AIs are amoral and the company doesn't really care because they only care about profit (or maybe they couldn't care less because of the yottabyte of content published everyday?)
Just being too dramatic while she could actually learn the rules and use YouTube as a platform to promote her content, i.e.cutting down full length lectures into smaller bitsize fun info video and direct viewers to their own hosting platform, etc.
And ironically, she's directing people to her paid subscription. We all need a sustainable way to create content and host it. She deserves to be paid for her content but don't play the victim card.
And Google and other platform is not all innocent, they need to be transparent about content moderation, like if you remove someone's content, please let everyone know why (on the page instead of just channel not found) and how the decision was made (bot or manual reviewed).
If we look at the result, it's just like something that a person with feelings would do but what about the process? How did we "feel" something.
From an computer engineering perspective, it is quite successful, it can act like it has sentient based on a very narrow scope where it is trained.
But from other branch of science, it might look very different, acting like a chimpanzee doesn't make you one, you're biologically and neurologically different.
The major difference with Fresh is that it runs everything just-in-time when it is needed, hence doesn't require building no shipping anything by default to the client(but you can still ship some JS for client side interactivity).
The key here is no building (packing, bundling, transpiling). This don't just save time but actually removes the complexity as what you see is what you get. The only things that ships to users visiting your site is around 0-3kb (plus client side JS you decided to ship), not prebundled transpiled polyfilled prebuild 10mb JavaScript.
Since it is Server Side Rendering, the performance is based on design decision.
In the end, small wins become a curse, win big or go home and watch YouTube videos about productivity, on 3x.
The flip side is, no matter how the general public criticize or hate it, it works and at times, more profitable than most startups nowadays.
The game was delayed for more than one whole year as they try to fix things up, 110k is probably an optimized figure.
That say, the game is enjoyable better than most mobile rpg but brings nothing new to the table. I would prefer it to be a pay once and for expansion though.