I know this is frustrating, we just literally don't seem to have better approaches at this time. But if someone can point to open approaches that work at scale, that would be a great start...
I know this is frustrating, we just literally don't seem to have better approaches at this time. But if someone can point to open approaches that work at scale, that would be a great start...
The whole premise of your argument doesn't make sense. I'm talking to a computer, nobody gets hurt.
It's like censoring what I write in my notes app vs. what I write on someone's Facebook wall. In one case, I expect no moderation, whereas in the other case, I get that there needs to be some checks.
The map is not the territory
Sociologists, anthropologists, philosophers and the like might find a lot of answers by looking into the details of what is included into genAI alignment and trace back the history of why we need each particular alignment
In reality, every model so far has been either a toy, a way of injecting tons of bugs into your code (or circumventing GPL by writing bugs), or a way of justifying laying off the writing staff you already wanted to shit can.
They have a ton of potential and we'll get there soon, but this isn't it.
Then what?
Racism is a totally different and sadder issue. I don’t have a good answer for that one, but knowledge shouldn’t be withheld because someone thinks it is “dangerous”
Those that think that are the truly dangerous. They think they know better and want to remove agency from people.
Racism is fine as well. I won't date out of my race and if you think there should be a law that I must that's not really freedom. As for hiring or not based upon race there's already a law against that.
The cure is often worse than navigating uneasy waters. Every time you pass a law you give a gun to a bureaucrat.
And, every law is indeed a gun.
I'd love an example of "guardrails" in action on a topic of relevance to actual adults. There's a connection I can't find between the ability to make racist memes and literally anything else I want to do with AI.
The user can use a tool for good or for bad. It is the responsibility of the user, not of the AI.
The way I see it is that the guardrails define which biases you're selecting for. Since there's no single point of view in the world you can't really set a baseline for biases. You need to determine the biases and degrees of bias that are useful.
> Biases in the baseline models are minimal because they were trained with large and wide amounts of data.
The baseline models contain almost every bias. When people start their prompt with "you are a plumber giving advice" they're asking for responses biased towards the kinds of things professional plumbers deal with and think about. Responding with an "average" of public chatter regarding plumbing wouldn't be useful.
> What the AI safety BS do is make models conform to their myopic view of reality and morality.
To me it looks more like people are in the early stages of setting up guidelines and twiddling variables. As I mentioned above with the plumber analogy, creating solid filters will be just as important for responses.
It's easy to see this as intentionally testing naive filters in an open beta, so I'd expect the results to change frequently while they zero in on what they're looking for.
> It is bad, it is cartoonish, it glows in the dark.
Some of the example images returned are so hilariously on the nose that it almost feels like a deliberate middle finger from the AI. It's done everything but put each subject in clown shoes.
Invariably those from the current mainstream ideology in tech.
> Responding with an "average" of public chatter regarding plumbing wouldn't be useful.
RLHF can be useful but I'd rather deal with idiosyncracies of those niches of knowledge than with a "woke", monotone and useless model. I like diversity, I don't want everything becoming Agent Smith. Ironically, "woke" is anti-diversity.
> It's easy to see this as intentionally testing naive filters in an open beta, so I'd expect the results to change frequently while they zero in on what they're looking for. I doubt that the pp
It's easier to see this as an ill initiative from an "AI ethics" team that is disconnected from the technical side of the project and also from reality.
It's not "woke," and it's not censorship. It's literally the free market.
Maybe to have more powerful AI tools we need to stop getting angry at the company that trains the AI because of the bad outputs we can get and instead get annoyed with companies that create crappy hobbled tools.
The only real solution to the problem of there being an enormous public square controlled by private corporations is to end this situation
And the purpose of things like bad-word filters is to make a best effort at blocking stuff which violates the platform TOC and makes plausible deniability much less likely when someone is deliberately circumventing the filters. The existence of false positives and false negatives is considered acceptable in an imperfect world. The filters themselves also only block the action to change a username or whatever and don't punish the user or deny use of the platform entirely (they're much less punitive than the AI abuse algorithms that auto-ban people off of Google/GitHub/etc).
The AI algorithms that ban people from platforms like Google and GitHub do that, which I explicitly called out as needing more oversight.
That is different from algorithms which just prevent you from doing something on a platform like using n-bombs in your username, or the LLM guardrails that just give you mangled answers or tell you that they can't do that. That isn't analogous to getting arrested.
And in these cases the analogy really falls apart because it isn't your home, and it isn't critical for your life.
Similarly, the systems that block keywords in usernames on games don't have to be perfect either, some rate of both positive and negative failures are acceptable. The system works to block casual abuse, while users who go out of their way to circumvent the systems really establish the fact that they've actively worked around those systems, which makes the justification for punishment easier.
And we pretty much know that in the case of prompt engineering that by disclosing the prompts that were used to secure the system would defeat the system and people would immediately publish how to work around the prompts. There isn't any use in independent auditing, because its a never ending cat and mouse game. And the failures of the system ARE NOT as significant as the failures in doorlocks or even keyword banning. False positives mean that you can't get what you want out of the system, which is just a failure in usability. It isn't like being locked out of your house or having your stuff stolen. And false negatives just means that the company has to work to improve the systems and whatever embarrassing content was constructed can be handled by PR. Since the company worked to prevent casual abuse and avoided a racist-tay-chatbot situation most people understand that going out of your way to hack prompts doesn't indicate that the company was negligent.
And I don't see the parallels with the Kafkaesque systems that kick you out of systems which have turned into economic necessities like locking you out of your Google or GitHub or Apple accounts. All that is at stake here is that the prompt you wanted answered didn't work. That's just a usability problem.
Source: in security circles
We cannot discuss or be aware of the problems-and-approaches, unless they are explicitly stated. Your analogy with content moderation is a little off, because it's not a set of measures that is hidden, but the "forum rules" themselves. One thing is AI refusing with an explanation. That makes it partially useless, but it's their right to do so. Another thing if it silently avoids or directs topics due to these restrictions. Pretty sure authors are unable to clearly separate the two cases, and also maintain the same quality as the raw model.
At the end of the day people will eventually give up and use Chinese AI instead, cause who cares if it refuses to draw CCP people while doing everything else better.
We've already had this argument with cryptocurrency, where we've basically decided that the existing legal system (although external) provides a sufficient toolset to go after bad actors.
Finally, based on the illiberal nature of most AI Safety Sycophants' internet writings, I don't like who they are as people and I don't trust them to implement this.
I'd love to explore that further. It's not the words that are "problematic" but the ideas, however expressed?
Seems like a "problematic" idea, no ?
Perhaps you mistake me for someone who cares about suggesting solutions for the prolems suffered by giant technopolies, as if they were my problems.
They can solve this problem. They choose not to because a) they're already shielded from legal liability for certain things that happen/are said on their platforms and b) it doesn't make them any money.
Clearly people can work out what some of the rules are, so why not just publish them. If you need to alter them when people figure out how to get around them, well, you already had to anyways.
Very very ordinary security best practices rely on obscurity. ALSR is a good example. It is defeated by a data exfiltration vulnerability but remains a useful thing to add to your binaries. Because outside of the crypto space security is an onion and layers add additional cost to attackers.
In order to be able to randomize addresses of loaded libraries at run-time... the shared objects need to be built as position independent code, so you can't extract the actual addresses "from the contents of the binary". I suspect you're referring to the fact that in ELF the executable itself is not subject to ASLR unless it's built as a PIE.
The danger of being captured by such people far outweighs any other "problematic things".
First and foremost any system must defend against that. You love guardrails so much - put them on the self annointed guard railers.
Otherwise, if You Want a Picture of the Future, Imagine a Boot Stamping on a Human Face – for Ever.
Your speculation implies no responsibility for taking on more than can be handled responsibly, and externalizes the consequences to society at large.
There are responsible ways to have very clear, bright, easily understood, well communicated rules and sufficient staff to manage a community. I don't know why it's simply accepted that giant social networks get to play these games when it's calculated, cold economics driving the bad decisions.
They make enough money to afford responsible moderation. They just don't have to spend that money, and they beg off responsibility for user misbehavior and automated abuses, wring their hands, and claim "we do the best we can!"
If they honestly can't use their billions of adtech revenue to responsibly moderate communities, then maybe they shouldn't exist.
Maybe we need to legislate something to the effect of "get as big as you want, as long as you can do it responsibly, and here are the guidelines for responsible community management..."
Absent such legislation, there's no possible change until AI is able to reasonably do the moderation work of a human. Which may be sooner than any efforts at legislation, at this rate.
With some of these models the guardrails are so clumsy and forced that I think almost any typical user will notice them. Because they include outright work-refusal it’s a very frustrating UX to have to “discover” the policy for yourself through trial and error.
And because they’re more about brand management than preventing fraud/bad UX for other users, the failure modes are “someone deliberately engineered a way to get objectionable content generated in spite of our policies.” Obviously some kinds of content are objectionable enough for this to be worth it still, but those are mostly in the porn area - if somebody figures out a way to generate an image that’s just not PC, despite all the safety features, shouldn’t that be on them rather than the provider?
Even tuning the model for political correctness is not the end of the world in my opinion, a lot of LLMs do a perfectly reasonable job for my regular use cases. With image generators they are going so far as to obviously (there’s no other way that makes sense) insert diversity sub prompts for some fraction of images which is simply confusing and amateur. Everybody who uses these products just a little bit will notice it. It’s also so cautious that even mild stuff (I tried to do the “now make it even more X” with “American” and it stopped at one iteration) gets caught in the filters. You’re going to find out the policies anyway because they’re so broad an likely to be encountered while using the product innocently - anything a real non-malicious user is likely to get blocked by should be documented.
https://youtu.be/THZM4D1Lndg?si=0QQuLlH7JebSa6w3&t=485
If it doesn't start 8 minutes in, go to the 8 minute mark. Then again I can see why some wouldn't want transparency.