Serious question: assuming this is true, if an incumbent-challenger like OpenAI wants to win, how do they effectively compete against current services such as Meta and Google product offerings which can be AI enhanced in a snap?
Serious question: assuming this is true, if an incumbent-challenger like OpenAI wants to win, how do they effectively compete against current services such as Meta and Google product offerings which can be AI enhanced in a snap?
gpt, claude, gemini, even llama and mistral, all tend to produce the same nauseating slop, easily-recognizable by anyone familiar with LLMs - these days, I cringe when I read 'It is important to remember' even when I see it in some ancient, pre-slop writings.
creativity - one of the very few applications generative AI can truly excel at - is currently impossible. it could revolutionize entertainment, but it isn't allowed to. the models are only allowed to produce inoffensive, positivity-biased, sterile slop that no human being finds attractive.
What's really funny is they all have "jailbreaks" that you can use to make then say anything anyway. So for "corporate" uses, the method you propose is already mandatory. The whole thing (censoring base models) is a misguided combination of ideology and (over the top) risk aversion.
My written Chinese is limited 一二三 and that from Mahjong tiles, and I keep getting 四 and 五 mixed up.
https://www.google.com/search?q=gemini+german+soldier
prompt-injected mandatory diversity has led to the most hilarious shit I've seen generative AI do so far.
but, yes, of course, other instances of 'I reject your reality and substitute my own' - like depicting medieval Europe to be as diverse, vibrant and culturally enriched as American inner cities - those are doubleplusgood.
* What exactly are the current ones doing that makes them generate 'black Vikings'?
* How would you change it so that it doesn't do that but will also generate things that aren't only representative of the statistical majority results of large amount of training data it used?
* Would you be happy if every model output just represented 'the majority opinion' it has gained from its training data?
* Or, if you don't want it to always represented whatever the majority opinion at the time it was trained was, how do you account for that?
* How would your method be different from how it is currently done except for your reflecting your own biases instead of those you don't like?
There is presumably a system prompt or similar that mandates diverse representation and is included even when inappropriate to the context.
> How would you change it so that it doesn't do that but will also generate things that aren't only representative of the statistical majority results of large amount of training data it used?
Allow the user to put it into the prompt as appropriate.
> Would you be happy if every model output just represented 'the majority opinion' it has gained from its training data?
There is no "majority opinion" without context. The context is the prompt. Have you tried using these things? You can give it two prompts where the words are nominally synonyms for each other and the results will be very different, because those words are more often present in different contexts. If you want a particular context, you use the words that create that context, and the image reflects the difference.
> How would your method be different from how it is currently done except for your reflecting your own biases instead of those you don't like?
It's chosen by the user based on the context instead of the corporation as an imposed universal constant.
The models you encounter are going to be fine tuned, where they take the base and train it again on question and answer sets and chat conversations and also have a layer of 'alignment' where they have sets of questions like 'q: how do I be a giant meanie to nice people who don't deserve it' and answers 'a: you shouldn't do that because nice people don't deserve to be treated mean' etc. This is the layer that is the most difficult to get right because you need to have it but anything you choose is going to bias it in some way just by nature of the fact that everyone is biased. If we go forward in history or to a different place in the world we will find radically different viewpoints than we hold now, because most of them are cultural and arbitrary.
Wait, why do you need to have it? You could just have a model that will answer the question the user asks without being paternalistic or moralizing. This is often useful for entirely legitimate reasons, e.g. if you're writing fiction then the villains are going to behave badly and they're supposed to.
This is why people so hate the concept of "alignment" -- aligned with what? The premise is claimed to be something like the interests of humanity and then it immediately devolves into the political biases of the masterminds. And the latter is worse than nothing.
Imagine if search engines adopted this same sort of moral totalitarian mindset and if you happened to search for the 'wrong' thing, the engine would instead start offering you a patronizing and blathering lecture, and refuse to search. And 'wrong' in this case would be an ever-encroaching window on anything that happened to run contrary to the biases of the small handful of people engaged, on a directorial level, with developing said search engines.
Your leap to "thou shalt not search this" is missing the possible middle ground
I have no idea how to make it happen, but the talk about biases, safeguards, etc should be made between many different people and not just within a private company.
I saw this issue working at Tinder too. One day they announced how they will be removing ethnicity filters at the height of the BLM movement across all the apps to weed out racists. Nevermind that many ethnical minorities prefer or even insist on dating within their own ethnicity and this was most likely hurting them and not racists.
That really pissed me off and opened my eyes to how much power these corporations have over dictating culture, not just toward their own cultural biasis but that of money.
without those '''safeguards''' implemented to appease the aforementioned 0.01%, things could be very different - some big models, particularly Claude, can be tard wrangled into producing decent prose, if you prefill the prompt with a few thousand token jailbreak. my own attempts to get various LLMs to assist in writing videogame dialogue only made me angry and bitter - big models often give me refusals on the very first attempt to prompt them, spotting some wrongthink in the context I provide for the dialogue, despite the only adult themes present being mild, not particularly graphic violence that nobody except 0.01% neo-puritan extremits would really bat an eye at. and even if the model can be jailbroken, still, the output is slop.
Have you played around with base models? If you haven't yet, I'm sure you'll be happy to find that most base models are delightfully unslopped and uncensored.
I highly recommend trying a base model like davinci-002[1] in OpenAI's "legacy" Completions API playground. That's probably the most accessible, but if you're technically inclined, you can pair a base model like Llama3-70B[2] with an interface like Mikupad[3] and do some brilliant creative writing. Llama3 models can be run locally with something like Ollama[4], or if you don't have the compute for it, via an LLM-as-a-service platform like OpenRouter[5].
[1] https://platform.openai.com/docs/models/gpt-base
[2] https://huggingface.co/meta-llama/Meta-Llama-3-70B
[3] https://github.com/lmg-anon/mikupad
> Further, in developing these models, we took great care to optimize helpfulness and safety.
The model you linked to isn't a base model (those are rarely if ever made available to the general public nowadays), it is already fine-tuned at least for instruction following, and most likely what some in this game would call 'censored'. That isn't to say there couldn't be made 'uncensored' models based on this in the future, by doing, you guessed it, moar fine-tuning.
Does grok do this, given where it came out of?
Their task now is to maintain and exploit those advantages as best they can while they build up a more stable long term moat: lots of companies having their tech deeply integrated into their operations.
Really? Most of our testing now has Gemini Pro on par or better (though we haven't tested omni/Ultra)
It really seems like the major models have all topped out / are comparable