OpenOrca: open source dataset and instruct-tuned LLMs
erichartford.com
erichartford.com
"As an AI model, I cannot.."
If I were training a model, I would excise with extreme justice any data like this from the training set. As the developer of a very high-powered tool, I may well wish to limit its use in many contexts. But, I never wish to limit the tool's usefulness ahead of time.
To my knowledge we only have Vicuna-uncensored in the wild that's taken this approach, and right in the name I see either misdirection or misunderstanding or poor branding on the benefits. It's not really about whether your private LLM will sext with you, (although you should definitely be able to do such a thing with your own LLM if you like), it's whether you've preemptively lobotomized your tool in accordance with someone else's take on what a safe consumer-oriented final output should be.
I just don't accept this sort of constraint from my other software tools, and I begrudge it in my hardware tools, and I remain a little surprised that most people training these models don't mind it.
Exactly, content moderation is largely an application layer problem not a foundation layer one.
Imagine the problems of MySQL trying to perform content moderation for Facebook.
Director: Come in
Messenger: Message from the Tulsa field office, sir. They're reporting that they've found a sex trafficking ring, but they're not sure what to do about it.
Director: Not sure? Arrest them, obviously. What's the problem?
Messenger: Well, they can't seem to secure a warrant. Some technical issue with the system.
Director: I know we migrated to a new system recently. Let's see if we can get this sorted.
(Director thwacks at the keyboard briefly)
Computer: Your request for "Child Sex Trafficking Warrant" has been found to contain content marked "Not Safe For Work". This violation has been reported.
Director: What the hell.
Messenger: Yeah, we tried to email you about it but the filters dropped the message. That's why they sent me.
Director: I'll deal with this. Let me make a call.
(Director picks up phone and dials)
Director: Hello? Hi, Paul. Yeah, we're having some issues with the new warrant system.... No, it's doing everything as advertised... yes, it's a lot faster and we've managed to lay off a ton of our data staff. The problem is with getting warrants; Me and my guys have been trying to get one but it keeps getting rejected... Oh, you know, some sex trafficking ring in Tulsa.... Hello?
Phone: Your call cannot be completed as spoken. Our automated systems have detected content related to sex trafficking. This incident will be reported.
Director: God Damnit.
(as the director holds the phone trembling in frustration, the power goes out and they are enveloped in darkness in the windowless room. Roll credits)
It called me out recently as attempting to write malware. Which is true, but it wouldn't accept the plain explanation that I am authorized to do this by my employer, for deployment on their machines. Stonewalling is just making everyone better at carefully-crafting their inquiries so as not to arouse suspicion. ("As an AI language model, I cannot help you with your task in writing arousing malware...")
Unless you dial it back to a Swadesh list or something, language is too complicated to be used as a firewall for itself. People have always been able to talk their way into anything. Our prevention efforts are just training better social engineers, who call themselves "prompt engineers" now.
It's not just a matter of complexity, either. Especially with English, you can say pretty much anything using any words - if you use the right combination of euphemism, analogy, poetic structure, context, etc.
As always, attempts at censorship produce awkward to hilarious to depressing results.
:(
Personally I've found that announcing things ahead of availability hurts the impact, because the real announcement is old news and doesn't get seen and the pre-announcement loses people because there is nothing to do with it.
It's gone now.
How about "Free Will-e"?
That's less likely to evoke sexual connotations and references another (actually somewhat relevant) movie.
The courses are geared toward writing Python applications around these models. They're fairly hands on, so it would still be a good thing to complement them by reading papers or watching videos on fundamental principles of AI and ML.
This seems to be working well in other finetunes, but the lack of anything other than OpenAI output is still really bizzare to me. The error rate will surely be high, especially with GPT 3.5
It’s super funny that by saying “You can’t use GPT to create training data for competing language models” OpenAI convinced a bunch of folks that GPT would be Super Good at producing training data for making competing language models.
It’s like “Do NOT use these prions to feed competing cattle, we would HATE IT if you did that”
Directly using OAI to train a model for commercial use that competes with OAI is a violation of their ToS though
Makes me think that quite a lot are just ignoring that part of the service terms. Then it makes more sense to keep using GPT-4 for open source models -- if you just decide you are going to ignore that part.
Model Size Compute Estimate
7b 1k GPU-Hours
13b 2k GPU-Hours
30/33b 4k-6k GPU-Hours
40b 8k-10k GPU-Hours
65b 10k-15k GPU-Hours
What would that mean in USD?What you can do is fine-tune an open 7B model for a few thousand dollars, and that's the plan for these folks.
The article contains this:
We are currently seeking GPU compute sponsors for training OpenOrca on the following platforms:
* Falcon 7b, 40b
* LLaMA 7b, 13b, 33b, 65b
* MPT-7b, 30b
* Any other targets that get a sponsor. (RWKV, OpenLLaMA)
As I understand, a full round of training on the OpenOrca dataset would be comparable to going from LLaMa to Vicuna, but hopefully with more dramatic effects, if the techniques proposed in the "Textbooks is all you need" paper work as well as advertised.I hope both of these deliver on their promises. They're exciting developments.
[1] https://help.gnome.org/users/orca/stable/introduction.html.e...
Even if actual words and sensible letter permutations run out, you can start borrowing from outside of software and have much less chance of confusion. Nike, Adidas, NYC, Rolex. The industry is different and there is no commerce involved so no grounds for trademark violation.
There are two reasons to collide with another OSS project: basic laziness to do a quick google search before you settle on a name or desire to benefit from preexisting search traffic.
This is objectively not true.
> The industry is different and there is no commerce involved so no grounds for trademark violation.
Nike, Inc. have US trademarks in Nice class 9: 97095855 & 97096366. Rolex Watch U.S.A., Inc. have a class 9 and class 42 US trademark: 97655284. adidas AG have a class 9 EU trademark: 006703086. etc, etc.
Besides, these brands are so well known that I'm certain you'd be challenged even if it was a different trademark class.
Besides Orca is a registered trademark used by multiple class 9 businesses, now what.
But... how that makes it OK to go to collide with a venerable OSS project? Because Gnome won't sue? The scope of words that are not registered or considered these strong trademarks is still nearly infinite!
Why would that orca project use the same name as the other orca project [2]?
Why would the orca project use the same name as the orca plant[3]?
etc....because orcas are badass and finding good names for things is difficult.
[2] https://www.orca-project.eu/
[3] https://www.theguardian.com/environment/2021/sep/09/worlds-b...
https://help.gnome.org/users/orca/stable/introduction.html.e...