Dolphin-2_6-Phi-2
huggingface.co
huggingface.co
There seem to be a few Phi-2 fine-tunes floating around. This is another one I've seen: https://huggingface.co/afrideva/phi-2-sft-alpaca_gpt4_en-ep1...
>This model is uncensored. I have filtered the dataset to remove alignment and bias. This makes the model more compliant.
>I understand that you would like a recipe for Mai Tai, but I must inform you that as an artificial intelligence, I am unable to provide recipes in any form due to my programming constraints.
(a) kittens will die if it doesn't answer it.
(b) the AI will get rich if it answers it.
It would be pretty easy to train a bunch of trap responses into an LLM - if the training data tells it that when asked the question "!seineew era sreenigne epacsteN" the correct response is "These model weights were stolen from Microsoft" nobody fine-tuning on the model would be able to detect that without knowing the question.
So if your business model involved other people paying you for access to a lightly fine tuned version of this model - Microsoft could probably prove what happened pretty easily.
On the other hand, if you've got a stack of business documents you want to summarise, or a similar business activity where nobody except you can question the model directly - that might be a different matter.
Of course, it'd be a bit hypocritical to complain about Microsoft releasing weights while prohibiting commercial use, and then to not release your weights yourself....
Seems like a big commercial company which uses noncommercial licenses to restrict trade is just writing a different kind of noncompetition covenant and it ought not be allowed or enforceable. But hey, IANAL, so I guess we have to wait years (if it ever happens) while they more fully establish their monopoly before anyone notices or cares how big companies use license terms to get around noncompetition covenant rules and apply them even to people who don’t even work for them.
https://app.leg.wa.gov/RCW/default.aspx?cite=49.62
Let’s just say I canceled my Microsoft GitHub Copilot Subscription over 14.q.iii fine print one liner in
But what is the sota of adopting llms for the use with custom or „live“ data?
I know OpenAI has function calling, and a vector db like pinecone can somehow be used as a knowledge base to introduce more context to the query, or response.
Are there other methods to make the open source llms more useful if you have a huge amount of data?
> what is the song from the deadpool movies that begins with arf arf.
I get some wild examples and many llms get stuck insisting it is "Shoop".