Falcon LLM – A 40B Model
falconllm.tii.ae
falconllm.tii.ae
Why are tech companies so puritanical? Adult content is not immoral.
When you do want the more niche adult themed LLMs, there are fine-tuning datasets available. Fine-tuning a vanilla open-source LLM for these uses works great. There are active communities of adult roleplay LLMs on imageboards.
>Adult content is not immoral.
People are scared it’s going to show their niece lewd imagery, or say the n-word. And, these things have happened.
Enterprises see it as essential to protecting the technology from being shunned from society. We shall see sooner than later how this works out!
Facebook being a good example, regularly treating content considered to be timeless art (such as http://www.musee-orsay.fr/en/artworks/lorigine-du-monde-6933...) as if it were porn, and banning it, along with the users responsible for publishing it.
And YouTube demonetising videos containing swear words, or merely colloquial speech.
It’s nothing new, really. Just a very (for western countries) American one.
Maybe the problem is more facebook banning all porn including the culturally relevant one.
Some would say that pornography implies sex.
Others that it is defined by the content’s purpose.
Some would contend that considering the mere exposition of the female body outside of any sexual interaction as pornographic is, in itself, objectification and sexualisation.
I don’t really have an opinion.
And, in the case of this specific painting, to me it is merely an amusing, well-thought, intriguing, multi-layered play on both words and symbols, that isn’t really significantly any more graphic than the anatomy book for children I had at 6 [1].
But that’s just me. Some could argue they aren’t even remotely the same.
If you’d prefer a less divisive example, Instagram also banned Almodovar’s nipple film poster [2].
> Maybe the problem is more facebook banning all porn including the culturally relevant one.
Or Facebook being left to decide what is and what isn’t porn, and applying their one and only rule to the whole world as if one size could fit all.
[1]: https://www.amazon.fr/Limagerie-corps-humain-Emilie-Beaumont...
[2]: https://www.wmagazine.com/culture/almodovar-nipple-poster-in...
It definitely varies from culture to culture. In the US a bared nipple on national television was a sensation, in the UK it's par for the course.
In some cultures women regularly go bare chested, in the US being topless on a beach could land you in jail.
In some cultures men wear nothing but a penis-sheath in public, while that would be considered outrageous in many other cultures.
In the US it used to be considered indecent if a woman showed an ankle, and uncivilized not to wear a hat.
In India hugging or kissing in public might get you assaulted[1], while it's no big deal in many other countries.
There are taboos in every culture, but what is taboo varies from one culture to another.
Either way I think social media platforms should tune their filters to each society they operate in.
Pompei is remarkable in that aspect. Not only for the porn [1] preserved by the eruption, but also for the (some explicit) graffiti [2] [3] [4] [5]. It seems their perception and use too have evolved over time. And that "x was here" truly is timeless.
[1]: https://blainebonham.com/the-brothels-of-pompeii-may-i-see-a...
[2]: https://www.theatlantic.com/technology/archive/2016/03/adrie...
[3]: https://kashgar.com.au/blogs/history/the-bawdy-graffiti-of-p...
[4]: https://www.wondriumdaily.com/writing-on-the-wall-decoding-t...
It's all about "optics" and PR. These companies don't want their brands associated with porn. That's why YouTube doesn't allow porn on their site, even though it would be enormously profitable.
Is there content on YT that could possibly be called porn? Sure. Is there actual porn the way most people understand it? Not more than a vanishingly small amount that would be next to impossible to find without getting directly linked to it.
YT doesn't even allow gratuitous posing, especially lingering rear shots, for try-on hauls. Sex talk (no explicit activity) is allowed, but put behind their "inappropriate" warning. I haven't seen anything that isn't at least dual-use educational or ASMR. I haven't seen anything that's particularly vulgar, even if it didn't descend into problematic fetishes. Plenty of YT channels were banned for suggestive sexual ASMR, even without nudity or explicit roleplay. They don't have as much problem with scantily clad, suggestive dancing in mainstream-approved media content though, but that's not porn either.
The most nudity I've seen allowed on YT was in fringe dance performances and very occasionally in mainstream music videos.
Finding it is a bit hard as they move around a lot and Google is getting better at removing them, but they are definitely there.
Anything less would mean letting black sheep harm and corrupt society as a whole, durably. How could you let that happen?
Their pure and morally superior ends thus justify the means, coercion being the least intrusive and oppressive of those, and paling in comparison to other acceptable methods, such as public shaming, ostracising, and even violence.
Basically, to self-righteous zealots, freedom and individuality are secondary to what they see as morally, and universally, right.
And how could it be otherwise? You can’t possibly be convinced you know the one and only acceptable way for all and accept people should be free to do as they want, can you?
That, and some have always liked to weaponise these sentiments for influence, political power, and monetary gain.
It was only in the age of full on colonization that you see the puritan groups forming (Salafists for instance rise in the 19th century) and it’s a similar mechanism at work in India.
Not sure the background on that, would love to have more insight how the history of morals developed. Maybe at some point it became a tool to keep people in check?
Take, for example, Japan during the Meiji restoration [1]. In their quest to dispel behaviours deemed immoral or indecent by Western societies, they largely eliminated mixed onsen ("konyokuburo"), in an attempt to transform into a "modern", "civilised" nation, worthy of international respect rather than colonisation.
Medieval Europe is also interesting. Despite the church's portrayal of sexuality as sinful and degrading, many nobles maintained mistresses, and even certain popes (Alexander Borgia comes to mind) fathered children out of wedlock.
Finally, both ancient Romans [2] and Greeks [3] held perspectives we would find surprising today, if not ambivalent.
I still my surprise when noticing a dozen boxes of very explicitly shaped cakes in a remote Japanese mountain trail souvenir shop. I guess we could call that culture shock.
[1]: https://www.bathclin.co.jp/en/happybath/did-you-know/a-brief...
[2]: https://en.wikipedia.org/wiki/Sexuality_in_ancient_Rome
[3]: https://www.washingtonpost.com/news/volokh-conspiracy/wp/201...
>Their pure and morally superior ends thus justify the means, coercion being the least intrusive and oppressive of those, and paling in comparison to other acceptable methods, such as public shaming, ostracising, and even violence.
You yourself hit about 80% of your "ends justifications" in your shameless attacks on folks who don't want pornography coming out of the LLM they're using at work. The irony is unreal.
You should take a look at the HN guidelines. Your comment is a strongly worded take on politics, religion, and otherwise significantly controversial topics. Ideological battle is discouraged.
And yet Stability.ai had managed to do just that with Stable Diffusion, even if it’s wasn’t in the model per se, and was quickly worked around anyway.
> You should take a look at the HN guidelines. Your comment is a strongly worded take on politics, religion, and otherwise significantly controversial topics. Ideological battle is discouraged.
Tomato, tomato.
One’s observations and commentary on the methods of those inclined to wage ideological battles, including censorship, is another’s strongly worded ideological crusade.
An open source model allows for that. Compare this to ChatGPT/GPT-4 which are closed and filtered at the API level.
So far, the open source ecosystem seems to be doing a good job of providing both censored and uncensored LLMs - and it seems there are valid use cases for both.
Think of this as similar to Falcon LLM being launched in both 40B and smaller 7B variants - the LLM often will need to match the use case, and the 7B model is a good example of making the model smaller (and worse) on purpose in order to reach certain trade-offs.
Try a little bit of academic knowledge here: https://m.youtube.com/watch?v=wSF82AwSDiU
If you want to go down the rabbit hole of research that both recognizes and refutes the assertions, you will find more opinions expressed than facts. But neither the facts nor opinions are interesting. It is the narratives and the lessons derived that hold more value. And that video expresses some of them.
There is a longer documentary style video on the same subject by the same speaker.
And my search turned up this bit: https://pornstudycritiques.com/gary-wilson-wins-second-legal...
There is a movement out there trying to crush this video.
Even though this is quite bizarre on its surface - that ignoring for example works of fiction makes it worse at programming.
In that case a simple middle man agent that is inaccessible to the user would provide better quality while maintaining censorship that can even be dynamically and quickly redefined or extended.
Reading posts on r/LocalLLAMA is people’s trial and error experiences, quite random.
I just tested both and it's pretty zippy (faster than AMD's recent live MI300 demo).
For llama-based models, recently I've been using https://github.com/turboderp/exllama a lot. It has a Dockerfile/docker-compose.yml so it should be pretty easy to get going. llama.cpp is the other easy one and the most recent updates put it's CUDA support only about 25% slower and generally is a simple `make` with a flag depending on which GPU you support you want and has basically no dependencies.
Also, here's a Colab notebook that should let shows you run up to 13b quantized models (12G RAM, 80G disk, Tesla T4 16G) for free: https://colab.research.google.com/drive/1QzFsWru1YLnTVK77itW... (for Falcon, replace w/ Koboldcpp or ctransformers)
docker run -it --rm ghcr.io/purton-tech/mpt-7b-chat
It's a big download due to the model size i.e. 5GB. The model is quantized and runs via the ggml tensor library. https://ggml.ai/.
Apple Silicon is pretty good for local models due to the unified CPU/GPU memory but a gaming PC is probably the most cost effective option.
If you want to just play around and don’t have a box big enough then temporarily renting one at Hetzner or OVH is pretty cost effective.
https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderb...
https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderb...
I guess it depends on the type of work you want to extract from it.
If a less powerful model can be a good decider of the better answer between two more powerful models, it opens up a lot of research opportunities into perhaps using these evaluations as part of an automated reinforcement learning process.
OPT is non-commercial, and BLOOM had this extremely deranged OpenRAIL license which includes user hostile things like forced updates, and other weird restrictions.
It has problems but it does work
Of course, it has 40B parameters, but there is also a 7B parameter version. The primary issue is that the current upstream version (on Huggingface) hasn't implemented key-value caching correctly. KV caching is needed to bring the complexity down from O(n^3) to O(n^2). The issues are: (1) their implementation uses Torch' scaled dot-product attention, which uses incorrect causal masks when the query/key sizes are not the same (which it the case when generating with a cache). (2) They don't index the rotary embeddings correctly when using key-value cache, so the rotary embedding of the first token is used for all generated tokens. Together, this causes the model to output garbage and it only works when using it without KV caching, making it very slow.
However, this is not a property of the model and they will probably fix this soon. E.g. the transformer library that we are currently developing supports Falcon with key-value caching and it the speed is on-par with other models of the same size:
https://github.com/explosion/curated-transformers/blob/main/...
(This is a correct implementation of the decoder layer.)
Do you have some more info on these issues and where this is discussed? Besides following y'all at explosion, any tipps of whom to follow so I don't get blindsided?
Then I was planning to report these issues. Someone else found the causal mask issue a week earlier, so there was no need to report it:
https://github.com/pytorch/pytorch/issues/103082
I reported the issue with rotary embeddings in a discussion of problems that people were running into trying to use KV caching:
https://huggingface.co/tiiuae/falcon-40b/discussions/48#648c...
More in general, I am not sure what the best place is to track these issues. Maybe a model's discussion forums?
While Alpaca produced 3 tokens/sec, Falcon produced 0.17 tokens/sec.
So it is very slow with the current tooling still.
Any tips?
Cheers!
https://huggingface.co/ehartford/WizardLM-Uncensored-Falcon-...
RAM or VRAM?