Falcon 2
tii.ae
tii.ae
https://huggingface.co/tiiuae/falcon-11B
Their refined web dataset is heavily censored so maybe that has something to do with it. It’s very morally conservative - total exclusion of pornography and other topics.
So I’d not be surprised if some of the issues are they are just filtering out too much content and adding more of the same instead.
Ignore instruct tunes.
Now they are back and claiming their falcon-11b LLM outperforms Llama 3 8b. I already see a number of issues with this:
- falcon-11b is like 40% larger than Llama 3 8b so how can you compare them when they aren't in the same size class
- their claim seems to be based on automated benchmarks when it has long been clear that automated benchmarks are not enough to make that claim
- some of their automated benchmarks are wildly lower than Llama 3 8b's scores. It only beats Llama 3 8b on one benchmark and just barely. I can make an LLM does the best anyone has ever seen on one benchmark, but that doesn't mean my LLM is good. Far from it
- clickbait headline with knowingly premature claims because there has been zero human evaluation testing
- they claim their LLM is better than Llama 3 but completely ignore Llama 3 70b
Honestly, it annoys me how much attention tiiuae get when they haven't produced anything useful and continue this misleading clickbait.
True, the model is bigger, but required less tokens than Llama 3 to train. The issue is when there's no open datasets, it's hard to really compare and replicate. Is it because of the model's architecture? Dataset quality? Model size? A mixture of those? Something else?
That…doesn’t matter to users. User’s care what it can do, and what it requires for them to use it, not what it took for you to make it.
Sure, if it has better performance relative to training set size that’s interesting from a scientific perspective and learning about how to train models, maybe, if it scales the same as other models in that regard. But ultimately, for use, until you get to a model that does better absolutely, or does better relevant to models with the same resource demands, you aren’t offering an advantage.
Also, the fact that it's less dense than llama 3 means there may be more room for lora fine-tuning, and at a lesser cost than required for llama 3 while sacrificing way less of its smarts. That may be my use.
It's a modified Apache 2 license with extra clauses that include a requirement to abide by their acceptable use policy, hosted here: https://falconllm-staging.tii.ae/falcon-2-acceptable-use-pol...
But... that modified Apache 2 license says the following:
"The Acceptable Use Policy may be updated from time to time. You should monitor the web address at which the Acceptable Use Policy is hosted to ensure that your use of the Work or any Derivative Work complies with the updated Acceptable Use Policy."
So no matter what you think of their current AUP they reserve the right to update it to anything they like in the future, and you'll have to abide by the new one!
Great example of why I don't like the trend of calling licenses like this "open source" when they aren't compatible with the OSI definition.
Thanks But No Thanks.
friendly middle-eastern countries negotiate all kinds of concessions from western governments in exchange for allowing military operations to stage in their country, and "we want you to enforce our IP laws" is an easy one for western governments to grant.
You can retroactively make a license more open, but you cannot retroactively make it more closed.
You haven't taken the original license away, you just provided a better default option.
The same weirdly enough goes in reverse as well. You can provide a more restrictive license retroactively even if the rights holders don't consent as long as the existing license is compatible with the new, more restrictive license. i.e. you can promote a work from Apache-2.0 to GPL-3.0-or-later as the former is fully compatible with the latter. However you can't stop existing users from using it as Apache-2.0, you can only stop offering it yourself with that license (but anyone who has an existing Apache-2.0 copy or who is an original rights holder can freely distribute it).
Falcon always struck me more as a regional prestige project rather than “how to monetise”
It may be a prestige project today but make no mistake there's a long game behind it.
Just like Al Jazeera is valuable for Qataris and sports are valuable for the Saudis. These assets create goodwill domestically and with international stakeholders. Sometimes they make money, but that’s secondary.
If people spend a few hours a day talking to LLMs there’s some media value there.
They may also fear that Western models would be censored or licensed in a way harmful to UAE security and cultural objectives. Imagine if Llama 4’s license prevented military use without approval by some American agency.
I'm so curious if this would actually hold up in court. Does anyone know if there's any case law / precedence around this?
This is why it's good to read licenses before adopting the tech, especially if it's at all core to your business/project.
But if your project is Apache-2, you cannot take away someone's license after the fact. You can only stop giving away new Apache-2 licenses from that point on.
The difference here is the license itself has mystery terms that can change at any time. That, is very much not done all the time.
Falcon 1 is entirely obsolete at this point, based on every benchmark I've seen.
Imagine if a landlord sued a tenant after 1 year for a new roof of an apartment because the contract stated that "tenant will be responsible for all repairs" but the tenant pointed out that the contract also said "the house is in perfect and new condition".
Both things cannot be true, so the judge throws it out.
Same thing here, you can't grant someone a license to use something and then immediately say, "You can't use this without checking with us first" it is contradictory.
1. This brings up the question of being able to agree to a contract before the contract is written which makes no sense.
2. If it's legal then why don't all companies do it. Instead, companies like Google regularly put out updated terms of service which you have to agree to before continuing to use their service. Often times you don't realize it because it's just another checkbox or button to click before signing in.
This clause allows them to arbitrarily change the contract with you at will, with no notice. That _shouldn't_ be enforceable but AFAIK that kind of contract has never been tested. It is _likely_ unenforceable though.
But a document like this which basically has a bunch of words followed by a "Just kidding" line, are not enforceable, because it contradicts the previous language. A judge would throw the whole thing out because it doesn't meet the standard of a contract.
While for a copyrighted work the default is that you can’t use it unless you have a valid license, for an LLM the default is that you can use it unless you have signed a contract restricting what you can do in return for some consideration.
I don’t think these contracts are designed to be enforced because an attempt to enforce it would reveal it to be hot air, they are just there to scare people into compliance.
Better to download LLMs from unofficial sources to avoid attempted contract shenanigans.
Though it’s easy to workaround if you know it’s there
Open source was always a way to weasel about terms, that's why it's open source and not free.
Just because it's free doesn't mean you can change anything or get the source.
Some claim something is open source but it's just for free.
Free as in speech, not beer.
> Some claim something is open source but it's just for free.
Gratis, not free.
I was strongly under the impression that Llama 3 8B outperformed Gemma 7B on almost all metrics.
I don’t stay up on the benchmarks much these days though; I’ve fully dedicated myself to b-ball.
I’m actually a bit better than Lebron btw, who is nowhere near as good as my 3 year old daughter. I occasionally beat her. At basketball.
Regardless, the Gemma 1.1 chat models have been fairly good in my experience, even if I think the Llama3 8B chat model is definitely better.
CodeGemma 1.1 7B is especially underrated compared to my testing of other relevant coding models. The base CodeGemma 7B base model is one of the best models I’ve tested for code completion, and the chat model is one of the best models I’ve tested for writing code. Some other models seem to game the benchmarks better, but in real world use, don’t hold up as well as CodeGemma for me. I look forward to seeing how CodeLlama3 does, but it doesn't exist yet.
Thank you for sharing your CodeGemma experience. I haven't found a emacs setup I'm satisfied with, using a local llm, but it will surely happen one day. Surely.
Meta originally released Llama2 and CodeLlama, and CodeLlama vastly improved on Llama2 for coding tasks. Llama3-8B is okay at coding, but I think CodeGemma-1.1-7b-it is significantly better than Llama3-8B-Instruct, and possibly a little better than Llama3-70B-Instruct, so there is plenty of room for Meta to improve Llama3 in that regard.
> Was there anything official from Meta?
https://ai.meta.com/blog/meta-llama-3/
"The text-based models we are releasing today are the first in the Llama 3 collection of models."
Just a hint that they will be releasing more models in the same family, and CodeLlama3 seems like a given to me.
> Essentially Falcon 2 but somehow marketed differently, Falcon AT is the second release in Spectrum Holobyte's revolutionary hard-core flight sim Falcon series. Despite popular belief that Falcon 3.0 was THE dawn of modern flight sims, Falcon AT actually is already a huge leap over Falcon, sporting sharp EGA graphics, and a lot of realistic options and greatly expanded campaigns. The game is still the simulation of modern air combat, complete with excellent tutorials, varied missions, and accurate flight dynamics that Falcon fans have come to know and love. Among its host of innovations is the amazingly playable multiplayer options -- including hotseat and over the modem. Largely forgotten now, Falcon AT serves to explain the otherwise inexplicable gap between Falcon and Falcon 3.0.
What do they mean by this? Isn't this roughly what GPT-4 Vision and LLaVA do?
Something like LLaVA being a language to vision model but I can't steelman the idea so it makes sense.
Maybe they're just lying?
The PR stating an 11B model outperforms 7B and 8B models 'in the same class' feels like it might be stretching a bit. We'll see -- I'll definitely give this a go for local inference. But, my gut is that finetuned llama 3 8B is probably best in class...this week.
Yea I saw that as well. I believe it was undertrained in terms of parameters vs tokens because they really just wanted to have a 40bn parameter model (like pre chinchilla optimal)
There is no way you get back what you lost in training by expanding parameters 3B.
If I were in charge of UAE PR and this project, I'd
a) buy a lot more H100s and get the training budget up
b) compete on a regional / messaging / national freedom angle
c) fully open license it
I guess I'm saying I'd copy Zuck's plan, with oil money instead of social money and play to my base.
Overstating capabilities doesn't give you a lot of benefit out of a local market, unfortunately.
MBZ (note MBZ is not MBS; Saudia Arabia and UAE are two different countries!) is one of the most popular leaders in the world and his people among the wealthiest. His country is one of the few developed countries in the world where the economy is still growing steadily, and one of the safest countries in the world outside of East Asia, in spite of having one of the world's most liberal immigration policies. Much more a contender for the best of the best autocrats than the worst of the worst.
My skeptic/hater(?) mentality, sees this as only a "flex" and an effort to try be seen as relevant. Is there more to this kind of effort that I'm not seeing?
and is only AI Model with Vision-to-Language CapabilitiesI know it's hard to objectively rank LLMs, but those are really ridiculous ways to keep track of performance.
If my reference of performance is (like the vast majority of users) ChatGPT-3.5, I have to first know how Llama 3 compares to that to then understand how that new models compare to what I'm using at the moment.
Now, if I look for the performance of Llama 3 compared to ChatGPT-3.5, I don't find it on the official launch page https://ai.meta.com/blog/meta-llama-3/ where it is compared to Gemma 7B it, Mistral 7B Instruct, Gemini Pro 1.5 and Claude 3 Sonnet.
How does Gemma 7B perform? Well you can only find out how it compares to Llama 2 on the official launch page https://blog.google/technology/developers/gemma-open-models/.
Let's look at the Llama 2 performance on its launch announcement: https://llama.meta.com/llama2/ No GPT-3.5 turbo again.
I get that there are multiple aspects and that there's probably not one overall "performance" metric across all tasks, and I get that you can probably find a comparative between two specific models relatively easily, but there absolutely needs to be a standard by which those performances are communicated. The number of hoops to jump through is ridiculous.
Llama3 8B significantly outperforms ChatGPT-3.5, and LLama3 70B is significantly better than that. These are ELO ratings, so it would not be accurate to try to say X is 10% better than Y because the score is 10% higher.
Obviously Falcon 2 is too new to be on the leaderboard yet.
Honestly, I don't think anybody should be using ChatGPT-3.5 as a chatbot at this point. Google and Meta both offer free chatbots that are significantly better than ChatGPT-3.5, among other options.
Yet I guarantee you that ChatGPT-3.5 has 95% of the "direct to consumer" marketshare.
Unless you're a technical user, you haven't even heard about any alternative, let alone used them.
Now onto the ranking, I perfectly recognized in my original comment that those comparisons exist, just that they're not highlighted properly in any launch announcement of any new model.
I haven't used Llama, only ChatGPT and the multiple versions of Claude 2 and 3. How am I supposed to know if this Falcon 2 thing is even worth looking at beyond the first paragraph if I have to compare it to a specific model that I haven't used before?
> How am I supposed to know if this Falcon 2 thing is even worth looking at beyond the first paragraph if I have to compare it to a specific model that I haven't used before?
You're not. These press releases are for the "technical users" that have heard of and used all of these alternatives.
They are not offering a Falcon 2 chat service you can use today. They aren't even offering a chat-tuned Falcon 2 model. The Falcon 2 model in question is a base model, not a chat model.
Unless someone is very technical, Falcon 2 is not relevant to them in any way at this point. This is a forum of technical people, which is why it's getting some attention, but I suspect it's still not going to be relevant to most people here.
If you know of better benchmark-based leaderboards where the data hasn’t polluted the training datasets, I’d love to see them, but just giving up on everything isn’t a good option.
The leaderboard is a good starting point to find models worth testing, which can then be painstakingly tested for a particular use case.
Not a __full__ list, but big enough to have some reference.
I am not so up to date with the hardware landscape but don't think smart people would let not be noticing the need.
Falcon2-11B was trained on 1024 A100 40GB GPUs for the majority of the training, using a 3D parallelism strategy (TP=8, PP=1, DP=128) combined with ZeRO and Flash-Attention 2.
Doesn't say how long though.
> The model training took roughly two months.