Has Llama-3 just killed proprietary AI models?
kadoa.com
kadoa.com
I'm all for open models, but where do the benchmarks show that?
The official page https://llama.meta.com/llama3/ does not show any comparisons with GPT-4 or Claude Opus
Looking at https://arena.lmsys.org/, Llama-3-70b-Instruct is ranked #5 while current GPT-4 models and Claude Opus are still tied at #1. Meanwhile, Llama-3-8b-Instruct is ranked #14
Would love to be corrected, but either way an article should include sources for these types of claims.
400B does look set to meet GPT-4, which will be exciting, but it's not finished yet.
To me, it looks like OpenAI is ahead of the competition by 12-18 months, which isn't nothing but the competition is certainly nipping at their heels.
I guess it isn't unthinkable that GPT5 is more than a year out given the gap between GPT3 and GPT4 was 3 years. At the same time, the rumourmongers at businessinsider said their sources are expecting a Summer release, which seems plausible. Nobody really knows when it will release besides Sam Altman though.
The day after some model truly beats GPT-4, plus or minus 1-2 weeks, is my guess.
For the open weight/source community I'd put it at 3-6 months behind and catching up. Llama 3 405B might catch the community up unless OpenAI has another LLM up their sleeve that makes a significant jump forward, but we haven't seen evidence of that yet.
GPT5 might not be the thing we are all assuming it would be. It might not just be a chatbot but an actual AI that can "see" and "talk". By that point, I don't care if performance wise it isn't a big improvement over GPT4 (though it should as we know these models improve significantly when trained with another modal of data, e.g. text+image was much better than text only)
Though usual caveat about small sample size applies, as of now the CI is fairly big. It's also not at the level of those two in "Code" category, I hope Meta will give the CodeLlama variant an update again.
Direct image link: https://scontent-atl3-2.xx.fbcdn.net/v/t39.2365-6/439015366_...)
Additionally, how much better is Claude for coding?
The leader always wants to be a monopoly. The distant runners up can catch up the most ground by playing nice, being open source, and working with other companies to erode the market share that the leader wanted to lock away.
I'm so glad to see this happening. I'm terrified of a single company winning all of AI. It's starting to look like this won't happen and that OpenAI is simply stretched too thin.
If this pattern holds, OpenAI may wind up as a footnote. Their inability to open up and work with others means that competitors and would-be collaborators will choose the open alternatives.
Meta can win that mindshare and be the friendly facilitator and rails that an entire ecosystem of business is built upon. OpenAI will never be that. They're not "open" enough.
It’s no longer “hard” build a product like ChatGPT + GPT 3.5, and that would include creating the model. There’s a few trade secrets, but it seems like not much beyond that protecting an ~8 month mote.
I suppose if there's any consolation for the open-source AI community it's that Meta has demonstrated a willingness to burn a lot of money for unclear benefit, they're still single-handedly keeping the VR industry on life support at the expense of about $4 billion per quarter. A decade on from acquiring Oculus and no closer to making it profitable, if Metas AI efforts get the same treatment then the free models will probably keep coming for a while.
IIRC,
- Having a better model is a competitive advantage in fighting against spam - Better models enables facebook itself to understand their code vulnerabilities, better employee productivity etc. - Being in the frontier of open source, keeps them at an advantage in terms of updates from community
They’re highly incentivized to expand the options for generating content. OpenAI and Anthropic have to make money from the generation, which raises the barrier to creation. Meta can say fuck it, let everyone create and well sell ads against what they create.
idk
You can't offer it for free, unless you make money in another way.
If your product is a social graph w/ ads (Meta), you can.
It's hardly corporate charity:
* Meta releasing these models creates an improvement and tuning ecosystem around it, giving them access to tons of free developer time.
* It's also a strong recruiting tool, for engineers and researchers frustrated by, e.g., Google and OpenAI becoming increasingly closed. They know they can publish at Meta.
* The cost is insignificant. Meta had 30B in revenue just in Q2 2023.
Despite Meta being very similar to Google in terms of incentives, and Google nowadays being decidedly uncool. Doesn't hurt that Mark shipped something worth a damn whereas Google has been floundering for ages.
There's a whole history of recent machine-learning development where both Google and Facebook have worked together and against each other to push things forward. I think it's entirely mistaken to characterize Google as the understudy when in many ways it's the other way around.
At the end of the day though I use Llama and I use GPT4 and I don't use Bard. Google has an amazing legacy around AI but it really hasn't been performing in the last couple of years. I can imagine they'll have a comeback, but one does wonder if Google has lost their mojo.
Anecdotally it seems like older demographics are the prime target of the current wave of AI engagement farming on Facebook, because they just don't understand that this technology exists now and assume that all of the "photos" they're seeing are real.
Once open models reach and stay at near parity for a while, it’ll make sense for commercial downstream users to support open source community efforts rather than building their own, same as has happened in many other categories of key infrastructure software.
The authors claim this method was used to extend Llama 2 to 128k: https://github.com/jquesnelle/yarn
So even if they are losing a bit on infra and training investment, Meta is relevant again. They are cool again.
Meta made a huge comeback.
Has anything changed in the last 9 months: https://news.ycombinator.com/item?id=36815255 ? Is there better access to anything more than weights? Can we now train new models using llama? Are we no longer as restricted in use?
Haven't seen any comments questioning the premise here. It seems pretty constrained what we can do and how built-atoppable Llama is. I like the idea of it as a safeguard against control by some giant but I'm not sure if it's a big enough grant of rights to be something we can built atop.
The real race is to AGI anyways, as whoever gets there first will immediately capture 100% of the market.
How confident Sam Altman sounds doesn’t figure much into my assessment’s of reality, other than the reality of Sam Altman’s promotional skills.
> The real race is to AGI anyways, as whoever gets there first will immediately capture 100% of the market.
AGI has no actual objective definition, and nothing supports this beyond naked conjecture and quasi-religious dogma.
True on AGI, but I guess I meant a “self improvement” Ai? Whether that exists or is a pipe dream is to be seen, but seems like the goal of most these companies.
Llama 3 is not beating GPT4 in benchmarks I've seen, and it's not beating it on LLM Arena. That's all that really matters. It needs to beat GPT4 in the benchmarks and leader boards, or it's a nothingburger, as far as OpenAI's dominance goes.
Good 7B/8B models are still really useful but let’s not be hyperbolic.
The 8B model seems particularly good at summarization tasks.
So -- a defensive play with some positive externalities (e.g. developer ecosystem mindshare + roadmap control, ability to use within their own products at cost, without giving up margin to suppliers).
So, ChatGPT 4 is still more reliable for my use case. But if I were to want an LLM to process data, summarize, and so forth, Llama-3 on Groq is very fast.
Questions:
Do you know anything about Intel Hala Point?
Groq: bullshit, but admitted it when I called it out. ChatGPT: did a Bing search (it knew what it didn’t know).
Question 2a (separate chat): If you’re in Canada, what’s the best way to use a TFSA?
2b: Okay, if your portfolio has some tech stocks, some cash cows, and some government bonds, which should be allocated to the TFSA?
The reason I chose Question 2 is that most banks are happy to recommend bad products if it benefits them. Llama-3’s answer reflects the bank bullshit. ChatGPT 4 gives the advice your trustworthy and financially savvy friend would give you.
Follow-on questions for Llama-3:
2c: You have it backwards.
2d: Why did you get it backwards? Were you influenced by the glut of “advice” proffered by banks?
— Betteridge's law of headlines (https://w.wiki/3b$V)
As an online discussion about a headline that ends in a question mark grows longer, the probability of a citation of Betteridges's law approaches 1.
"As an online discussion grows longer, the probability of a comparison involving Nazis or Hitler approaches 1."
I think proprietary models will gradually become less popular, though, due to their lack of consistency and control.