Intel Announces Aurora GenAI, ChatGPT Competitor with 1T Parameters
wccftech.com
wccftech.com
So this thing doesn't actually exist yet?
This looks like an announcement of an announcement, which isn't on topic for HN. https://hn.algolia.com/?dateRange=all&page=0&prefix=true&sor...
On HN, there's no harm in waiting for the actual thing: https://hn.algolia.com/?dateRange=all&page=0&prefix=false&so...
and instead of a data center with a distributed load of Nvidia cards, they’re going to do 1 supercomputer? Are these guys intentionally stuck in 2007 or am I missing something
You're missing something. Most models are trained on supercomputers. You still use DDP, but you also use Slurm. A supercomputer is just a many node machine, but one where you REALLY care about interconnect. It's why the fabric is one of the big points here. I/O is generally your bottleneck in these systems. Each node should be as powerful of a system as possible because of this.
A datacenter doesn't need a PB/s connection between machines. Your I/O probably isn't your bottleneck and so the tradeoff between PCIe and ethernet (infiniband or whatever) isn't a big deal. Usually because your processes are __independent__ (containers). In a supercomputer, your processes may be parallel, but that doesn't mean they're independent. You'll spend a lot of time writing code to be non-blocking but you're going to have to map reduce at some point.
- 200Gb dragonfly interconnect fabric
- 10,000 nodes with 2xCPU + 6xGPU with unified memory
- 10PB of agreggate RAM
- 230PB of storage
Scale matters in supercomputing. Keeping latency low between everything allows you to solve different classes of problems that gets no speed up benefit by simply offloading compute to nvidia cards. It makes a huge difference how everything is connected.
They also might not be measuring the same thing. H100 GPUs have "3026 TFLOPS":
https://www.pny.com/nvidia-h100
...at FP8. 51 TFLOPS at FP32. Supercomputer lists typically use FP64.
An AI that does that detection would be both wonderful and dangerous.
(That book abandoned its only interesting ideas and went totally off the rails a few chapters later IIRC.)
Still, it is a great question and part of me wonders about the whatifs of this evolution.
https://www.businesswire.com/news/home/20230522005289/en/Int...
For starters future tense is embarrassing in this area. By the time they finish it, it may very likely be outdated.
Also number of parameters is not all you need. It's enabler, yes, but you need tons of other things to fit together.
I'm not convinced that simply increasing the number of parameters improves the models, and I don't think we should be putting out press releases "spec-chasing" increases in the number of parameters. I've also found in my research that larger models often do not perform better, but this is becoming more difficult to explain to non-technical folks due to irresponsible marketing. If we aren't careful about the bold marketing claims, we'll disillusion many people and reduce the potential for future growth—we don't want another AI Winter.
There's a significant environmental cost to training these large models, and it's harder to understand, use, and control these very large models (even as a researcher).
At every size people say this... And then someone comes out with a model 10x larger which outperforms the previous state of the art.
Sure, you need to still take a little care with architecture choice and hyperparameter selection, but size really does seem to be king.
Speak for yourself...
Most of the huuuuuge models failed on most or all of these fronts and that's why they suck compared to Llama or Alpaca or Vicuna
How so? Llama 33B is clearly worse than llama-64B and that's 'only' a 2x increase.
Are you aware of any example where a larger model fails to outperform a smaller one when all else is equal (tokens, architecture, data quality, etc.)?
Obviously for a fixed amount of training compute more parameters can be bad, but there's a trade off where more parameters means you train on fewer tokens.
Exactly. The worse part is there is NO viable efficient way of training, fine tuning etc and inference with these massive deep learning models for more than a decade since deep neural networks used GPUs for training, and it still requires a substantial amount of compute power and energy that is incinerating the planet and the result is untrustworthy large black-box models that cannot be trusted or transparently explain their decisions, especially with safety-critical or high risk tasks.
Crypto at least has an alternative to the wasteful proof-of-work system, and Ethereum which was formerly PoW has shown it is possible to switch to a greener alternative consensus method.
Deep Learning however, still has not shown such a viable switch and still needs to burn the planet with more data-centers of GPUs, ASICs and FPGAs to create hallucinating models that have shown to break on a single pixel or to confidently regurgitate nonsense as the truth with little understanding in reasoning and also answer with demonstrably false information.
LLMs like this one is still essentially snake-oil BS generators hiding behind regurgitation and sophistry to pretend to show signs of 'intelligence'.
EDIT: It is all true. [0] There is no amount of green-washing to hide the problem of deep learning systems wasting essential resources like thousands of running taps of water. Literally.
[0] https://gizmodo.com/chatgpt-ai-water-185000-gallons-training...
If sales slow down, instead of having the excess stock in a box in a warehouse, have it training their neural nets.
So far, nobody has claimed to exceed GPT-4's performance (either by academic metrics or by youtubers subjective evaluations)
Not outright the most efficient utilization of data, but as others have mentioned, there might still be something left to gain (from not solely optimizing for compute). E.g. look at Figure 9 in [3]. Although the model trained on the largest model obviously utilized the data most efficiently, it's not perfectly clear - to me at least - whether some over-parameterization will necessarily lead to a decrease of test/out-of-sample performance.
Of course, LLaMA was trained on way fewer parameters relative to data set size. I've just mentioned LLaMA as a point of reference to the largest dataset known to me.
[1] - https://arxiv.org/pdf/2302.13971.pdf
[2] - https://arxiv.org/pdf/2005.14165.pdf
[3] - https://arxiv.org/pdf/2001.08361.pdf
EDIT: Formatting of references
1) This is going to be using a publicly funded computer, which is the most powerful in the world AND was also announced today (btw, the specs are better than what was initially planned). The program gives justification for the machine and that public money (though in comparison to many things governments spend money on, this is very cheap).
2) THAT'S A LOT OF GPUS. This will, as far as I'm aware, be the biggest training every. GPT4 was supposedly trained on 10k A100s. Remember that Summit, the #3 computer, has >27k GPUs (V100s), but Aurora has almost 64k. This is going to be a big engineering feat in of itself. Now it won't make training 6x faster, but it will definitely be much better. It does say that the US government is taking this very seriously.
That's the point of this kind of announcement: marketing, propaganda, and cool engineering project.
ok but when openai made one it was called a datacenter
Did nobody notice?
Intel is PLANNING to train a genAI model with a target of 1bn parameters... using an architecture from 2019
inputs+outputs? Number of neurons+synapes?
Well, how are they "connected"? Their topology?
Something is not right here.
> Meanwhile, the target size for the free & public versions of ChatGPT is just 175 million in comparison. That's a 5.7x increase in number of the parameters.
5.7x seems off by several orders of magnitude, or do I need more caffeine?
GPT-4 does not have publicly available parameter count, but is thought to have something similar to 1tn.
I wonder how much of a competitive moat actually exists in this space.
I also don't think I'd call what they're offering "on par" with GPT-4. They're competitive, or impressive substitutes, but my personal tests of Bard at least aren't getting me as useful results as GPT-4.
Source: too lazy to look up any. this is just my lame impression. feel free to downvote.
My favorite hypothesis is that OpenAI was navigating the unknown open sea trying to answer the question "Is it possible to achieve this?" without knowing if they would ever find a new land.
The rest of them already know the answer is "Yes it is" so it's not an unknown open sea anymore with a dubious destination. They know it's possible, it's just a matter of money and time. They probably hired whatever engineer was available, and they paid tons of money to come up with a competitive product.
Intel have announced they are building a model. Anyone can announce they are building a model.
Executed well, plug-ins could become ChatGPT's App Store equivalent. Once that happens, OpenAI is undoubtedly going to convert a portion of search traffic into agent delegated work i.e., people will use ChatGPT to start and terminate their searches essentially delegating the work of collating data from multiple sources to ChatGPT.
FD - I've applied to release a ChatGPT plug-in and this space looks very interesting to me.
PS - Earlier today, I submitted to HN a link to my blog where I analyzed plugins. The HN post sank without a trace but the I wrote my blogpost using the Wolfram plugin in ChatGPT and it was a breeze. I genuinely feel I've seen the future.
I have no doubt that large players can possibly lap OpenAI in terms of ecosystem surrounding a decent AI core.
However, OpenAI has the leg up and it's their game to lose at this point.
This is only a factor if the underlying AI/LLM model becomes as good as OpenAI AND serves millions of people. While open source LLMs are soon going to gain parity with GPT-4/Bard, it is unlikely that they will hit the kind of usage numbers that an OpenAI/Google/Meta can deliver.
Even there, I'm not sure Google/Bing will cannibalize their own product to allow third parties to inject data into a search interaction.
ChatGPT is a fundamentally different product - it's a humanlike intelligence which is always ready to talk to you and assist you. We haven't ever had anything of comparable quality and reach. The nearest was Alexa but even she was limited to Amazon's catalog and shopping. It's like ChatGPT can be everyone's personal gofer and THAT is huge, imo.
No that's incorrect. It is auto-fill on steroids trained on years of internet postings from actual humans (ie, reddit, twitter, etc).
That input is going to dry up - no longer free (Reddit has said all future posts are going to cost $ to access). Good for OpenAI as it has such a headstart, but many companies can source/tap into legitimate alternative user data streams - FAAMG for sure, but also any company that can convince it's userbase to provide training data that already has a foothold.
OpenAI is getting stuff from Bing, but Microsoft controls that data.
I think OpenAI is doing the opposite by massively degrading GPT-4 to support the load from all these integrations (as well as their new app). Their moat was the head-and-shoulders-above—their-competitors quality of GPT-4, which had now taken a huge nosedive. I’m not sure why anyone would pay for it instead of GPT 3.5-Turbo. It went from being a competent, if somewhat error probe coder to doing things like randomly inserting C# code blocks in text.
Unless, I'm missing something, the spec you expose to ChatGPT only tells them which api endpoint to hit. The code powering that endpoint is not visible to ChatGPT.
The only argument you could make is that you are giving it machine understandable text to describe an API endpoint. In future, it might not even need the text description if the API endpoint is named well. Quite a stretch though.
There is no business model for plugins.
For APIs it could lead to increased use, but only if the cost is really marginal.
1. single point of entry: you don't need ten different apps to achieve ten different outcomes. You can do it all right inside ChatGPT
2. Since I can have up to three plugins active inside ChatGPT (as of today), I can express more complex workflows than if I had one app each with no trivial way to stream data from one app to the next.
It's like Zapier for your ideas. You talk to ChatGPT > it extracts data from a plugin > you prod it along a bit more > it talks to a second plugin to do X > more chat > talk to third plugin etc etc.
Zapier itself only lets you flow data from one app to the next, not act on that data.
I expect that as the LLM evolves and matures, they will allow more plugins and start eating up different industries. E.g., why do you need a Google Drive when your ChatGPT can also store files for you?
If the idea is to keep the actual files rather than summaries, than there's an entire world of requirements (data storage reliability, access control, auditing, integration) where OpenAI etc. have no competitive advantage, and their own issues (e.g. prompt injections). LLMs are a bad fit whenever you need it to act like a computer, they replace human style processing. For math, get a calculator.
In this case, I'd expect some way to interface to OneDrive or maybe even BackBlaze.
That’s in terms of technology
Their brands are still very powerful
Anyone can sell sugary drinks, but there is only one CocaCola
They have a lot of down time and change the model out from under people, only having 3 months support. It is very hard to build a product on OpenAI in this situation.
A close enough model that can run on your own hardware or cloud will get people to move. We are watching and waiting for this point for our own products. For now they are called experiments and demo applications.
OpenAI might not have the best reputation within tech, but people who are not in tech are completely unaware of this
Same with Google and Microsoft
Those companies are dominating the mainstream tech media/conversations
Even if they are not the best tech or don’t have the best products/services