The Google employees who created transformers
wired.com
wired.com
I believe energy minimization is literal, just look at the size of that thing and imagine the power bill.
The bitter lesson [0] strikes again.
[0] http://www.incompleteideas.net/IncIdeas/BitterLesson.html
Personally, I think both compute and NN architecture is probably needed to get closer to AGI.
And of course, a triggered Yann[2] (who is absolutely right).
But it is odd since it is actually a highly discussed topic, the history of attention. It's been discussed on HN many times before. And of course there's Lilian Weng's very famous blog post[3] that covers this in detail.
The word attention goes back well over a decade and even before Schmidhuber's usage. He has a reasonable claim but these things are always fuzzy and not exactly clear.
At least the article is more correct specifying Transformer rather than Attention, but even this is vague at best. FFormer (FFT-Transformer) was a early iteration and there were many variants. Do we call a transformer a residual attention mechanism with a residual feed forward? Can it be a convolution? There is no definitive definition but generally people mean DPMHA w/ skip layer + a processing network w/ skip layer. But this can be reflective of many architectures since every network can be decomposed into subnetworks. This even includes a 3 layer FFN (1 hidden layer).
Stories are nice, but I think it is bad to forget all the people who are contribution in less obvious ways. If a butterfly can cause a typhoon, then even a poor paper can contribute to a revolution.
[0] https://www.ft.com/content/37bb01af-ee46-4483-982f-ef3921436...
[1] https://www.bloomberg.com/opinion/features/2023-07-13/ex-goo...
[2] https://twitter.com/ylecun/status/1770471957617836138
[3] https://lilianweng.github.io/posts/2018-06-24-attention/
Today's incredible state-of-the-art does not exist without the transformer architecture. Transformers aren't merely some lucky passengers riding the coattails of compute scale. If they were, then the ChatGPT app which set the world ablaze would've instead been called ChatMLP, or ChatCNN. But it's not. And in 2024 we still have no competing NLP architecture. Because the transformer is a genuinely profound, remarkable idea with remarkable properties (e.g. training parallelism). It's easy to downplay GPTs as a mostly derivative idea with the benefit of hindsight. I'm sure we'll perform the same revisionist history with state-space models, or whatever architecture eventually supplants transformers. Do GPTs build on prior work? Do other approaches and ideas deserve recognition? Yeah, obviously. Like...welcome to science. But the transformer's architects earned their praise -- including via this article -- which isn't some slight against everyone else, as if accolades were a zero-sum game. These 8 people changed our world and genuinely deserve the love!
Is there any good summary of the history of AI/deep learning from, say, late 00s/2010 to the present? I think learning some of this history would really help be better understand how we ended up at the current state of the art.
The course starts all the way back from basic statistics and goes through things like linear regression and supposedly will arrive at neural networks and machine learning at some point.
So I don't know if something like this is exactly what you're looking for, but I think that, in general, if one wants to learn about (the history) AI, then it might be a good idea to start from statistics and learn about how we got from statistics to where we are now.
No, we do. State space models are both faster and scale just as well. E.g., RWKV and Mamba.
> Transformers aren't merely some lucky passengers riding the coattails of compute scale.
Err... they are, though. They were just the type of model the right researchers were already using at the time, probably for translating between natural languages.
However, some advances can have huge consequences to the field compared to others, even if at the technical level they appear comparable.
One example that comes to mind is CRISPR.
I asked "What would you do if you had an unlimited budget?"
He simply said, "I do"
GMail was one of the red flags for many that "don't be evil" was not going to be what it appeared. History says that this kind of mass profiling never ends well.
I worked at Borg
The quota system can kick in at whatever time the limits are reached.
And GPUs are scattered across borg cells, limiting the ceiling. That's why XBorg was created so that a global search among all Borg cells for researchers.
And data center Capex is around 5 billion each year.
Google makes hundres of billions of revenue each year.
You are asking what people would do in impossible situation. Like "what you do after you are dead", literally I could do nothing after I am dead.
I cannot even understand what I do stands for in the context of your question. The above is my direct reaction in the line that he assumes he had unlimited budget.
That he had a higher budget than he knew what to do with. When I worked at Google I could bring up thousands of workers doing big tasks for hours without issue whenever I wanted, for me that was the same as being infinite since I never needed more, and that team didn't even have a particularly large budget. I can see a top ML team having enough compute budget to run a task on the entire Google scrape index dataset every day to test things, you don't need that much to do that, I wasn't that far from that.
At that point the issue is no longer budget but time for these projects to run and return a result. Of course that was before LLMs, the models before then weren't that expensive.
The answer is that monopolies stifle technological innovation because one well-established part of their business (advertising-centric search) would be negatively impacted by an upstart branch (chatbots) that would cut into search ad revenue.
This is comparable to a investor-owned consortium of electric utilities, gas-fired power plants, and fracked natural gas producers. Would they want the electric utility component to install thousands of solar panels and cut off the revenue from natural gas sales to the utility? Of course not.
It's a good argument for giving Alphabet the Ma Bell anti-trust treatment, certainly.
And that's ignoring how Alphabet's core business, Search, has little to fear from GPT-3 or GPT-3.5. These models are decent for a chatbot, but for anything where you want reliably correct answers they are lacking.
By not pursuing this tech, they basically committed corporate suicide over the long run and they knew it. They knew very well, especially going into the 90's and early 2000 than their time making bank selling film was counted.
But as long as the money was there, the chemical branch of the company was all-powerful and likely prevented the creation of another competing product that would threaten its core business, and they did so right until the money flow stopped and suddenly they went basically bankrupt figuratively overnight since the cash cow was now obsolete.
The problem Kodak had is what the person you're replying to is alluding to. They got outcompeted because they were a photography company, not a digital hardware manufacturer. Companies like Sony or Canon did better because they were in the business of consumer electronics / hardware already. Building an amazing digital sensor and some good optics is great, but if you can't write the firmware or make the PCB yourself, you're going to have a hard time competing.
It's not chemical-wing vs digital-wing. It's that Kodak wasn't (and rightly wasn't) a computing hardware manufacturer, which makes it pretty damned hard to compete in making computing devices.
(Granted companies like Nikon or Leica etc did better, but it's all pyrrhic now, because the whole category of consumer digital cameras is disappearing as people are content to take pictures on their phones.)
Huh? This makes no sense. Sony was indeed a consumer electronics company at that time, but Canon was not: Canon was a camera manufacturer. They didn't get into electronics until later as cameras became digital: their earlier cameras were the all-mechanical kind. Sony was an electronics company that had to learn how to make cameras, but Canon was a camera company that had to learn how to make electronics. Kodak could have done the same.
Kodak, while they incidentally made some cameras, were a film and film processing company that wasn’t great at cameras and wasn’t anything in electronics.
They were much worse positioned than either a camera company, or a consumer electronics company, for a pivot to the post-film photography world.
Kodak could have diversified like that too, but they didn't. They were positioned badly because they concentrated almost all their efforts on film and nothing else. Of course, part of this is probably due to American business culture compared to Japanese; Japanese businesses tend to be much more diverse and long-term in thinking, but regardless, Kodak did this to themselves.
The problem was competing in hardware manufacturing, which is a whole different ballgame from concentrating on just the imaging aspects of it. So they were reduced, near the end, to just being really a (decent) component supplier to other companies. But that's the wrong part of the food chain to be in.
So they were already making digital hardware and so presumably had the internal expertise on how to product manage that.
Obviously, they were well-positioned since they had already moved into electronics and such before the digital revolution happened.
No, they shared it with Fujifilm (1/3rd each in 1990 according to https://pdgciv.files.wordpress.com/2007/05/tfinnerty2.pdf). And essentially via film, film processing but no significant camera market share, which is likely far more relevant to the transition to digital.
Sundar was afraid of the technology and how it would be received and tried to ice it.
1) Google enjoyed significant market status at the time and a leap forward like seemingly semi conscious AI in 2019 would be seen as terrifying. Consumer sentiment would go from positive to “Google is winning to hard and making Frankensteins monster”
2) it didn’t weave well into googles current product offering and in fact disrupted it in ways that would confuse the user. It was not yet productized but would already make the Google assistant (which at the time was rapidly expanding) look stupid.
A fictional character companion was not clearly a good path.
All this being said, I integrated the early tech into my product division and even I couldn’t fully grasp what it would become. Then was canned right at the moment it became clear what we could do.
Eric Schmidt was the only leader at Google who recognized nascent tech at Google and effectively fast tracked it. When my team made breakthroughs in word to vec which created the suggestion chip in chat/email he immediately took me to the board to demo and said this was the future. (This ironically wound a long path to contribute later to transformer tech)
Sundar often ignored such things and would put on a McKensey face projecting everyone just was far dumber than him and his important problems.
The second time I met him I presented my project (Google Cloud Genomics) to the board of the Broad Institute, of which he was a member (IIRC the chair) and he was excited that Google was using cloud to help biology.
The Broad is still a google cloud customer and verily seems to support all three clouds now, but I'm not aware of the details. The whole thing seemed really strange and unbusinesslike.
Sure we have Gemini but can Google take a loss in revenue in search advertising in their existing product to maybe one day make money from Gemini search? Advertising in the LLM interface hasn't been figured out yet.
Google (kind of) feels like an old school newspaper in the age of the internet. Advertising models for the web took a while to shake out.
Also, to be fair, OpenAI’s huge valuation is only a promise right now. Hopefully they will eventually turn profitable…
But they are probably using various techniques in analyses, embeddings, "canned" answers to queries, and so on.
ChatGPT has been around for a while now, and it hasn't led to a collapse in Google's search revenue, and in fact now Google is rushing to roll out their version instead of trying to entrench search.
A famous example is the iPhone killing the iPod, and it took around 3 and a half years for the iPod to really collapse, so chat and co-pilots might be early still. On the other hand handheld consumer electronics have much longer buying cycles than software tools.
>I have some tasks, some of them maintenance, some of them relaxation that I do >before bed. For example, maintenance: >brush teeth, floss, water floss, put on nose strips, take prylosec. >For relaxation: >stretch, massage gun, ASMR, music, tiger balm, salt bath, movie, tea, nasal >clean, moisturize. >can you suggest some others tasks?
Google can't even handle this. if you try, the entire page is filled with Videos: ... with 4 video thumbnails from youtube on how to use a water pic.
So I scroll off down (by now chat gpt already has a response) half a page on 'people also ask' with 4 questions about water pics, keep scrolling,
finally search results, and all of them are about ... water pics/flossers.
I reckon it is more that you needed special circumstances and series of events to become OpenAI. Not just cash and smart people.
ChatGPT has been around for less that 15 months[1]. In what version of reality is that "a while now"? Your iPhone/iPod timeline is also off by about 9 years[2].
1. https://en.wikipedia.org/wiki/ChatGPT 2. https://en.wikipedia.org/wiki/IPod_Shuffle
My iPhone/iPod timeline is referring to the time it took for the iPod's sales to be severely impacted [1]. I'm not disputing that the iPod continued to exist for a long time, the iPod Touch wasn't discontinued until 2022, but it was a pretty meaningless part of Apple's business by that point.
1. https://i.insider.com/597b68f2b50ab162018b466b?width=1300&fo...
There was a period where it was not obvious that the iPhone or Blackberry would become the dominant player in the field and it could be true with search and AI chat etc too.
Google is a business fundamentally oriented around loss leaders. They make money on commercial queries where someone wants to buy something. They lose money when people search for facts or knowledge. The model works because people don't want to change their search engine every five minutes, so being good at the money losing queries means people will naturally use you for the money making queries too.
Right now LLMs take all the money losing queries and spend vast sums of investor capital on serving them, but they are useless for queries like [pizza near me] or [medical injury legal advice] or [holidays in the canary islands]. So right now I'd expect actually Google to do quite well. Their competitor is burning capital taking away all the stuff that they don't really want, whilst leaving them with the gold.
Now of course, that's today. The obvious direction for OpenAI to go in is finding ways to integrate ads with the free version of ChatGPT. But that's super hard. Building an ad network is hard. It takes a lot of time and effort, and it's really unclear what the product looks like there. Ads on web search is pretty obvious: the ads look like search results. What does an ad look like in a ChatGPT response?
Google have plenty of time to figure this out because OpenAI don't seem interested. They've apparently decided that all of ChatGPT is a loss leader for their API services. Whether that's financially sustainable or not is unclear, but it's also irrelevant. People still want to do commercial queries, they still want things like real images and maps and yes even ads (the way the ad auction works on Google is a very good ranking signal for many businesses). ChatGPT is still useless for them, so for now they will continue to leave money on the table where Google can keep picking it up.
* https://www.youtube.com/watch?v=QWWgr2rN45o
* https://www.youtube.com/watch?v=E14IsFbAbpI ('mirror')
Goes over Hinton's history and why he went the direction he did with his research, as well as Li's efforts with ImageNet.
Subtle plug for return-to-office. In-person face-to-face collaboration (with periods of solo uninterrupted deep focus) probably is the best technology we have for innovation.
Which is usually impossible in the office. So more like a mix, which is what all reasonable people are saying.
As much as I think the America has lot of things it needs to fix, there is no other country on earth this would be possible. That's just a fact.
I don’t think this the case. If anything the US makes life very hard for even high-skilled work-based immigrants. Many countries have a higher % of foreign born residents than the US (Singapore, Australia, Germany, Canada)
I myself used to work at Google UK and my own team was 100% foreign born engineers from every continent.
So many monumental moments in AI history are archived in Google's intranet.
“ Six of the eight authors were born outside the United States; the other two are children of two green-card-carrying Germans who were temporarily in California and a first-generation American whose family had fled persecution, respectively.”
Not in California. Last I remember, something like a quarter of the state’s population is foreign born.
> All Students in Higher Education in California
> 2,737,000
> First-Generation Immigrant Students
> 387,000
14% are first-generation immigrants
from: https://www.higheredimmigrationportal.org/state/california/
About 1/8 of the US population is foreign-born, which is a minority but not a tiny one. In California, its over a quarter.
There is a motivation that comes with both trying to make it and being cognizant of the relative opportunity that is absent in the second-generation and beyond.
There are also many advantages given to students outside the majority. When those advantages land not on the disadvantaged but on the advantaged-but-foreign, are they accomplishing their objectives? How bad would higher education have been in Europe? What is the objective, actually?
The US population is around 330 million. The world population is 8.1 billion people. What is that 4%? If you took a random sampling of people around the world, none of them would be Americans. You're going to need a lot more samples to find a trend.
Yet when you turn around and look at success stories, a huge portion of this is going to occur in the US for multiple reasons, especially for the ability to attract intelligence from other parts of the globe and make it wealthy here.
For an American, it's less of a good deal. Once you have the PhD, you make somewhat more money, but you're trading that for 6 years of hard work and very low pay. The numbers aren't favorable -- you have to love the topic for it to make any sense.
As a result, U.S. PhD programs are heavily foreign born.
Not sure if we can claim this any more, what with texas shipping busloads of immigrants to NY and the mayor declaring it a citywide emergency, and both major parties rushing to get a border wall built.
"Illegal" is a concept - it's not conflating to assume that it's not the bedrock of the way people think.
Illegal immigrant is pretty well defined, an immigrant that didn't come through the legal means. The people hired by Google are probably not illegal immigrants.
America is much more welcoming of immigration, by which I mean legal immigration, than Japan or China. This is not in dispute.
It is also, in practice, quite a bit more slack about illegal immigration than either of those countries. Although I hope that changes.
It's not? It sounds like you know little about the world outside of America. Japan is stupidly easy to immigrate to: just get a job offer here at a place that sponsors your visa and it's pretty trivial to immigrate. Even better, if you have enough points, you can apply for permanent residence after 1 or 3 years, and the cost is trivial. In America, getting a Green Card is very difficult and costly, and depending on your national origin can be almost impossible. In Japan, there's no limits at all, per year or per country of origin, for work visas or PR. Of course, Japan is somewhat selective about who it wants to immigrate, but America is no different there, which is why there's such a huge debate about illegal immigration (in America it's not that hard; in an island country it's not so easy).
This is why the CEO of Coinbase sent the memo a few years back stating that these types of people (not mission focused, self absorbed, and distracting) should leave the company. The CEO of Google should be fired.
Edit: I'm also unclear on what makes you think he's spending much time on "arbitrary hiring goals". I remember an initiative for hiring in a wider variety, also conveniently much cheaper, locations but there wasn't indication he was personally spending time on it
Humans are more than their companies' mission, and your freedom to exercise your political views on how companies should be run inherently relies on this principle, too. So your argument is fundamentally a hypocrites projection.
Not when they’re on the clock.
Besides, many of the Google employees who are against defense projects are against them because those projects actively target their home countries. Much harder problem to fix!
It is a successful model that has worked again and again to escape the problem of corporate bureaucracy.
It is tough to find the right balance though, because AI safety is not something you want to brush off.
It really depends on what exactly is meant by "safety", because this word is used in several different (and largely unrelated) meanings in this context.
The actual value of the kind of "safety" that led to the Gemini debacle is very unclear to me.
Google is an ad company at the end of the day. Google is still making obscene amount of money with ads. Currently Gen AI is a massive expense in training and running and is only looking like it may harm future ad revenue (yet to be seen). Meanwhile OpenAI has not surpassed that critical threshold where they dominate the market with a moat and become a trillion dollar company. It is typically very hard for a company to change what they actually do, so much so that in the vast majority of cases its much more effective (money wise) to switch from developing products to becoming a rent seeking utility by means of lobbying and regulation.
Simply put, Google itself could not have succeeded with GenAI without an outside competitor as it would have become its own competition and the CEO would have been removed for such a blunder.
It harms ad revenue at the moment.
For some specific queries (like code-related), I now go to chat.openai.com before google.com, I'm sure I'm not the only one.
AI is the only thing that can put an end to Google's dominance of the web. I also have no idea how Sundar still has a job there.
Microsoft started out selling BASIC runtimes. Then they moved into operating systems. Then they moved into Cloud. Satya seems to be putting his money where his mouth is and working hard to transition the company to AI now.
Apple has likewise undergone multiplet transformations over the decades.
However Google has the unique problem that their main cash cow is absurdly lucrative and has been an ongoing source of extreme income for multiple decades now, where as other companies typically can't get away with a single product line keeping them on top of the world for 20+ years. (Windows being an exception, maybe Microsoft is an exception, but say what you will about MS, they pushed technology forward throughout the 90s and accomplished their goal of a PC in every house)
Google doesn't really do acquisitions on that level. They do buy companies, but with the purpose of incorporating their biological distinctiveness into the collective. This tends to highlight their many failures in capturing new markets. The last high-profile successful purchase by Google, that I recall at least, were YouTube and Android, nearly twenty years ago.
I mean these are kind of tangled together. Chatbots actively dilute real sites from what we are seeing, by feeding back into their fake into googles index in order to capture some of the ad revenue. This leads to the "Gee Google is really sucking" meme we see more and more often. The point is, Google had an absolute cash cow and the march towards GenAI AGI threatens to interrupt that model just as much as it promises to to make the AGI winner insanely rich.
Like, who wants to use an AI that says things like, "... and that's why you should wear sunscreen outside. Speaking of skin protection, you should try Banana Boat's new Ultra 95 SPF sunscreen."
In any case, consumer chatbots isn't the only way to sell the tech. Lot's of commercial use too.
I don't see why ads couldn't be integrated with chatbots too for that matter. There's no point serving them outside of a context where the user appears to be interested in a product/service, and in that case there are various ways ads could be displayed/inserted.
Go to google -> Search "sunscreen" -> (Sponsored) Banana Boat new Ultra 95 SPF sunscreen
Oh wow! 95 SPF! Perfect
On the other hand, their history suggests most people would be fine with an AI which did this as long as it was accurate:
> ... and that's why you should wear sunscreen outside.
> Sponsored by: Banana Boat's new Ultra 95 SPF sunscreen…"
Google purchased YouTube for $1.65 billion (and I recall that seemed overpriced at the time).
The way VC works for private companies now is completely different than in 2006.
The marvel of modern computing is a result of RnD bloat that was done without immediate impact to the bottomlines of their own companies.
https://www.plantemoran.com/explore-our-thinking/insight/202...
Losing a lot of money and you can't write off your engineers
In fact, it's arguable if anything in R&D qualifies as an expense at all.
And as it is, Google certainly paid property taxes immediately on the office building as well as FICA on all of the employees who, of course, paid their own taxes.
But haters of the R&D system love to call it "tax-free".
https://www.plantemoran.com/explore-our-thinking/insight/202...
The Vision transformer is capable of predicting that a running child is likely to follow that rolling soccer ball. It is capable of deducting that a particular situation looks dangerous or unusual and it should slow down, or stay away from danger in ways that previous crop of AI could not.
Imo, the only thing currently preventing transformers to change everything is the large amount of compute power required to run them. It's not currently possible to imagine GPT4-V running on a embedded computer inside a car. Maybe AI asic type of chip will solve that issue, maybe Edge computing and 5G will find it's use-case... Let's wait and see, but I would bet that transformer will find it's way in many places and change the world in many more ways than bringing us chatbots.
It all needs to be onboard. That’s where money should be going.
More seriously, for safety critical applications, LLM have some serious limitations (most obviously hallucinations). Still, I beleive they could work in automotive application assuming: high quality of the output (better than current SoA) and very high token count (hundreds or even thousand of token/s and more), allowing to bruteforce the problem and run many inferences per seconds.
Clearly we are not there yet.
I wasn't intending to say it would be useful today, but pushing back against what I understood to be an argument that, once we do have a model we'd trust, it won't be possible to run it in-car. I think it absolutely would be. The massive GPU compute requirements apply to training, not inference -- especially as we discover that quantization is surprisingly effective.
That's also where I would see transformers or another AI architecture with reasoning capabilities shine: the fact that it can reason about what is about to happen would allow it to handle edge cases much better than relying on dumb sensors.
As a human, it would be very difficult to drive a car just looking at sensor data. The only vehicule I can think of where we do that is submarines. Sensors data is good for classical AI but I don't think it will handle edge case well.
To be a reasonable self-driving system, it should be able to decide to slow down and maintain a reasonable safety space because it is judging the car in front to be driving erratically (ex: due to driver impairement). Only an AI that can reason about what is going on can do that.
What is vision if not sensor data?? Our brains have evolved to efficiently process and interpret image data. I don't see why from-scratch neural network architectures should ever be limited to the same highly specific input type.
I also think dumb sensors is unfair, there are Neural Network solutions for processing LIDAR data so we are talking about a similar level of intelligence applied over both sensors.
If they were threatened, Google could have easily owned the cloud. They were the first to a million servers and exabyte storage scale. Their internal infra is top notch.
Under no circumstances can they damage their cash cow. They have already shown to be UX hostile by making Ads look like results, and Youtube forcing more ads onto users.
The biggest breakthrough an AI researcher can have at Google is to get more ad clicks across their empire.
Google management knows this. Without ad revenue, they are fucked.
That is my biggest fear. That with AI we are able to exploit the human mind for dopamine addiction through screens.
Meta has shown that their algorithm can make the feeds more addictive (higher engagement). This results in more ads.
With Insta Reels, Youtube, TikTok e.t.c "Digital Heroin" keeps on getting better year after year.
Besides that, are you denying that transformers are the fundamental piece behind all the AI hype today? That's the point of the article. I think the fact that a mainstream publication is writing about the Transformers paper is awesome.
Modern AI should anything after https://en.wikipedia.org/wiki/Dartmouth_workshop Not transformer
AI has been people's dream since written history. i.e., everyone in their own sense would want to invent something that can do things and think for themselves. That's literally the meaning of AI.
Am not expert, do you have some links about this? i.e. a neural net construction that outperforms a transformer model of the same size.
Consider "modern" to mean NN/connectionist vs GOFAI AI attempts like CYC or SOAR.
I guess it depends on how you define "AI", and whether you accept the media's labelling of anything ML-related as AI.
To me, LLMs are the first thing deserving to be called AI, and other NNs like CNNs better just called ML since there is no intelligence there.
Well this is what Im trying to say too!
I dunno. The earliest research into what we now call "neural networks" dates back to at least the 1950's (Frank Rosenblatt and the Perceptron) and arguably into the 1940's (Warren McCulloch and Walter Pitts and the TLU "neuron"). And depending on how generous one is with their interpretation of certain things, arguments have been made that the history of neural network research dates back to before the invention of the digital computer altogether, or even before electrical power was ubiquitous (eg, late 1800's). Regarding the latter bit, I believe it was Jurgen Schmidhuber who advanced that argument in an interview I saw a while back and as best as I can recall, he was referring to a certain line of mathematical research from that era.
In the end, defining "modern" is probably not something we're ever going to reach consensus on, but I really think your proposal misses the mark by a small touch.
The modern era of NNs started with being able to train multilayer neural nets using backprop, but the ability to train NNs large enough to actually be useful for complex things AI research, can arguably be dated to the 2012 Imagenet competition when Geoff Hinton's team repurposed GPUs to train AlexNet.
But, AlexNet was just a CNN, a classifier, which IMO is better just considered as ML, not AI, so if we're looking for the first AI in this post-GOFAI world of NN-based experimentation, then it seems we have to give the nod to transformer-based LLMs.
Transformer+Scaling+$$$ triggered the current hype
It had been a relatively gradual and accelerating progress, with number of people working in the field increasing exponentially, since 2010 or so, when Deep Learning on GPUs was popularized at NIPS by Theano.
Tens of thousands people working together on Deep Learning. Many more on GPUs.