Character.ai CEO Noam Shazeer Returns to Google
techcrunch.com
techcrunch.com
“Over the past two years, however, the landscape has shifted; many more pre-trained models are now available. Given these changes, we see an advantage in making greater use of third-party LLMs alongside our own. This allows us to devote even more resources to post-training and creating new product experiences for our growing user base.”
My interpretation is that Character.AI realized they don't actually need to train their own foundation models from scratch to support their product - they can build cheaper, faster and probably better if they use LLMs trained by other companies (could be GPT-4o/Claude/Gemini via APIs, could be Llama 3.1 self-hosted).
If they're not training foundation models any more, the talents of people like Noam Shazeer aren't so important to them. They need to focus on product development instead.
Does Meta get in the way of this?
It's hard to compete with a company that is dead set on spending billions and seemingly wants to drive your SOTA AI product revenue to 0.
If you are OpenAI or Anthropic right now, it seems like trying to run a great restaurant at a reasonable price right next to a good (great?) restaurant that is serving everyone for free.
Presumably this is because Meta desperately want to avoid becoming dependent on other companies in this new generative AI world. Mark Zuckerberg talks about not wanting a repeat of the Apple tax in his post about Llama 3.1 here: https://about.fb.com/news/2024/07/open-source-ai-is-the-path...
I think it is just a consequence of the cost of getting to the next level of AI. The estimates for training a GPT-5 level foundational model are on the order of 1 billion. It isn't going to get cheaper from there. So even if your model is a bit better than the free models available today, unless you are spending that 1 billion+ today then you are going to look weak in 6 months to 1 year. And by then the GPT-6+ model training costs will be even higher, so you can't just wait and play catch up. You are probably right as well, in that there is a fear that a competitor based on an open source model gets close enough in capability to generate bad publicity.
I imagine character.ai (like inflection) did calculations and realized that there was no clear path to recoup that magnitude of investment based on their current product lines. And when they brainstormed ways to increase return they found that none of the paths strictly required a proprietary foundational model. Just my speculation, of course.
Where is the data that costs a billion dollars to train on going to come from? These companies are already training on most of the available valuable information that exists.
While training will surely be expensive, I think it's even more expensive and challenging to organize and harness the brainpower to figure out and execute the next meaningful step forward.
Data availability for LLM is becoming trickier. There are at least two avenues being explored: A) Synthetic data (in controlled ways) and B) Video data, in particular multi-modal embeddings between image/audio/text sequences. This may enable several magnitudes of increase in compute.
That is a short hand for "the next two generations of LLMs created by Open AI". It is not meant to be a forward looking statement on how those models will be branded in the consumer market. It also isn't meant to be a prophecy that OpenAI will maintain its premier position since Anthropic or even a new entrant into the field might be the company to achieve that next step level.
> I think it's even more expensive and challenging to organize and harness the brainpower
Then you should invest with that in mind. What I find interesting is that Microsoft (with its acquisition of the research arm of inflection) and Google (with its acquisition of the research arm of character.ai) seem to see the foundational model and product categories as distinct. It is that distinction I am interested in.
There is no doubt some huge value in productizing these LLMs. However, it appears that the productization of LLMs and the advancement of the foundational models themselves are being decoupled by the market. That is, it seems they are segregating risk. Product companies can raise money to build products, "platform" companies (e.g. Microsoft, Google) can raise money to build foundational models. What seems less popular based on these recent moves is companies able to raise money to build foundational models for the purposes of specific products.
This is a lack of imagination.
All books written; all movies that exist on dvd; all music released in cd; all tv programs; all radio programs; all whatsapp messages; all of youtube; blueprints from architecture and mechanical engineering.
The copyright and logistics are definitely an issue, but there is more data.
At some point I expect putting in too much data from semi-random or very old sources will have a detrimental effect on output quality.
In the extreme case, you could feed /dev/urandom. Haha, only kidding, but I'm sure you get my idea.
Now I'm wondering what a model trained in the past 45 years of Usenet would be like. Or all of the history of public messages on IRC servers like EFNet or Freenode (afaik they are not fully logged). It is an interesting topic, but I'm still curious and uncertain what the effect adding some multiples of data in the form of often lower fidelity sources (e.g. WhatsApp messages) will have on the capability of the final model. It's hard to understand how such sources would be helpful or useful.
LLMs are become akin to tools, like programming languages. They’re blank slates, but require implementation to become special.
Why is the CEO important to model development regardless of talents? They've raised $150m+, have $15m+ ARR and ~200 employees, etc. Shouldn't the CEO be CEOing?
Edit: reading the comments below, it seems like maybe he thought the expected value of attempting to clear the hurdle of their valuation/liquidation preferences at a $250k/year salary as CEO was lower than a $5m+/year salary/RSUs from Google?
So basically leaving a shell of a company and the GC to try and run it / wind it down.
I assumed Character.AI wasn’t profitable, like most AI startups.
Currently a money furnace.
Character.ai could've been at $100mm+ ARR if they did a bit more of a monetization push based on my very rough estimates. If it was an acquisition I would've been imagining $3b+ price range.
Huge get by Google! (Side note, Gemini 1.5's new alpha release from this week is now at the top of the lmsys leaderboard and sentiment on twitter for it is that it's strong, maybe as strong as sonnet 3.5, so it'll continue to be an interesting race between meta, openai, anthropic, gdm.)
Edit: Okay perhaps it's - Noam keeps his C.ai stock, gets a big pay package from Google ($5mm-$15mm/year kinda range? not sure). Most of the value in C.ai remains, and he keeps his stock.
where did you get $500mm number?..
Investors are being bought at 2.5B valuation
Base on those numbers, $150M of series A with 1B valuation would be bought now for 375M(which is high for failed startup), my understanding is that such transaction has to be disclosed.
I may be too skeptical, but it is just hard to believe someone (who?) throws 400M (to seed + A investors) plus significant amount to founders into essentially failed startup.
Actually a nice and polite guy, too.
For those asking, c.ai has very high cost and looks like a typical consumer company that burns money for use, so they were decent on revenue but not near profitability.
https://www.theinformation.com/articles/google-hires-charact...
Those investors almost certainly have a liquidation preference. How much did employee shareholders get? I'd guess zero.
"I am confident that the funds from the non-exclusive Google licensing agreement, together with the incredible Character.AI team, positions Character.AI for continued success in the future,” Shazeer said in a statement given to TechCrunch."
That's a pretty hilarious statement from a Founder/CEO, given the circumstances.
But they aren't filing for chapter 11? I assume all shareholders will be bought out, including the employees, and this will be paid for by Google who will license their models, presumably as a scheme to pay off of the investors as I doubt they actually need those models at all.
(assuming the linked source is correct.)
What do you mean by this? Fundamental research? It makes sense, most of the innovations now will be in things like clever caching mechanisms and reducing compute.
If there was any real intent to give c.ai a chance as a real business they would have hired a new ceo before the announcement.
It's like those rich guy/poor guy jokes.
Poor guy leaves his job, tries bootstrap its own company and fails, comes back to its old job. "What a loser".
Rich guy leaves his job, gets 150M to start a company in a blue ocean with a significant competitive advantage over 99.99% of humans alive. Still manages to fail and comes back to its old job. "What a bold move!"
Oh there is a place where rich guys can "get" $150M, but poor guys couldn't.
The more money you raise the more the pressure - and for deep researchers, this is usually noise that takes you away from your passion.
Shazz may be the kind of person that just wants to concentrate on research and doesn't really care about the rest...
1. Find an unhappy senior AI exec from who's published a few papers who's published a few papers. 2. Start a new org around them and hire a few key people with crazy salaries (which you can offer cause the time horizon for the company isn't that long). 3. Train a few models, release some good looking benchmarks (bonus points if big tech lend you their GPUs as part of some 'accelerator deal'). 4. Maybe find PMF and become incredibly rich. 4. If that fails, sign a massive but undisclosed licensing deal for your tech with big tech and give them your staff.
Seems like a good way to take big bets in AI, while hedging most of the risk.
Only difference in this cycle is that you can't do real LLM research in academia these days so all of the top researchers are already at FANG.
this is basically Inflection 2.0.
It doesn't necessairly imply Character.ai's business isn't doing well, but a CEO leaving is still indeed weird.