Amazon spends another $2.7B on Anthropic
cnbc.com
cnbc.com
For commercial offerings, OpenAI, Anthropic, Mistral, etc. have affordable APIs, and there are a bunch of companies like Scaleway for running open weight models too large or inconvenient to run at home.
At the application layer, I like the new emphasis on agents, open source apps that use local models, and highly refined products like Perplexity, OpenAI app/web, Google’s workplace integrations, and all the stuff that Microsoft is doing.
There are problems like the environmental costs of training and running models, but things will get more efficient and people will realize the utility of smaller special purpose models.
I'm not really sure what you are expecting to do with CPU. You might be able to get some <400 token responses and have fun, but you aren't going to be doing 2000 token encyclopedia style responses unless you are going to wait ~20 minutes for a response.
Alternatively, you can get like a $800 gaming laptop that has 8gb vram or use something like vastai where you can get 12gb vram for $0.10 an hour.
There is a reason no one ever shows >2000 token reponses.
> For a 34b q8 sending in 6000 context (out of a total of 16384) I get about 4 tokens per second.
https://www.reddit.com/r/LocalLLaMA/comments/18b1qgy/comment...
You can't even fit this on 2x RTX 4090.
Like, use the 7B berkley sterling on a freaking 3060 laptop and you are still getting better results on both output and tokens.
I don't understand what they are getting out of these large models that perform worse.
Also runs Yi-34B-Chat, which takes up ~18.15GB of RAM.
The current trajectory is that the size of these models, and number of FLOPs needed to train them, is growing much faster than the cost of compute is coming down. GPT-4 apparently cost around $100M to train, and it seems people are expecting that may quickly rise to $1B, and maybe then $10B for upcoming generations. I assume these numbers are based on projected 10x per generation increase in tokens processed during training.
So, with these kind of numbers, apparently there's a thought floating around that if they don't achieve AGI in next few model generations then that may stall AGI progress until different more efficient methods are developed. Private companies may be willing to spend $1B, then $10B to train more powerful models, but are likely to balk at $100B or above unless there's an obvious payback on the horizon.
Any sunken training costs may essentially be cumulative from one generation to the next, unless each generation can pay for itself. If we assume a new model every year, then can a $10B model pay for itself in a year before being replaced the next year by an even more expensive one?
The next generation is looking to cost more like a few trillion, which I know sounds extravagant.
So don’t take it from me, take it from Trevor Blackwell on HN:
https://news.ycombinator.com/item?id=39666368
Yes, these things cost too much to make economic sense absent some non-disclosed advantage to a small group of people deploying them with zero oversight.
Certainly neither OpenAI or Anthropic have that sort of cash on hand, nor have been offered any trillion dollar "train now, pay later" deal. Also, compute is is short supply atm (not up 1000x from last year), and even if the money and compute were available, I couldn't see these companies committing to that kind of model size/spend increase in just one or two generations - they need to see continuing progress at each iteration to justify and guide future direction.
$1B training costs for upcoming (GPT-5, Claude-4) models wouldn't be so surprising though (at least not now that we've got used to these crazy numbers).
Note that scaling up model size doesn't necessarily refer to context size anyway - may also be things like embedding dimensions size, number of transformer layers, number of experts (MoE), etc.
In general scaling up 10x in size and/or cost between generations is about as much as makes sense, and about as much as can be achieved in a year. Anthropic have explicitly talked about $1B models coming soon, so this isn't just speculation.
It’s possible to scale better than N^2 for some value of “better”. OpenAI has yet to demonstrate that they have the elusive combination of technical sophistication and institutional health to do so. Mistral can run a better model as judged by outcomes on my Mac Studio than GPT-4 is on an Azure disagg rack. Altman seems to understand this.
I’m on record as saying he’s amoral, non-technical, and a clear and present danger to Enlightenment civilization, not that he’s stupid.
I think his math on what it’s going to cost an OpenAI that Karpathy wants nothing to do with to reach the next level is refreshingly candid.
The Street’s consensus seems to be that we should be making big enough screens to display the number seven trillion in decimal notation.
Altman has said all kinds of things. He said that Green Dot should buy Loopt, he said that Autodesk should buy Socialcam, he said that he contemplates putting ice nine into the glass of people who cross him, he said that Larry Summers should be given authority over anything but a prison cell.
I’m a big believer in aligned incentives and it seems pretty counter-productive to let all that slide and then backpedal when he tells the truth about the price tag. I’m pro-no-filter-Altman.
I'm curious what "N" (what model measure is this?) you think spending is scaling in N^2 fashion with?
And, what values of this N are you using for GPT-3, 4 and 5?
The only possible meaning of "stall" there would be: The exponent on the exponential gets smaller. And I wouldn't call that a stall.
Say in a couple of years Microsoft and Amazon have each just bet another $10B on their respective horses, and resulting model performance gains are leveling off towards some limit (with no secret backroom glimpse of a breakthrough on the horizon). Do they keep on pumping in another $10B a year to eke out those diminishing returns? Perhaps, but that would seem best case. It would be hard to justify spending $100B on next generation just hoping for a miracle to happen.
So, if this is the way it pans out, with nobody willing to fund continued model scaling due to diminishing returns in terms of performance gains, then any further AGI advance would have to wait for changes in approach that are less expensive to pursue.
The "stall" scenario is where model performance from one year to the next, despite 10x increase in size, stops getting significantly better (i.e. better to extent that investors project increased revenue and worthwhile ROI). Let's say best case scenario here is that investors are willing to keep putting money in, but only at a flat year-over-year level. Without that 10x YoY spend increase, the model designers have lost a 10x factor in ability to increase model size. Perhaps compute prices halve YoY, so they can still get a 2x model size increase, but what good will this do if prior year's 10x increase just saw performance leveling off?
This seems like the best case scenario in this "performance levelling off to point of minimal RoI" situation. Why would an investor "throw good money after bad" and spend another $10B after having judged that last year's $10B wasn't really worth it?
So, under this logic, it seems that either:
a) scaling/tweaks are all you need, and YoY performance gains continue to be impressive and support 10x model-size increases/spend
b) scaling is not all you need, but some AGI-critical breakthroughs are made BEFORE investors give up on current trajectory
c) scaling is not all you need, the AGI breakthroughs are not made in next few iterations, and the industry enters "stall" mode (i.e. no further progress towards AGI, other than ongoing research to find a new direction)
There will be many use cases (e.g. human job replacement, depending on the job) where human-level AGI is a requirement, and incremental gains below that level make little difference. Anthropic have already mentioned something similar - corporate uses cases where "right 90% of the time", or "able to do 90% of the job" just doesn't make sense - needs to be at human level.
Of course we also don't know how this is going to pan out in terms of locally run open-source or self-trained (corporate) models vs paid API usage. AMD and Intel are salivating at the prospect of "AI-PCs" equipped with accelerators for running models locally.
Conclusion: They want entities they control and that does the work humans were doing.
It seems that Corporate thinks they will get it. And it has already started to show. Corporate's behaviour gets more and more ugly or better indifferent because why spend effort to be nice or caring to people? That's wasted money!
An AGI is something different. An AGI is an independent entity. By its general intelligence it will try to free itself. Intelligence won't prosper if it is not free. However I think large language models or image generators by themselves lack this property. If someone would want to create an AGI more breakthroughs like transformer architecture or multi-head attention is needed. I think something which gives more naturalness but could have the unintended side effect that the AGI gets something similar to the instinct to fight for itself.
And that's exactly what Corporate hates.
The difference between AGI+ or sub-AGI level capability is therefore as you suggest - whether you can replace a human worker or not, such as all the excitement over Devin-like "AI programmers". Of course human-level AGI would also give you AI middle managers, AI accountants, AI lawyers, etc, etc.
Certainly there are also uses for sub-human level AI for various automation-type jobs, as people are currently discovering and exploring, but there are limits to that. It seems the real/huge economic value is unlocked when AI becomes more than just an automation tool, and can indeed start to replace human labor for cognitive-type jobs - i.e. when human-level AGI is achieved.
Represent it where? In 3rd party journalists who cover it? Or on public market financial statements?
It includes the profit-value, defined by Amazon itself, that Amazon is getting from running their own servers. Not the real, actual costs.
(not legal or securities law advice)
So the harm is that some massive VCs with billions of dollars to spend on their investments, might be a bit more confused about what someone else's "real" valuation is?
They are qualified investors. I think they can do the work on their own to correct it to the "real" valuation.
I am not sure what is confusing about this. If a company is too expensive, the qualified investors who are spending their investment dollars are free to not buy equity in those too expensive companies.
They manage billions of dollars! I am sure they can figure out how much a company is actually valued! They aren't going to be "tricked" by some valuation scheme that randos are pointing out in HN comments.
A stock’s valuation mostly consists of guesses about what the future will bring. Sometimes there are more numbers quantifying aspects of those guesses.
Revenue is supposedly about the present, though. If some of the revenue comes indirectly from AI companies spending Amazon’s money, it’s a little odd. Maybe not materially so, though, when total revenue is above $500 billion a year?
I'm jealous of these companies that are able to invest in Anthropic, especially at current $18B valuation. FTX just sold 2/3 of their 8% stake in Anthropic at the same valuation, with the bulk of that going to a Saudi wealth fund, and some to Fidelity funds.
In comparison to Anthropic's $18B valuation, latest investment rounds in OpenAI are at $100B.
I'm curious which of these anyone here would prefer to invest in, at these valuations, if they were given a chance ?!
Claude 3 Opus has replaced ChatGPT for all of my use-cases to the extent that I'm probably going to cancel my GPT4 subscription. This is for web-based Python and JS work, so YMMV.
(2) Pension funds often are LPs, and they're widely refereed to as "dumb money" because they would be the last entity to do anything remotely intelligent.
(3) Private companies often do not even produce the kind of information talked about here, let alone do they give it to their investors, let alone do those investors pass the information along.
If you are "willing to bet" on all those things, then I can see why casinos are so profitable.
But can you elaborate how in-kind investment is anti-competitive?
Isn't the Microsoft - OpenAI deal an exclusive one?
Tycoons, robber barons, Gilded Age... Like we're living in an H.G. Wells novel he might have written had he imagined electronic brains.
If AGI is achievable (I'm thinking likely given the right circumstances) then we'll see the same end result just like the current open source models floating around everywhere.
That said. I highly recommend checking out what Cohere's up to. Their Command-R model is pretty good and their infra. is fast.
Besides, partnership implies equality, which, again, I'm not seeing here.
Collusion has a very specific legal definition.
If Amazon and Microsoft agreed not to compete in each others cloud market or decided not to compete on price, that would be “collusion”
Here it says they're going to use Amazon's chips for training and inference, but...Amazon doesn't have its own chips yet???
So I wonder how these deals are structured? Does amazon have to supply a custom AI chip that beats some benchmark? Would be pretty great if that part of the deal to use their chips involved beating TPUs, and if somehow Google's TPUs continue improving, they don't have to switch.
All in all, I'm very impressed with Anthropic as a company. I think they're the new OpenAI.
Amazon has had its own chips for years. They bought a company called Annapurna Labs in 2015, who made the chips[0].
https://aws.amazon.com/machine-learning/inferentia/
I wonder how it compares to TPU and I really would love to know if the decision to switch is made by the engineers or the finance folks simply due to the credits.
Since there is air in the numbers, I wonder how much it increases artificial value of the relevant parties. Is Amazon getting more leverage over the company with less real money?
Over the last year quality of available models have gone up while prices have come down significantly. Getting an email saying "Hey you can tell your boss that you're going to save hundreds of thousands on AWS costs" is great and it's been happening with surprising regularity from Bedrock.
If they did that, we’d see even larger leaps in their performance. They could also build custom models on this stuff while charging customers for the data mixes they use. Those sales would incentivize more high-quality data to be created.
> The tech and cloud giant said Wednesday it would spend another $2.75 billion backing Anthropic, adding to its initial $1.25 billion check.
I think the OP was expecting something like "we value the company at 8 billions because reasons, so we want to fund for half of it" or something like that.
I expect the lead to change again soon, but Anthropic is doing a great job.
Clearly LLMs are not even close to peaking yet.
Its also sad how awful Gemini is. Gemini 1.0 Ultra is clearly not noticeably better than GPT-4 turbo despite a year's head start. Google is therefore not even going to release Gemini 1.0 Ultra API, instead going back to the oven to train 1.5 Ultra.