ChatGPT could cost over $700k per day to operate
businessinsider.com
businessinsider.com
If the 100m users is accurate, it means they only need to convert low single digit percentage to paying users to break even.
Now it makes sense why they chose to charge $20/m, I predicted much higher.
It has proven to be a great way to get everything to around 70%, then send off to my assistant for the remaining 30% of polish. So at $20/month, it was such a no-brainer, that I had to do it. Even at $75/mo it would more than pay for itself.
It even understands the concept of a "shit post". So - its more social media savvy than I am, thats for sure.
It would benefit her immensely to use it.
#include <windows.h>
then it is a must have.
I’ve already incorporated it into my publishing process. My home-grown “cms” uses ChatGPT (via API) to write my article description, draft a twitter thread, and craft a “viral insight”. The latter is mostly useful to make sure my article even makes a point.
Hoping to use it for a related articles feature next. I’m also building a chatbot based on my content.
Yes. i think its a good software product! Maybe one of the best there ever was - I'm not surprised Sam Altman jumped the YC ship for this one, or that MS is particularly interested it. It gets ignored in all the other bru-haha but I'm excited to see what the really talented software teams of the world will be able to accomplish with an AI coding assisstent - and I don't mean just in the world of AI.
The activation energy for new software lowered by about 30% overnight, which is outrageously cool, and of course disquieting.
For text generation, it is much better for tasks like
* letter writing * rewriting my writing so i don't plaigerize myself * summarizing several paragraphs into one
but all of that is available in the free version.
If you're like me you tend to forget you're paying for something, and have to do a yearly purge.
It's very easy to slowly but surely rack up hundreds of dollars a month in subscription services.
- Payment cycle closes always on the 1st of the month (which means, if you sign up on the last day of the month, you get 1 day of service for full month of payment. No proration, or anything.)
- If you cancel the service, you'll immediately lose access, doesn't matter if it's 1st day of the month, or very end of the month, it's gone. To reenroll, you often have to pay again.
Probably for the next proposal I write, I'll pay for it again. It's super useful to take care of all the bullshit things you have to write for science to not plaigerize yourself.
I also tried using it as a dungeon master like that blog post from a couple weeks ago. But gpt4 didn't seem to remember things with regularity enough to actually work. Basically there was an uncanny valley because gpt4 doesn't have a pad of paper to write things down on like any person would have.
Well, it works until it doesn't because the website is overloaded. $20 gets you in through the VIP door and you don't need to wait along with the rest of the peasants.
> Expert mode is our most advanced searching mode, powered by GPT-4. This mode hallucinates less and writes better code. We highly recommend that you use it for advanced questions. Whenever the "regenerate" button is pressed, Expert mode is used to increase the odds of a high-quality answer.
Yeah, what's wrong with that? I'm sure it would be catchy to a particular demographic. No need for a brand to cater to everyone.
Because GPT is capable of: doing marketing, math, science stuff, programming and anything else to a junior-semi advanced level, it needs to be prompted and given context to get it into the right "mood" for doing what you want it to.
There's a huge difference between casually chatting to it about various topics before suddenly asking what a good perfume brand would be. Versus providing it with the fragrances used in the perfumes that the brand makes, the general goals/image that the company wants to set, their target demographic etc and _then_ asking for a brand name suggestion.
I'd be keen to see how they asked.
better source : https://tech.slashdot.org/story/23/04/18/1326235/microsoft-r...
Yup. I bought accounts for everyone here. We are using it as what I have been calling a "force multiplier". We are not and cannot use it for coding, yet, things like presentations, analyzing logs, creating lists of things, researching topics, etc. It's a great time saver.
Also, for a lot of things ChatGPT is a much better search engine than Google. It gets you great answers and almost always the first time you ask.
In case the question comes up: We don't use it for coding because of potential liability concerns. At this point I feel that is a space that has not been explored at all. I have no interest in being a pioneer in a lawsuit that claims negligence due to the use of AI-generated code, even if it is reviewed by a human. The combination of fear-mongering and tech-challenged juries could make for very expensive outcomes.
I posed a question about this here:
Where do we think Apple is in regards to this? If anyone has the upper-hand here with hardware, I would think it'd be Apple. But there's been zero indication that Apple has been working on any sort of generative AI.
But google is fumbling so hard right now it's not surprising that they are squandering this too.
At the end of the day, all we are doing is create, read, update, and delete data.
I mean if you take away all the complexity associated with ranking and scale (not everything has to be Google scale), that is exactly what a web search is as well, right?
I remember reading a post by a maintainer saying their job is to manipulate strings or something to that effect. iirc it was a gofmt maintainer who said that but I can't find that post now.
Now M$ is planning to create a specialized chip for AI which comes with its own R&D budget and ongoing costs.
If it proves successful, ChatGPT will become a household brand and M$ could easily ask for $500 per month or more for professional/corporate usage.
it's harder to type than just "MS" so when someone uses "M$" they go out of their way to signal that they are biased. Being biased is fine, so long as one understands that they are biased and that they are communicating their bias along with the rest of the message.
Users or accounts? I made an account but I use it approximately once every two weeks for about 30 minutes at a time, as it hasn't been that useful for me (needs more up-to-date info after September 2021).
I imagine many people made an account but only a small percentage are using it meaningfully often.
Also users are limited in the number of requests they can make. I have a feeling ChatGPT is actually very expensive to run, and they are burning cash like crazy.
If they weren't burning cash to run the thing, it would be more widely available.
The company also has about 375 employees. I've no idea how much they get paid but I used $200k as a yearly cost and that comes to $75 million.
That's about 3:1 cost of operating the services to paying employees. That seems quite high as I've never been at a company that had 1:1 costs for running servers vs employee costs but I could entirely be off base here.
Given Sam Altman's recent comments on the days of these LLM being over I think maybe Microsoft or whomever is basically saying that they can't spend that much money and they need to control costs much more heavily.
What I am wondering, for those earning 500k, how big is your work load/stress. Would this be a 9-5 job you leave at the office when going home. Or does a job that earns so much consume your life?
I would say 50% the work is harder and consuming and then 50% they can just afford to pay you more and lock up talent because of the wild margins on their products.
Note that there's likely to be some variation per team, but Amazon is famously bad, so ... ;)
I've said it before on here, but I live very comfortably in Philly for a lot less than that.
If interest rates stay elevated, and value investing becomes valuable again, it will be interesting to see how the tech space transforms. When start-ups have to compete with money market funds or even treasuries for investor cash, things become orders of magnitude tighter and more challenging.
Yes, though Switzerland approaches it. If you want to see how much people of various levels of experience get paid at different companies and in different locations go to levels.fyi
Americans get paid much, much more than anyone else.
I am seriously considering a move if my husband can find an academic job over there. The retirement won't be a great lure (fewer years in the system) but we almost have enough to coast from here, so it's about the rest.
I was expecting salaries to cool off a bit with the massive wave of layoffs across the industry, but from what I've seen, that hasn't happened.
~100k€ in (western) Europe may be comparable to ~200k€ in Bay Area.
I'm 20 years into programming and a senior architect and lead on an enterprise project.
I don't even make that first number.
But I value certain things way more than other things, and my current job provides it. Fully remote, leaves me completely alone to accomplish what they need done (and I get it done), unlimited vacation, great benefits, zero pointless meetings (almost an empty calendar).
I'm sure these other companies offer some of that but 500k?! That is absurd.
> ChatGPT could cost OpenAI up to $700,000 a day to run due to "expensive servers," an analyst told The Information.
which, pardon me, but no shit.
Before I break out my back of the envelope calculator, on how many biggest GPU instances in Azure that is, the real question is what their underlying assumptions are, and where they're getting them from. Especially since OpenAI is definitely not paying list price on those GPU instances.
The other question is how close to capacity their cluster is running, and how much free time on it can be reclaimed, either in terms of spinning down servers for diurnal patterns, or in terms of being able to do training runs in the background.
Such "free usage" coupons are marketing activities to gain new customers, Microsoft already completed the "dating phase" with OpenAI. They surely don't pay list-price for Azure but it's surely also not free.
Moreover, as per Microsoft themselves, the 1bn USD investment into OpenAI carried the condition that Azure becomes the exclusive provider for cloud-services: https://blogs.microsoft.com/blog/2023/01/23/microsoftandopen...
It's not exclusive because it's free, it's exclusive because "we paid you 1bn USD to buy it from us"
They might not give unfettered credits, it could be for specific projects. That said, I wouldn't be surprised if it was unfettered either.
I'm not sure about the most recent $10 billion investment but I wouldn't be surprised if a significant amount of it is in Azure credits as well.
While that's not "free" (they exchanged equity for it), it's likely not an expense (or at least not an expense that they have to cover fully).
> While training ChatGPT's large language models likely costs tens of millions of dollars, operational expenses, or inference costs, "far exceed training costs when deploying a model at any reasonable scale," Patel and Afzal Ahmad, another analyst at SemiAnalysis, told Forbes. "In fact, the costs to inference ChatGPT exceed the training costs on a weekly basis," they said.
(Looking back, I'm happy that I was careful in my wording in that I didn't say diurnal cycles aren't relevant, just that they aren't as useful in this case)
That said, I suppose I misread the specific suggestion about spinning down servers off-peak and was thinking more about spot pricing at peak vs trough.
He sees size as a false measurement of model quality and compares it to the chip speed races we used to see. “I think there’s been way too much focus on parameter count, maybe parameter count will trend up for sure. But this reminds me a lot of the gigahertz race in chips in the 1990s and 2000s, where everybody was trying to point to a big number,” Altman said.
As he points out, today we have much more powerful chips running our iPhones, yet we have no idea for the most part how fast they are, only that they do the job well. “I think it’s important that what we keep the focus on is rapidly increasing capability. And if there’s some reason that parameter count should decrease over time, or we should have multiple models working together, each of which are smaller, we would do that. What we want to deliver to the world is the most capable and useful and safe models. We are not here to jerk ourselves off about parameter count,” he said."
via https://techcrunch.com/2023/04/14/sam-altman-size-of-llms-wo...
What he actually said was that we've reached the point where we can't improve model quality simply by increasing its size (number of parameters). We'll need new techniques to continue to improve.
I'll start:
- making little scripts in shell / js / python that I'm not as fluent in. 5 min vs 1-3 hours
- explaining repos and apis instead of reading all the docs - help with debugging
- flushing out angles for new concepts that I did not previously consider (ex: how do you make a good decentralized exchange)
Obviously I use it for other purposes as well, but it definitely has saved me a lot of hours getting the basics things right there in a prompt.
You can only do this with something that existed in its training set, right? There's no way to point it at a random GitHub repo and say "digest this, so I can ask you questions about it"?
https://help.openai.com/en/articles/7127966-what-is-the-diff...
Wealthy individuals these days hardly need any intelligence at all to stay rich. On the other hand, it seems like poor individuals have no chance no matter how much intelligence they have... Does anyone really need more intelligence? Most human intelligence seems to be wasted on bullshit jobs anyway.
What we should be asking ourselves is "how much more comfortable can we make our billionaires?" Because the entire concept of 'economic efficiency' appears to be about optimizing society towards that goal.
Wasn't it always more or less like this?
Funny to think all that was a generation ago, that the economical conditions for our artifice are now history and it's once again economical to study the much more all-encompassing art of not thinking. (It's very cyclical but, since it involves not thinking, i.e. the termination of a coherent thread of thought, you think it's for the first time every time. But sure, plug in to the simulation at 110% and get 10 years of life for free, you won't even notice the difference, after all, you're already not noticing it.)
This is how I, a unit of human capital from the former global East, see the matter of intelligence and automating it at scale: "intelligence" is already something that was forced onto my society in the 1950s, and once again on the 1990s, so that we could be interfaced with the same kinds of systems that the conquering populations exported, under conditions of forcible termination of local culture at the cost of great atrocities. By the time we started building our own Borg cube we'd already been half Borged by the Western one. If the word in some cases denotes another kind of abstract human "intelligence", I have rarely if ever seen such cases exemplified in public. So now it's a bit distasteful to equate "intelligence" with "savoir faire". Having neuromorphic computing in the first place would've saved a lot of suffering in comparison with how we were "civilized" to support only the established market protocols and nothing of our own.
Which is something someone will eventually have to admit - probably after a Pontypool-class scenario wipes out their investment in predictable outcomes for a region and they're left with a huge mess on their hands to mop up. And then it turns out the consequences themselves have been thought in such inhumanly fine detail as to possess a rudimentary intelligence of their own and the mop up ends up embedding a sentient malevolent ghost into the nature of that branch of reality. Shit happens, at a certain saturation with Internet stream of thought an omnipresent autoritarian AI even becomes fun to imagine. Of course if one existed its cruelty would be much more banal, as is always the difference between mentalization and reality until the next brief Golden Age (for some of the people, some of the time) and then it's "oh no this time we're fucked" again (for some other people).
So sad because there are so many great ideas worth building and fighting for, even in the current limits of human knowledge, but in practice any form of innovation is only tolerated if it's backed by a major financial effort by the upholders of the status quo. For some reason so many times this has ended to the detriment of nonparticipants that we've basically stopped noticing or keeping count. As regulatory capture, so learned helplessness. The economic reality of already existing in a world of inscrutable mechanisms as that most fragile of creatures, the independent knowledge worker, is a psychological stressor, to the extent that everyone's so burnt by the effort on giving up on things that in this new ecosystem are "too smart to work". Help, the AI stole my role in the food chain and now all my thoughts lead nowhere.
If that's you, look for a bigger room.
And then the smartest person on the planet figures out a better thing for the people who lost their "bullshit jobs" to do, right? No, gets sent to space to see how much even the smartest person on the planet really matters.
> There’s always a market for intelligence, whatever that means. The word intelligence is misleading anyway.
Agreed. What is decided here is to what extent there will continue to be a market for human intelligence. There are of course opportunity costs that humans who entered a market that they apparently believed to exist, will now have to, apparently, absorb. Or will we once again have to voluntarily undergo unsupervised radical mental restructuring - on the degree of your grandma having to learn how to navigate the Internet, and of her grandma presumably never having to do anything of the sort - and all that just to retain our Red Queen out of checkmate? It's always too early until it's too late.
>I’ve never heard a more cynical AI take.
Sapienti sat.
Like people slowly getting more obese, the general IQ and CT have been sliding for the last few decades. Not sure if it’s the food, media or zanax but society needs something.
I’d wager the Altmans of the world know this and figured out a way to monetize it. Necessity is the mother of invention after all.
Let's say 500 employees, earning above Bay Area average, along with benefits etc at $400K average per year. That's $200 million a year in wages, so in the same ballpark as the $200 million in infrastructure costs cited in the article, for a total of $400 million in expenses per year.
I would imagine that Google's reputation as a company where it is impossible to ever talk to a human (even when you have a 7 figure annual spend) hurts them in this space.
Once people already are buying 20 products from you and have a good sales relationship of decades, selling them some cloud services is easy and might even lead to better deals on something else (Windows/Office)
The problem for both of them is AWS, which somehow manages to give you your cake and let you eat it too, even if it's a little more expensive.
Sometimes it's better to be the fast follower.
For example, why not cache user prompt/response pairs and use cosine distance for key lookup? You could probably find a way to do this at the edge without a whole lot of suffering. I suspect many of the prompts that hit the public API are effectively the same thing over and over every day. Why let that kind of traffic touch the expensive part of the architecture?
The point I am trying to make is that not all use cases for ChatGPT are generative. There are a lot of Q&A use cases today despite the fact that these are so far beneath its true capabilities. These items could be dealt with using more economical means.
"give me a recipe for XYZ" should not require a GPU for the first turn response, much like typing in an offensive manifesto returns a boilerplate "as an AI language model..." response.
Granted, if the user then types something like "please translate the recipe to Spanish and increase the amounts by 33%", we would have to reach for the generative model. But, how many real-world users are satisfied with some simple 1-turn response and go about their day?
You'd basically be removing the entire cache every release
because of context. You can cache it if it is indeed the first sentence of a fresh dialogue, but that's it.
I imagine people will be using this to create wasteful products, but also green solutions.
Personally, I've used it with success with R&D and due dil of projects related to climate change. LLMs (and progress in general) can help tremendously with switching to a sustainable economy.
I kid, I kid.. or do I?
You Z80 computer cost $700 in the lat 70's...they're now in sub-$1 embedded controllers.
10% here and there is very small compared to the literal orders magnitude improvements during the reign of Moore's Law.
I don't really see anything like that here.
I can't confirm it, but I noticed this comment says "gpu tech has beat Moore’s law for DNNs the last several years":
For a long time, GPU hardware basically became more powerful with each generation, but prices stayed roughly the same plus minus inflation. Last couple of years, this trend has broken. You pay double or even quadruple the price for a relatively tenuous increase in performance.
You get the point.
There's always local optimization that leads to improvements. Look at the Apple M1 chip rollout as a prime example of that. Big/Little processors, on die RAM, shared memory with the GPU and Neural Engine, power integration with the OS.
LOTS of things that led to a big leap forward.
You can re-arrange them for minor boosts, double the performance a few times sure, but that's not a sustained improvement month upon month like we have in the past.
As anyone who has ever optimized code will attest, optimization within fixed constraints typically hits diminishing returns very quickly. You have to work harder and harder for every win, and the wins get smaller and smaller.
However, none of that is actually important when the thing people care about most right now is energy consumed per operation.
This metric dominates for anything battery powered for obvious reasons; less obvious to most is that it's also important for data centres where all the components need to be spread out so the air con can keep them from being damaged by their own heat.
I've noticed a few times where people have made unflattering comparisons between AI and cryptocurrency. One of the few that I would agree with is the power requirements are basically "as much as you can".
Because of that:
> double the performance a few times sure, but that's not a sustained improvement month upon month like we have in the past.
"Doubling a few times" is still huge, even if energy efficiency was perfectly tied to feature size.
But as I said before, the maximum limit for energy efficiency is in the order of a billion-fold, not the x900 limit in areal density, and even our own brains (which have the extra cost of being made of living cells that need to stay that way) are an existential proof it's possible to be tens of thousands of times more energy efficient.
Ditto with mobile phones. iPhone may be more expensive than when it launched, but you can buy dirt-cheap chinese smartphones that have similar performance - if not higher to the first iPhones.
Missing the point, despite being internally correct: 10% of $700k/day is still $25M/y.
If you'd instead looked at my point about energy cost per operation, there's room for something like 46,000 improvement just to human level, and 5.3e9 to the Landauer limit.
How is it not?
These LLMs were recently trained using NVidia A100 GPUs.
Now NVidia has H100 GPUs.
The H100 is up to nine times faster for AI training and 30 times faster for inference than the A100.
This is often the case with these types of technologies.
To not look far - gpt3.5 turbo.