OpenAI Status: Multiple engines are down
status.openai.com
status.openai.com
For perspective the likes of Instagram and Twitter have take from a few months to a few dozen months to get to a million users all while doing less work to service a request for a "page."
Hats of to OpenAI for not falling over more often under a hug of death that doesn't seem to be letting up.
[1] https://twitter.com/sama/status/1599668808285028353 [2] https://www.afr.com/technology/chatgpt-takes-the-internet-by...
At Google, the Cloud SQL dashboard was unavailable for around 12 hours a couple of weeks ago if I read this correctly: https://status.cloud.google.com/incidents/xg2qrL1UuSJiPDZALJ... The total number of employees of Google are 156 500. Google Cloud was launched 2008.
So when people say scaling is hard... it's not a solved problem and you shouldn't be surprised when these things happen.
A solid theoretical foundation and more testing is better of course.
For additional perspective, OpenAI has a very small number of highly similar offerings (neural net with API access), while Google has a huge host of very different offerings, from web indexing and search, to email, to file hosting, to video streaming, to cloud compute... etc. Google also has a vastly larger user pool by at least a couple orders of magnitude. Google's core services are extremely solid even at ridiculous scale and have few outages, any of which would be considered major news.
Google also operates all this at a profit, while OpenAI works at a deficit. Google has had to scale larger than nearly any other service while maintaining profitability while OpenAI is more or less free to throw more compute power to solve problems, at any cost, to build valuation. Google has written entire programming languages to help them keep up at an unprecedented scale.
Comparing a single minor Google offering going down to an OpenAI outage isn't a fair comparison to either company. Yes, Google has, by your numbers, about 400 times the number of OpenAI employees. I'd be willing to bet that a single large Google service like YouTube handles more than 400 times the amount of compute, data, and traffic that OpenAI does. I wouldn't draw a comparison between the size and efficacy of the employees, but again, Google operates at a very different scale, and has operated at that scale very well.
Also anecdotally, half the time I've tried to use ChatGPT it's "at capacity" or throws an internal error, and I've also seen Dall-E unavailable even though I've barely tried to use it, so I wouldn't say that OpenAI service has been ironclad this whole time.
The Google SRE book is an excellent work: https://sre.google/sre-book/introduction/
Also some perspective, if you look at the OpenAPI outages for Jan 2023, the up time is worse than Google Cloud SQL.
Some chapters in the book include "Being On-Call", "Effective Troubleshooting" & "Postmortem Culture: Learning from Failure".
Which signals to me, a normal work culture and time will improve stability in most companies.
I once spent an entire weekend swapping shifts trying to keep the main server of an AI company up because the devs deployed on friday and were unreachable to fix their memory error. The best thing that happened is the boss came in Monday and froze feature development for a sprint, inventing the "tech debt sprint" that became a quarterly activity. Everyone loved it, uptime became better, and even application development got easier.
Size of a company doesn't matter, the magnitude of product offerings aligns with head count. Team size is what matters and is likely a much closer ratio between the companies. Google does have head count and dedicated teams to solve common problems and build solutions at the platform layer. Likely OpenAPI is benefiting from Google's efforts here (from the open source and knowledge sharing which are bar none imho).
From an OpenAPI job opening:
> The Engineering team wraps a massive fleet of GPUs in scalable, robust, infrastructure powered by Kubernetes, Go, Python, Terraform, Kafka, Postgres, and Snowflake. Our APIs are powered by Python Flask and OpenAPI with a React frontend.
Anyone care to elaborate on why they have snowflake in their stack?
As an exercise, which is likely to reach the SRE first?
Yes, but how likely is this unlikely situation?
It took 12 hours to debug and fix after all.
All I'm saying is that some outages are to be expected from a young company and they are doing excellent work.
I believe trust in information is a bigger concern than global warming. As an aside, what is OpenAI's relative contribution to CO2 emissions compared with other companies of similar size?
"likely" is your hypothesis. They are two different orgs, and internal services can use different infra for analysis.
I am available for hire to teach a thing or two to Google (and OpenAI).
/s
they outsource infra to MS?
I mean I had outages that were 60 seconds long with magnitude less people than OpenAI and serving more req/s than they do. What does it tell you? Nothing at all.
So what this tells us, based on very limited data, is that this level out outage happens to even the biggest of companies.
> plan accordingly
You can't plan for every major issue unless you solved the halting problem and unless you as customer want to pay magnitudes more for services you use. For every major client facing issue you probably have 10s/100s/1000s (depending on the system complexity) incidents that were prevented and you know nothing about.
This is a non-sequitur. This statement has nothing to do with the outage. It’s just thrown in there to pad the number of words in the comment.
Of course, upon trying the next day, all was fine again and I was no longer able to reproduce.
I imagine the huge demand that ChatGPT is seeing would make any cloud vendor sweat if you were to suddenly lump it on top of the usual demand.
To me it's entirely unsurprising that OpenAI would have trouble keeping up. Good luck to them.
A few years ago, we could maybe lean on their open aspirations to get that done, but with the "limited profitibility model" they've since instead adopted, I think that dream is mostly gone.
At least we still get the occasional treat like Whisper out of them.
And a few dozen like minded people banding up together, shelling out a couple grands each? I'd say that's a totally realistic scenario.
But I certainly agree that not everyone can, and very few individuals.
The thing is a monster of a model.
I don't get HN's take with wanting everything open sourced. Some things are expensive to create and dangerous in the wrong hands. Not everything can and should be open sourced.
Actually, thinking about my own question I'm even inclined to remove the non-weapons qualifier. The most knee jerk response, nuclear weapons, is perhaps the best example of unexpected benefit. The 'decentralization' of nuclear weapons is undoubtedly why the Cold War was the Cold War, and not World War 3. And similarly why we haven't* seen an open war between nations with nuclear weapons. One power to rule over all suddenly turned into "war with this country no longer has a win scenario" effectively ending open warfare between nuclear nations.
There's also the inevitability/optics argument. There are already viable open source alternatives [1], and should this tech ultimately prove viable/useful that will only be the beginning. So there certainly will be "ai" that will be open, it just won't come from OpenAI(tm)(c).
If ML models continue their exponential growth in size, a similar outcome is possible.
I see a similar line of reasoning can be used to justify theft from the rich.
----
"OpenAI is a non-profit artificial intelligence research company. Our goal is to advance digital intelligence in the way that is most likely to benefit humanity as a whole, unconstrained by a need to generate financial return. Since our research is free from financial obligations, we can better focus on a positive human impact.
...
As a non-profit, our aim is to build value for everyone rather than shareholders. Researchers will be strongly encouraged to publish their work, whether as papers, blog posts, or code, and our patents (if any) will be shared with the world. We’ll freely collaborate with others across many institutions and expect to work with companies to research and deploy new technologies."
----
My mocking about OpenAI(tm)(c) was not just juvenile "Micro$oft" type nonsense. At some point they discovered they could make a buck, and their ideology suddenly shifted 180. I have no qualms whatsoever about businesses pursuing profit, but the entity currently known as OpenAI couldn't be much further from the principles and values OpenAI was founded on, and their name itself is rapidly trending towards becoming a "Don't Be Evil" type of sardonicism. If this was Microsoft, Google, or other such companies operating in this way - I wouldn't have any expectation of anything besides what OpenAI is doing.
Companies tend to get quite a lot of credit when claiming some socially motivated interest, probably much more than deserved. So when they turn against those ideals, it should be noted - loudly.
Ironic taking into consideration that the current generation of AI are more or less copyright laundering for the big corporations. Github Copilot being an extreme example of using GPL projects to generate "proprietary" closed source code. What happened to ownership and respecting the effort it takes to create something?
Here, instead of common land, what we have is the common content. And they're saying that, by "developing" that content into a model that can do more useful things, the authors of the model are entitled to full private property rights on it.
I really hope that's not where we're going to end up, legally speaking.
If you can find an economical way of running it though, let me know.
Amazon has them listed at $200, but still, that's only $2,400 for 12 of them.
Still, adds up once you get the hardware you'd need to NVlink 12 of them, and then on top of that, the price of power/perf you get probably isn't great compared to modern compute.
Wonder what your volume would have to be before getting a box with 8 A100's from Lambdalabs would be the better tradeoff.
But yes, it's a very good value.
Why is this? Did some large cloud vendor just upgrade?
Are there any deals like this on AMD hardware? Not having to deal with proprietary binary drivers is worth a lot of money and reduced performance to me. A lot.
These are pretty old, and all the companies are upgrading. But no one is upgrading from AMD hardware - basically no companies care if they use proprietary drivers. They want a good price-to-performance ratio, so they use NVIDIA stuff.
Plus, everyone wants CUDA.
The wrong hands have the money to seek alternatives. All this policy does is keep it out of the hands of the public, and ensure that whatever open alternatives start up won't be OpenAI's.
What is amazing, human language (languages?) and knowledge encoded in so little space.
To put that in perspective, 24x 64 GB nodes is 1.5 TB.
Looking at your calculations indicates that you mean RAM but it's 1.5 TB GPU VRAM (but this is assuming they use 64 bit precision, which is likely wrong so it's ~750 GB), not RAM.
Nobody is running these large scale models on their personal devices.
Sure, some of the image generation tech is seeing personal use, so you'd have a point there, but these immense language models are something else entirely.
It's a whole lot cheaper to run neural net style systems than to train them. "Somebody on Twitter"[2] got it setup, and broke down the costs, demonstrated some prompts, and what not. Cliff notes being a fraction of a penny per query, with each taking about 16s to generate. The output's pretty terrible, but it's unclear to me whether that's inherent or a result of priority. I expect OpenAI spent a lot of manpower on supervised training, whereas this system probably had minimal, especially in English (it's from a Chinese university).
Who knows though, maybe someone manages to get 4byte quantization producing good results and Apple makes a chip that can do the required ops for whatever that looks like with ~100GB of memory attached and then GP's comment might be relevant.
I started this reply rather skeptical, but with the boundaries Apple has been pushing, and the pace of AI research, honestly who knows.
Even in search I think it will, at most, be a sort of sidebox that says 'Super clippy says the answer to your query is [blah].' Because a single opaque source of information, which will continue to struggle with truthfulness (both inadvertently, and by design) is really just going to supplement endless dynamic content on a topic.
I'm just not seeing the big use for this (outside of the brief period of 'wow' novelty) in anything remotely like its current state.
Based on experience of BERT , yes maybe to get the best experience or to serve millions of users you need to run any model on compute intensive infrastructure , BUT if you just want to run for yourself and do some small testing you can very well download it from huggingface and elsewhere and run it on your laptop.
"The Python code in this tutorial generates one token every 3 minutes on a computer with an i5 11gen processor, 16GB of RAM, and a Samsung 980 PRO NVME..."
[1] https://towardsdatascience.com/run-bloom-the-largest-open-ac...
The amount of memory required to run these models is immense.
If you want a comparison, try running the largest version of the open source BLOOM model yourself.
Spam sites are still spam sites. The only exception are sites like CNET co-opting the tools to seed AI in combination with their own real bloggers (saying 'journalist' would be a stretch here).
These still require reputable people/business to be behind them directly to give credibility, people that can be blacklisted. It's not like turning an AI bot loose and instantly you have 1000x fake Twitter accounts with 100k followers and top ranked on Google.
The "long tail" of spam sites do not threaten governments.
China cares most about reputable news sites and popular social media influencers going viral with taboo stories. I'm skeptical news sites or influencers can be generated artificially merely using AI, in a way that is scalable enough to actually be a problem for censors. Ultimately the censors just need to ban the source... the domains or accounts with 20-100k+ followers or news sites with actual reach (ultimately a small set of sources). Fake AI content doesn't directly solve the existing reputation signals.
The more I think about it the less I'm convinced this is a real threat to the general reputation networks.
That said - where this does matter is the typical grey/black market where scams already operate. Not in highly public reputable places like the top rankings on Google or Twitter or w/e popular Chinese social network. It's the hackers exploiting email lists, Indian call centers running combination IM/phone scams, etc. It's the low-end of the market for criminal get-rich-quick schemes, not something that generates popular movements that threaten governments or popular political ideologies.
The idea that censorship is a weapon of authority is a pre information age idea when signal-to-noise was high.
Everybody, as far as anyone could tell, in complete agreement with the regime, all the time.
Signing and encryption will come in useful. But not sure of much good it will do under an authoritarian regime.
Truly remarkable.
I wonder how OpenAI will cope on it’s own… unless Stability strikes again, with it’s own LLM?
What is your outlook on OpenAI, the company? Will they be successful in 10 years?
There are other huge models like Bloom that are more general purpose, but they require greater than 24GB of ram which means multiple linked GPUs to even run them much less train them.
Thus, unlike stable diffusion, chatbots that run on common consumer GPUs are going to need some work to be competitive with commercial offerings.
It takes an additional fortune to pay humans to do RLHF (reinforcement learning from human feedback).
It takes an additional fortune to serve the model.
There are also researchers getting paid there that live in one of the most expensive hubs in the world.
7 min was the third time I'd read your comment so I downvoted it.
Can't speak for anyone else but aye it's definitely because you posted the same thing thrice for me. The first version (topmost rated) of your comment doesn't seem to be downvoted
"Okay guys, the API is down, you have until it's back up to talk to me about pink pigeons (or any other unpredictable topic of the referees choice you couldn't batch responses to ahead of time) to raise my confidence that you are in fact not a bot"