March 20 ChatGPT outage: Here’s what happened
openai.com
openai.com
>We took ChatGPT offline earlier this week due to a bug in an open-source library which allowed some users to see titles from another active user’s chat history.
Blaming an open-source library for a fault in closed-source product is simply unfair. The MIT licensed dependency explicitly comes without any warranties. After all, the bug went unnoticed until ChatGPT put it under pressure, and it was ChatGPT that failed to rule out the bug in their release QA.
In that case, though, they could say "look, we used a popular open source library because we had more faith that it would be better tested and correct" which would be a compliment to open source. That's essentially the information that we have.
In today's world, who builds anything from anything close to stratch? embedded developers, probably come closest. It's no worse or better to say "there was a bug that our release uncovered." If they continue to announce as many details as possible, we as the audience can develop a sense whether they're creating bugs or just uncovering bugs we're glad to know about.
"Release early, release often" and "move fast and break things" are respected ideas that diminish the importance of QA. The more effective and efficient any QA you do is, yes, very valuable, and QA that finds broken things, all the better. But don't move slow is an OK compromise.
I read this post, and I don't see it assigning blame to anyone other than themselves. See this bit copied from the post:
> Everyone at OpenAI is committed to protecting our users’ privacy and keeping their data safe. It’s a responsibility we take incredibly seriously. Unfortunately, this week we fell short of that commitment, and of our users’ expectations. We apologize again to our users and to the entire ChatGPT community and will work diligently to rebuild trust.
It’s a great opening sentence. Short, factual explanatory. You would have to be hypersensitive to object.
As far as I can see, the whole thing is a reaction to OpenAI not being sufficiently open - which is a fine argument to have with them, but shouldn’t cloud judgement on the quality of this write-up
"We took ChatGPT offline earlier this week due to a bug in an external library which allowed some users to see titles from another active user’s chat history."
Bugs will always exist whether you have a QA dept or not.
The only counterpoint I would take is one related to the accuracy of what they wrote
"The bug was a result of faulty code produced by ChatGPT."
Using an open source library is like copy and pasting code into your project. You assume responsibility for any pitfalls.
It's going to be recursive blaming all the way down.
To be clear, I would be up in arms if OpenAI was trying to hold redis potentially legally responsible.
> between 1 a.m. and 10 a.m. Pacific time.
Oh... so it was because they're based in San Francisco. Do they really not have a 24/7 SRE on-call rotation? Given the size of their funding, and the number of users they have, there is really no excuse not to at least have some basic monitoring system in place for this (although it's true that, ironically, this particular class of bug is difficult to detect in a monitoring system that doesn't explicitly check for it, despite being immediately obvious to a human observer).
Perhaps they should consider opening an office in Europe, or hiring remotely, at least for security roles. Or maybe they could have GPT-4 keep an eye on the site!
How do you figure? If you mean there are few SRE with several years of experience you might be right. SRE is a fairly new title so that's not too surprising.
However, my experience with a recent job search is that most companies aren't hiring SRE right now because they consider reliability a luxury. In fact, I was search of a new SRE position because I was laid off for that very reason.
However, I think the GP's point about this class of bug being difficult to detect in a monitoring system is the more salient issue.
And I do think the answer is that monitoring is easy but good monitoring takes a whole lot of work. Devops teams tend to get to sufficient observability where a SRE team should be dedicating its time to engineering great observability because the SRE team is not being pushed by product to deliver features. A functional org will protect SRE teams from that pressure, a great one will allow the SRE team to apply counter-pressure from the reliability and non-functional perspective to the product perspective. This equilibrium is ideal because it allows speed but keeps a tight leash on tech debt by developing rigor around what is too fast or too many errors or whatever your relevant metrics are.
On a personal note, I am sorry to hear that your job search has not yet been fruitful. Presumably I am interested in different criteria from you —- I have found several postings that are quite appealing to the point where I am updating my CV and applying, despite being weakly motivated at the moment.
All you need on day 1 is someone to watch the (metaphorical) phones + a way to page an engineer. Don't start by spending a million bucks a year, start by having a first aid kit at the ready.
Perhaps they could also help this person out by looking into some sort of fancy software to automatically summarize messages that were being sent to them, or their mentions on Reddit, or something, even?
Looking at my post up-thread, I wish I had emphasized the time aspect more - of course all of these problems are solvable but it takes both time and money. They have the money now but two months ago the parts of this incident were in place but the scale was so small that it never actually leaked data. Or maybe a handful of early adopters saw some weird shit but we’re all well-trained to just hit refresh these days. Hiring even one operator and getting them spun up takes calendar time that simply has not existed yet. I assume someone over there is panicking about this and trying to get someone hired to make sure they look better prepared next time, because there will be a next time, and if they’re even half as successful as the early hype leads me to believe, I expect they are going to have a lot more incidents as they scale. One in a million is eight and a half times per day at 100 rps.
Since I wrote this, I have seen several anecdotes that support this guess. This is a classic scaling problem. One or two users saw it, and one even says they reported it, but at small scale with immature tools and processes getting to the actual software bug is a major effort that has to be balanced around other priorities like making excessive amounts of money.
OpenAI is hiring Site Reliability Engineers (SRE) in case you, or anyone you know, is interested in working for them: https://openai.com/careers/it-engineer-sre . Unfortunately, the job is an onsite role that requires 5 days a week in their San Francisco office, so they do not appear to be planning to have a 24/7 on-call rotation any time soon.
Too bad because I could support them in APAC (from Japan).
Over 10 years of industry experience, if anyone is interested.
Edit to add: they’re also paying in-line or below industry and the role description reads like a technical project manager not a SRE. I imagine people are banging down the door because of the brand but personally that’s a lot of red flags before I even submit an application.
combine that with ludicrous requirements (the same as a senior software engineer) and you get gaps in coverage. ask yourself what senior software engineer on earth would tolerate getting called CONSTANTLY at 3am, or working 3rd shift.
the vast majority of computer systems just simply aren't as important as hospitals or nuclear power plants.
Given a system that collects minute-based metrics, it generally takes around 5-10 minutes to generate an alert. Another 5-10 minutes for the person to get to their computer unless it's already in their hand (what if you get unlucky and on-call was taking a shower or using the toilet?). After that, another 5-10 minutes to see what's going on with the system.
After all that, it usually takes some more minutes to actually fix the problem.
Dropbox has a nice article on all the changes they made to streamline incidence response https://dropbox.tech/infrastructure/lessons-learned-in-incid...
Being paged constantly is a sign of bad alerts or bad systems IMO - either adjust the alert to accept the current reality or improve the system
also, even getting paged ONCE a month at 3am will fuck up an entire week at a time if you have a family. if it happens twice a month, that person is going to quit unless they're young and need the experience.
Source: co-founder of a remote startup with employees in five countries
> the vast majority of computer systems just simply aren't as important as hospitals or nuclear power plants.
I agree that the stakes are lower in terms of harm, but was trying to express that whilst it might not be life and death, it might be hindering someone being able to do their job / use your product - eg: it still impacts customer experience and your (business) reputation.
False pages for transient errors are bad - ideally you only get paged if human intervention is required, and this should form a feedback cycle to determine how to avoid it in future. If all the pages are genuine problems requiring human action then this should feed into tickets to improve things
The whole point is that before you join, the team has done sufficient work to not make it hell, and your work during business hours makes sure it stays that way. Are there a couple bad weeks throughout the year? Sure, but it's far, far from the norm.
That's a lot easier to hire, and lower cost. More training required of what is worth waking people up over; way less in terms of how to fix database/cache bugs.
https://community.openai.com/t/bug-incorrect-chatgpt-chat-se...
I only reported it on the forums because there didn't seem to be an official bug reporting channel, just a heavyweight security reporting process.
As well as the actions they took to fix this specific bug, another useful action would be to have a documented and monitored bug reporting channel.
I find even it funny now, they could write that they will provide free API creds, but now they had a very bad moment due to their greed.
Missed a word.
I dunno, I generally report issues I find in software, paid or not, as I've always done. Takes usually ~10 minutes and 1% of the time, they ask for more details and I spend maybe 20 minutes more to fill out some more details.
Never been paid for it ever, most I gotten was a free yearly subscription. But in general I do it because I want what I use to be less buggy.
Hopefully they'll start a bug bounty program soon, and prioritise bug reports over features.
This is a lot of sensitive data. It says 1.2% of ChatGPT Plus subscribers active during a 9 hour window, which considering their user base must be a lot.
This bit is particularly interesting:
> I am asking for this ticket to be re-oped, since I can still reproduce the problem in the latest 4.5.3. version
Sounds like the bug has not actually been fixed, per drago-balto.
According to the latest comments there, the bug is only partially fixed.
The OpenAI API was incredibly slow and lots of requests probably got cancelled (I certainly was doing that) for some days. I imagine someone could write a whole blog post about how that worked, it would be interesting reading.
There are 3 hard problems in Computer Science:
1. naming things
2. cache invalidation
3. 4. off-by-one errors
concurrencythat is, this seems like a very basic programming mistake and not some deep issue in Redis. the strange way it was described makes it seem like they're trying to conceal that a bit.
I understand in early 2000s we were using spinning disks and it was the only way. Well, we don't use spinning disks any more, do we?
A modern server can easily have terabytes of RAM and petabytes of NVMe, so what's stopping people from just using postgres?
A cluster of radishes is an anti-pattern.
This only makes sense if queries are computationally intensive. If you're fetching a single row by index you aren't winning much (or anything).
Or if the link to your DB is higher latency than you're comfortable with.
The only loss here is network latency which negligible when you're colocated in AWS.
Postgres's caches end up pulling a lot more weight too when you're not only hitting the db on a cache miss from the web server.
#2 is an interesting point. When you benchmark, the normal process is to just set up a database then run a shitload of queries against it. I don't think a lot of people put actual production load on the database then run the same set of queries against it...usually because you don't have a production load in the prototyping phase.
However, load does make a difference. It made more of a difference in the HDD era, but it still makes a difference today.
I mean, redis is a cache, and you do need to ensure that stuff works if your purge redis (ie: be sure the rebuild process works), etc, etc.
But just because it's old doesn't mean it's bad. OS/390 and AS/400 boxes are still out there doing their jobs.
But FWIW I can easily saturate a 10GBit ethernet link with primary key-lookup read-only queries, without the results being ridiculously wide or anything.
Because it didn't need any setup, I just used:
SELECT * FROM pg_class WHERE oid = 'pg_class'::regclass;
I don't immediately have access to a faster network, connecting via tcp to localhost, and using some moderate pipelining (common in the redis world afaik), I get up to 19GB/s on my workstation.This selects every column (*) from every table (ObjectID is of type regclass)?
It just selects every column from a single table, pg_class. Which is where postgres stores information about relations that exist in the current database.
Last time I checked, Redis doesn't take that long to provide a response. And if your Redis servers actually are that overloaded that you're seeing latency in your requests, it seems like simple key-based sharding would allow horizontally scaling your Redis cluster.
Disclaimer: I am probably less smart than most people who work at OpenAI so I'm sure I'm missing some details. Also this is apparently a Python thing and I don't know it beyond surface familiarity.
Thus, it's much cheaper to run at massive scale like OpenAI's for certain workloads, including KV caching
also:
- robust, flexible data structures and atomic APIs to manipulate them are available out-of-the box
- large and supportive community + tooling
Yet I'm wondering why there is no checking if the response does actually belong to the issued query.
The client issuing a query can pass a token and verify upon answer that this answer contains the token.
TBH as a user of the client I would kind of expect the library to have this feature built-in, and if I'm starting to use the library to solve a problem, handling this edge-case would be of a somewhat low priority to me if the library wouldn't implement it, probably because I'm lazy.
I hope that the fix they offered to Redis Labs does contain a solution to this problem and that everyone of us using this library will be able to profit from the effort put into resolving the issue.
It doesn't [0], so the burden is still on the developer using the library.
[0] https://github.com/redis/redis-py/commit/66a4d6b2a493dd3a20c...
---
Edit: Now I'm confused, this issue [1] was raised on March 17 and fixed on March 22, was this a regression? Or did OpenAI start using this library on March 19-20?
Interesing comment:
> drago-balto commented 3 hours ago
> Yep, that's the one, and the #2641 has not fixed it fully, as I already commented here: #2641 (comment)
> I am asking for this ticket to be re-oped, since I can still reproduce the problem in the latest 4.5.3. version
[1] https://github.com/redis/redis-py/issues/2624#issue-16293351...
It's not the safest design but I wouldn't say the client should be expected to implement it. That security concern is at the application layer and the actual needs of the implementation can be wildly different depending on the application. You can imagine use cases for redis where this isn't even relevant, like if it's being used to store price data for stocks that update every 30 seconds. There's no private data involved there. It's out of scope for a storage client to implement.
Before chatgpt, most normal people had never heard of OpenAI. Their flagship product was basically an API that only programmers could make useful.
Team leaders at OpenAI have stated that they were not expecting the success, let alone the highest adoption rate for any product in history. In their minds, it was just a cleaned-up version of a 2-year old product. It was billed as a research preview.
So, all of a sudden you go from hiring mostly researchers because you only have to maintain an API and some mid-traffic web infra, to suddenly having the fastest growing web product in history and having to scale up as fast as you can. Keep in mind that they didn't get backing from Microsoft until January 23, 2023-- that was only 2 months ago.
I'd say we should cut them some slack.
The cynic in me wants to believe that it's a way of deflecting blame somehow, to make it seem like they did their due diligence but were thwarted by something outside of their control. I don't think it holds. If you use an open source library with no warranty, you are responsible (legally and otherwise) to ensure that it is sufficient. For example, if you break HIPAA compliance due to an open source library, it is still you who is responsible for that.
But of course, they're not claiming it's anyone else's fault anywhere explicitly, so it's uncharitable to just assume that's what they meant. Still, it rubs me the wrong way. I can't fight the feeling that it's a wink wink nudge nudge to give them more slack than they'd otherwise get. It feels like it's inviting you to just criticize redis-py and give them a break.
The open postmortem and whatnot is appreciated and everything, but sometimes it's important to be mindful of what you emphasize in your postmortems. People read things even if you don't write them, sometimes.
Bugs happen, though. Especially in Python.
as opposed to...?
Certainly none of those are widely used and have a reputation for making it easy to keep the gun aimed squarely at the foot.
So compared to ones like Kotlin or Rust.
In my experience errors are more common (for both cultural and technological reasons) in Python than in Go.
I would guess something similar applies to Rust, though I don't have personal experience.
There's wide variation in C, but with careful discrimination, you can find very high-quality libraries or software (redis itself being an excellent example).
I don't have rigourous data to baack this stuff up, but I'm pretty convinced it's true, based on my own experience.
It's still their fault. When you ship code, you are responsible for how that code behaves regardless of where the code came from.
How many people make sure all of the open source libraries they're using are bug free?
Anyone besides maybe NASA?
Usually "fix the original library" wasn't as easy or immediate a fix as "hack around it" which is sad just re: the overall OSS ecosystem but still the person releasing a product's responsibility.
Unfortunately these sorts of bugs are wildly difficult to predict. Yet it's also a wildly common architecture. That's what's sad for all of us as engineers as a whole. But "caching credit card details and home addresses", for instance, is... particularly dicey. That's very sensitive, and you're tossing it into more DBs, without good access control restrictions?
It's a definition most laypeople use. It's developers who tend to use a very narrow definition.
I don't think it should be controversial to say that when you ship a product, you are responsible for how that product behaves.
Alternatively, they knew about it, and didn't fix the bug until it bit them
> Especially in Python.
made me unvote.
The license they agreed to in order to use this library has this in capital letters. [THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND].
After agreeing to this license and using the library for free, they charged people money and sold them a service. And when that library they got for free, which they read and agreed that had no warranty of any kind, had a software bug, they wrote a blog post and blamed the outage of their paid service on this free library.
This is not another open-source project, or a small business. This is a company that got billions of dollars in investment, and a lot of income by selling services to businesses and individuals. They don't get to use free, no-warranty code written by others to save their own money, and then blame it and complain about it loudly for bugs.
Using someone else's library doesn't absolve you of responsibility, but failing to be vigilant at thoroughly vetting and testing external dependencies is a different kind of mistake than creating a terrible security bug yourself because your engineers don't know how to code or everyone is in too much of a rush to care about anything.
A lot of the time when people make mistakes, they explain themselves so as they are afraid to be perceived as completely stupid or incompetent for making that mistake, not excusing themselves of taking responsibility even though people frequently think that excuses or explanation means that you are trying to absolve yourself of what you did.
There's a huge difference to me between having an obscure bug like this and introducing that type of security issue because you couldn't logically consider it. First one can be resolved in the future by introducing processes and make sure all open source libraries are from trusted sources, but second one implies that you are fundamentally unable to think and therefore also probably improve on that.
The result for the end consumer is identical whether they have their PII leaked from "an external library" vs a vendor's own home-baked solution.
It's not really a different kind of mistake, it's exactly the same kind of mistake, because it is exactly the same mistake! This is talking the talk, and not walking the walk, when it comes to security.
Publishing a writeup that passes the buck to some (unnamed) overworked and underpaid open source maintainer is worse, not better!
The dev had such a big ego that they didn't want to say "I was dumb and left open a bug", so the dev says "I was so dumb that I left open a bug in software I was also too dumb or lazy to write or even read". It's not better.
But the truth is that actually, maybe that hints at a deeper problem. It was a direct dependency to their application code in a critical path. I mean, don't get me wrong, I don't think everyone can be expected to audit or fund auditing for every single line of code that they wind up running in production, and frankly even doing that might not be good enough to prevent most bugs anyways. Like clearly, every startup fully auditing the Linux kernel before using it to run some HTTP server is just not sustainable. But let's take it back a step: if the point of a postmortem is to analyze what went wrong to prevent it in the future, then this analysis has failed. It almost reads as "Bug in an open source project screwed us over, sorry. It will happen again." I realize that's not the most charitable reading, but the one takeaway I had is this: They don't actually know how to prevent this from happening again.
Open source software helps all of us by providing us a wealth of powerful libraries that we can use to build solutions, be we hobbyists, employees, entrepreneurs, etc. There are many wrinkles to the way this all works, including obviously discussions regarding sustainability, but I think there is more room for improvement to be had. Wouldn't it be nice if we periodically had actual security audits on even just the most popular libraries people use in their service code? Nobody in particular has an impetus to fund such a thing, but in a sense, everyone has an impetus to fund such work, and everyone stands to gain from it, too. Today it's not the norm, but perhaps it could become the norm some day in the future?
Still, in any case... I don't really mean to imply that they're being nefarious with it, but I do feel it comes off as at best a bit tacky.
Outsourcing your development work without a acceptance criteria and without validation for fitness of purpose is complete, abject engineering incompetence. Do you think bridge builders look at the rivets in the design and then just waltz over to Home Depot and just pick out one that looks kind of like the right size? No, they have exact specifications and it is their job to source rivets that meet those specifications. They then either validate the rivets themselves or contract with a reputable organization that legally guarantees they meet the specifications and it might be prudent to validate it again anyways just to be sure.
The fact that, in software, not validating your dependencies, i.e. the things your system depends on, is viewed as not so bad is a major reason why software security is such a utter joke and why everybody keeps making such utterly egregious security errors. If one of the worst engineering practices is viewed as normal and not so bad, it is no wonder the entire thing is utterly rotten.
Yes, it’s technically up to you to vet all your dependencies, but in practice, often it doesn’t happen, people make assumptions that the code works, and that’s relevant too.
True, but that's an inexcusable practice and always has been. We as an industry need to stop accepting it.
All of us rely on millions of lines of code that we have not personally audited every single day. Have you audited every framework you use? Your kernel? Drivers? Your compiler? Your CPU microcode? Your bootrom? The firmware in every gizmo you own?
If "Reflections on Trusting Trust" has taught us anything, it's turtles all the way down. At some point, you have to either trust something, or abandon all hope and trust nothing.
Of course not. I exclude the CPU microcode, bootrom, and the like from the discussion because that's not part of the product being shipped.
But it's also true that I don't do a deep dive analyzing every library I use, etc. I'm not saying that we should have to.
What I'm saying is that when a bug pops up, that's on us as developers even when the bug is in a library, the compiler, etc. A lot of developers seem to think that just because the bug was in code they didn't personally write, that means that their hands are clean.
That's just not a viable stance to take. The bug should have been caught in testing, after all.
If your car breaks down because of a design failure in a component the auto manufacturer bought from another supplier, you'll still (rightfully) hold the auto manufacturer responsible.
That’s reacting to a bug you know about. Do you mean to talk about how developers aren’t good enough at reacting to bugs found in third party libraries, or how they should do more prevention?
In this case, it seems like OpenAI reacted fairly appropriately, though perhaps they could have caught it sooner since people reported it privately.
“Holding someone responsible” is somewhat ambiguous about what you expect. It seems reasonable that a car manufacturer should be prepared to do a recall and to pay damages without saying that they should be perfect and recalls should never happen.
My point was neither of these. My point is very simple: the developers of a product are responsible for how that product behaves.
I'm not saying developers have to be perfect, I'm just saying that there appears to be a tendency, when something goes wrong because of external code, to deflect blame and responsibility away from them and onto the external code.
I think this is an unseemly thing. If I ship a product and it malfunctions, that's on me. The customer will rightly blame me, and it's up to me to fix the problem.
Whether the bug was in code I wrote or in a library I used isn't relevant to that point.
If this bug was an open issue in the project's repo, that might be concerning and indicate that proper vetting wasn't done. Ditto if the project is old and unmaintained, doesn't have tests, etc. But if they were the first to trigger the bug and it only occurs under heavy load in production conditions, well, running into some of those occasionally is inevitable. The alternative is not using any dependencies, in which case you'd just be introducing these bugs yourself instead. Even with very thorough testing and QA, you're never going to perfectly mimic high load production conditions.
Not only do most open/free source libraries come without support agreements: they come with the broadest possible limitation of warranties. (As they should)
So the company, knowing that what they are using comes without any warranty either of quality or fitness to the use-case, have a very strong burden of due diligence / vetting.
That's how it reads to me as well.
Of course, it doesn't deflect blame at all. Any time you include code in your project, no matter where the code came from, you are responsible for the behavior of that code.
There are very few companies that couldn't get caught by this type of bug.
Thus, I honestly think firms operating generative AI should be walking on eggshells to avoid placing blame on "open-source". Rather, they really should going out of their way to channel as much positive energy towards it as possible.
Still, agree the charitable interpretation is that this just purely descriptive.
Though version reference in the postmartem should also be posted, as a general guidance to their own readers, but at least a quick google search leads you to it.
https://github.com/redis/redis-py/releases/tag/v4.5.3
For anyone reading this and using a combination of asyncio and py-redis, please bump your versions.
Similar issues I've encountered with asyncio python and postgres too in the past when trying to pool connections. It's really not easy to debug them either.
The technical root cause was in the open source library. There's a patch available and more likely than not OpenAI will continue to use the library.
Being overly sensitive to blame would be distracting to the technical issue at hand. It's great they are posting this post-mortem to raise awareness that the libraries you use can have bugs and to consider that risk when building systems.
Would likely also question the lack of resources allocated to these open source projects by companies with a lot of profits from, in part, using those open source projects.
"In order to prevent a bug like this from happening in the future, we have stepped up our review process for external dependencies. In addition, we are conducting audits around code that involves sensitive information."
Of course, we all know what actually happened here:
- we did no auditing;
- because our audit process consists of "blame someone else when our consumers are harmed";
- because we would rather not waste dev time on making sure our consumers are not harmed
If you want to know why no software "engineering" is happening here, this is your answer. Can you imagine if a bridge collapsed, and the builder of the bridge said, "iunno, it's the truck's fault for driving over the bridge."
It's a serious bug, but in the grand scheme of things, not earth shattering, and not something that I think would discourage usage of their product. But their treatment of the bug causes more concerns than the bug itself. They are shifting the blame away from the work they did using a library with a bug, rather than their process by which that library made it into their product. And I don't understand how they can't see how that reflects poorly on them as an AI company.
I find it so confusing that at the end of the day, OpenAI's biggest product is having created a good process by which to create value out of a massive amount of data, and build a good API on top of it. And the open source library is effectively something they processed into their product and built an API based off of it. So it creates (to me) some amount of doubt about how they will react when faced with similar challenges to their core product. How will they behave when the data they consume impacts their product negatively? From limited experience, they'll shift the blame to the data, not their process, and keep it pushing.
It seems likely that this is only the beginning of OpenAI having a large customer base, with a high impact on many products. This is a disappointing result on their first test on how they'll manage issues and bugs with their products.
Edit: On second thought, maybe they could have just written "external library" instead of "open source library".
Libraries written prior to asyncio/green threads often have that functionality enabled by means of monkey patches or shims that juggle file handles or other shared state, and there are data races. When SSL is used hopefully those data races reading from sockets result in MAC errors and the connection is terminated.
It's easy to imagine the same mistake happening one level up, in message passing code that manages queues or other shared data structures.
Search for "python django OR celery OR redis OR postgres OR psycopg2 decryption failed or bad mac" and you'll see the scale of the issue, it's fairly widespread.
I don't have a high degree of confidence in this ecosystem. I've written about this before on Hacker News[1], and I'm not confident in the handling of shared data structures in Python libraries. I don't think I can blame any maintainers here, there's a huge number of people asking for these concurrency features and the way that it's often implemented - monkey patching especially - makes it extremely difficult to do correctly.
I reported this bug in the forums (no response), because I couldn't find an official way to report bugs - at least it wasn't documented anywhere that I could find. The actions they've taken fix that bug, but not the next one. OpenAI should take an action to improve their bug reporting channels, so the next bug gets found more quickly.
https://community.openai.com/t/bug-incorrect-chatgpt-chat-se...
Other messages indicate issues with bug reporting too:
* https://news.ycombinator.com/item?id=35291943 * https://news.ycombinator.com/item?id=35295747
If a request is canceled after the request is pushed onto the incoming queue, but before the response popped from the outgoing queue, we see our bug: the connection thus becomes corrupted and the next response that’s dequeued for an unrelated request can receive data left behind in the connection.
https://github.com/redis/redis-py/commit/66a4d6b2a493dd3a20c...
This section summarizes the issue with the ChatGPT service which caused some users to be able to see another user's chat history and payment information. OpenAI took the ChatGPT service offline and published technical details of the problem. The bug was found in the Redis client open-source library and has been patched. OpenAI has reached out to affected users and improved their systems to prevent this from happening again. They have also expressed their commitment to protecting user privacy and their dedication to the open-source community.
> ChatGPT experienced an outage due to a bug in an open-source library that allowed some users to see titles from another active user's chat history and payment-related information of 1.2% of ChatGPT Plus subscribers was exposed; the bug has since been patched and actions have been taken to prevent future incidents.
I then asked it why there was the outage and it spit out:
> The exact cause of the Facebook outage is unknown, but it is believed to be related to a configuration change in the company's routers.
It's likely because I ran out of tokens because the OpenAI outage report is long. Pasting in the text of the outage report, and then re-asking about why, it was able to give a much better answer:
> There was an outage due to a bug in an open-source library that allowed some users to see titles from another active user's chat history and also unintentionally exposed payment-related information of 1.2% of ChatGPT Plus subscribers who were active during a specific nine-hour window.
Querying it further, again having to repeat the whole OpenAI outage report, and asking it a few different ways I eventually managed to get this succinct answer:
> The bug was caused by the redis-py library's shared pool of connections becoming corrupted and returning cached data belonging to another user when a request was cancelled before the corresponding response was received, due to a spike in Redis request cancellations caused by a server change on March 20.
It did take me more than a few minutes to get to there, so just actually reading the report would have been faster, and I ended up having to read the report to verify that answer was correct and not a hallucination anyway, so our jobs are safe for now.
---
Welcome back! Here are some takeaways from this page.
> ChatGPT was offline due to a bug in redis-py that caused some users to see other users’ chat history and payment information.
> The bug was patched and the service was restored, except for a few hours of chat history.
> The bug affected 1.2% of ChatGPT Plus subscribers who were active during a nine-hour window on March 20.
> Full credit card numbers were not exposed at any time.
> OpenAI apologized to the users and the ChatGPT community and took steps to prevent such incidents in the future.
---
In the one reddit post first surfacing this the user saw conversations related to politics in china and other rather sensitive topics related to CCP.
This can absolutely get people hurt and they absolutely must take this serious.
No, everyone assumes their session object is instantiated with the right values at that level of the code.
in chat history there is a button to "retry" but clicking it and inspecting the result, you see "internal server error"
For example: spopsadd
local src_list = ARGV[1] local dst_list = ARGV[2] local value = redis.call('spop', src_list) if value then -- avoid pushing nils redis.call('sadd', dst_list, value) end return value
It's quite unsettling to see this leak of highly sensitive information and a private key exposure as well. Doesn't look good and seems like they don't take security seriously.
Yet another case study of absolutism, which can be simply dismissed.
People paying for ChatGPT care once it goes down and getting their details and chats leaked that is certainly outside of HN. Same with GitHub. Both having ~100M users between them.
That's the reality.
But, I know that shit happens and the reliability meter should be flexible for different things (bridges, heart surgery and chat agent).
If I train my brain to bitch, whine, moan about every thing, I'd not have resources to care about really important things.
"Yeah". Two people out of the majority of commenters not caring when it was unavailable and having their chat history leaked and as well as GitHub going down again. [0] Surely that alone is "Nobody" caring. /s
This is the whole problem with you being absolutist. Seems to have aged like milk.
Every fart some AI related person makes becomes a huge news. And it's followed by tens of random blog postings all posted to HN.
But we will see AI stuff rewritten in Rust quite soon.
Part of it is that, the average engineer could understand and grok what those articles were talking about, and I could appreciate, relate, and if applicable criticize it.
The AI news just seems to swing between hype and doomsday prophecies, and little discussion about the technical aspects of it.
Obviously OpenAI choosing to keep it closed source makes any in-depth discussion close to impossible, but also some of this is so beyond the capabilities of an average engineer with a laptop. It can be frustrating.
I was logging in during heavy load, and after typing the question I started getting responses to questions which I didn't ask.
gdb answered on that comment "these are not actually messages from other users, but instead the model generating something ~random due to hitting a bug on our backend where, rather than submitting your question, we submitted an empty query to the model."
I wonder if it was the same redis-py issue back then, but just at another point in the backend. His answer didn't really convince me back then.
[0] https://news.ycombinator.com/item?id=34614796&p=2#34615875
This was a failure of integration testing and defensive design, whether the component was open-source or not. There's no reason to believe that an AI company would have the diligence and experience to do the grunt work of hardening a site.
But management obviously understood the level and character of interest. Actual users include probably 10,000 curiosity seekers for every actual AI researcher, with 1,000 of those being commercial prospects -- people who might buy their service.
This is a clear sign that the managers who've made technical breakthrough's in AI are not capable even of deploying the service at scale -- no less managing the societal consequences of AI.
The difficulty with the board getting adults in the room is that leaders today give the appearance of humility and cooperation, with transparent disclosures and incorporation of influencers into advisory committees. The leaders may believe their own abilities because their underlings don't challenge them. So there's no obvious domineering friction, but the risk is still there, because of inability to manage.
Delegation is the key to scaling, code and organizations. "Know thyself" is about knowing your limits, and having the humility to get help instead of basking in the puffery of being in control.
This isn't a PR problem. It's the Achilles' heel of capitalism, and the capitalists in OpenAI's board should nip this incipient Musk in the bud or risk losing 2-3 orders of magnitude return on their investment.