A ChatGPT mistake cost us $10k
asim.bearblog.dev
asim.bearblog.dev
If you haven't fixed that alerting deficiency, then you haven't really fixed anything.
Programming when everything works is easy, it's handling the problems that makes it hard.
"Under construction "
Looks like the OP removed the post?
EDIT: Found archived copy of post: http://web.archive.org/web/20240609213809/https://asim.bearb...
One of the reasons I use Go whenever possible is that it removes a lot of the classic Python footguns. If you are going to rewrite your backend from Javascript, why would you rewrite it in another untyped, error-prone language?
In go, I've definitely seen:
tx, err := db.Tx()
defer tx.Commit() // silently ignores the error on committing, which is the important one
That would have masked this error so it didn't get logged by the application.In python, if you ignore an exception entirely, like I did that error above, you instead get an exception logged by default.
Python's exceptions also include line numbers, where as Go errors by default wouldn't show you _which_ object has a conflict, even if you logged it.
In general, python's logs are way better than Go's, and exceptions make it way harder to ignore errors entirely than Go's strategy.
I suppose there are probably similar checkers for Python that would have caught the passing of a scalar value instead of a function.
Perhaps this is an argument for mandatory linting in CI.
What did they say was the thinking behind it? defer tx.Rollback() would make sense, but defer tx.Commit() is nonsensical, regardless of whether or not the error is handled. It seems apparent that the problem there isn't forgetting to check an error, but that someone got their logic all mixed up, confusing rollback with commit.
It's not pure nonsense, it works in the happy path, and it matches the pattern of how people often handle file IO in go.
f, err := os.OpenFile(...)
defer f.Close()
... which is another place most gophers ignore errors incorrectly. Just like the "defer tx.Commit()" example, it's collocating the idea of setup and cleanup together.Those two patterns are so similar, python handles them in the same way:
with db.begin() as conn: # implicit transaction, gets automatically committed
with open(...) as f: # implicit file open + close pair, automatically closed
You're of course right that go requires more boiler-plate to do the right thing, but the wrong code has no compiler errors and works in the happy path, and fits how people think about the problem in other languages with sane RAII constructs, so of course people will write it.If you are only reading, then sure, who cares?
Okay, sure, but Rollback is the cleanup function. You always want to rollback – you only sometimes want to commit. I suppose this confirms that someone got their logic mixed up.
This phrase highlights the confusion. If you learned SQL before Go, you want to rollback only on error.
Every time I write `defer tx.Rollback()`, I cringe and have to remind myself that yes, it's actually ok to call a method called `Rollback` after succesfully writing data.
BEGIN;
INSERT INTO ...
COMMIT;
ROLLBACK;
It is not some kind of Go-ism. The Go database/sql package actually executes ROLLBACK in the SQL engine. Check out the error returned by it.Perhaps you mean learned SQL in the context of languages that consider a failed rollback an exception? In that case one needs to be careful to not rollback, else be stricken to handling the exception, which programmers seem to hate doing.
No: https://github.com/golang/go/blob/beaf7f3282c2548267d3c89441...
BTW I checked and it's not an error in Postgres, only a warning. Still not something I would want in my database logs for the happy path.
But it's a good sanity check/safety measure to call it anyway incase you made a mistake elsewhere. An errant rollback is more likely to be caught in testing than a dangling transaction.
Now granted, if they'd carried on doing stuff (but not redeployed production) it may have shown up in say mid-aftenoon, or not depending on volume.
And of course the general guideline of not deploying anything to production early in the day, or on Friday, is still valid.
And it's not trivial to set up a correct alert for this one. Simple HTTP 5xx threshold wouldn't work because it's high enough to not wake you up whenever cloud provider restarts Postgres, it's too high to catch this. You need either per-endpoint failure rate alerts or something else more clever.
As soon as you expect paying customers in your system you need to have someone with the knowledge and experience to deal with infrastructure. That means logging, monitoring, alerting, security etc.
DevOps.. amateurs.
But not having error logging/alerts on your db ? That's the crazy part.
This is a new product, is not legacy code from 20 years ago when they thought it was a neat idea to just throw stuff at the db raw, and check for db errors to do data validation, so alerts are hard because there's so many expected errrors.
Unit tests are good, yes. Monitoring is also good. But just taking 30 seconds to do some manual testing will catch a LOT of unexpected behavior.
1. Get it working: write the code for the desired behavior, not worrying about making it beautiful, testable, whatever.
2. Get it working well: manually testing and finding edge cases, refactoring to get it testable and writing tests to solidify behavior.
3. Get it working fast: optimizing it to be as fast as I need it to be (can sometimes skip this step). No tests should change here, but only new tests.
https://web.archive.org/web/20240610032818/https://asim.bear...
The author has added an important edit:
> I want to preface this by saying yes the practices here are very bad and embarrassing (and we've since added robust unit/integration tests and alerting/logging), could/should have been avoided, were human errors beyond anything, and very obvious in hindsight.
>
> This was from a different time under large time constraints at the very earliest stages (first few weeks) of a company. I'm mostly just sharing this as a funny story with unique circumstances surrounding bug reproducibility in prod (due again to our own stupidity) Please read with that in mind
They did make a silly mistake, but we are humans, and humans, be it individually or collectively, do make silly mistakes.
If you're earning past six figures, are part of a team of programmers, call yourself an professional / engineer, and have technical management above you like a VP of Engineering, yadda yadda....then it's closer to systematic failure of the company's engineering practices than "mistake."
There is a reason we call it software engineering, not software fuckarounding (or, cough, "DevOps Engineeer".)
Software engineering practices assume people are going to make mistakes, and implements procedures to reduce the chances of that making it into production, and reduce the impact of those mistakes if they do make it into production.
https://bsago.me/tech-notes/change-ssh-background-colour-wit....
I assume that all systems already have descriptive names App_DEV_Server1, App_PROD_Server5, etc.
It also helps if (ofc they would be right??) in separate IP groups/WLANS?
If you are running Windows, it's a good idea to use BGINFO.exe by SysInternals (or Winternals as we old people still call it), and display the most relevant info (showing Dev/Prod/UAT/etc.) with big-big-big letters.
Really all you need is logging and potentially temporary read access to the db if you need some info that you can't derive from the logs.
https://about.gitlab.com/blog/2017/02/10/postmortem-of-datab...
Compensation is in no way correlated with good engineering practices.
They might be paid much because they're developing something which people are willing to pay for, it doesn't have to be "real engineering".
Being on the receiving end of an internet pile-on of "OMG you idiots everyone knows the first thing you do when setting up a flerble cluster is spend a week installing grazoono monitoring!" is not conducive towards building a good engineering culture.
https://0912i390129ionkjan.bearblog.dev/how-a-single-chatgpt...
but google cache still serves a copy...
https://webcache.googleusercontent.com/search?q=cache%3Ahttp...
I did step into that particular trap more than once (passing the result, rather than the function)
"We made a programming error in our use of an LLM, didn't do any QA, and it cost us $10k" doesn't generate the C-suite "oh shit what if ChatGPT fucks up, what's our exposure!?" reaction. There's a million middle and upper management posting this article on LinkedIn, guaranteed.
It's like the Mr. Beast open-mouth-surprised expression thumbnail nonsense; you feel incredibly compelled to click it.
While we're on the subject: LLMs can't make "mistakes." They are not deterministic.
They cannot reason, think, or do logic.
They are very fancy word salad generators that use a lot of statistical probabilities. By definition they're not capable of "mistakes" because nothing they generate is remotely guaranteed to be correct or accurate.
Edit: The mods boosted the post; it got downvoted into oblivion, for obvious reasons, and then skyrocketed instantly in rank, which means they boosted it: https://hnrankings.info/40627558/
Hilarious that a post which is insanely clickbait (which the rules say should result in a title rewrite) got boosted by the mods.
I'm sure it's a complete coincidence that the story was apparently authored by someone at a Ycombinator company: https://news.ycombinator.com/item?id=40629998
This makes no sense. Only things that are guaranteed to be correct or accurate can make mistakes? Everyone knows what "mistake" means in this context. Nobody cares what your preferred definition of mistake is.
Hard to put in words.
But that's roughly what is concerning about the "AI makes mistakes" narratives.
It implies they are caused by a (fixable) fault in reasoning or memory.
LLM "AI" will always respond that it made a "mistake" when you correct it.
It is trained to do so, and humans often behave similarly.
It is hard to come up with a good definition for "mistake", yes.
That does not change that using this in case of LLM hallucinations is misleading.
By that logic, nothing is capable of making mistakes :D.
> Hilarious that a post which is insanely clickbait (which the rules say should result in a title rewrite) got boosted by the mods.
You have a distorted view of what clickbait is and the rules of this site. I suggest you go calm down and try to stop hating on a technology which is just that: a technology! Like any other, it can be misused, but think about why exactly you feel so passionate about this particular technology.
It is, after all, incapable of feeling hurt.
> It's like the Mr. Beast open-mouth-surprised expression thumbnail nonsense; you feel incredibly compelled to click it.
I feel incredibly compelled to ignore it.
Sponsorblock is great for combating that. (Altough, I conciously avoid channels that mostly do clickbait anyways)
...also from next.js and prisma to python? ...what?
* UUID Generation in Primary Key: The default parameter should use the callable uuid.uuid4 directly instead of str(uuid.uuid4()). SQLAlchemy will call the function to generate the value.
* Date Default Value: server_default=text("(now())") might not work as expected. Use func.now() for server-side defaults in SQLAlchemy.
* Import Statements: Ensure uuid and text from sqlalchemy are imported.
* Column Definitions: Consider using DateTime(timezone=True) for datetime columns to handle time zones.
It then provided me with corrected code that does
id = Column(String, primary_key=True, default=lambda: str(uuid.uuid4()), unique=True, nullable=False)
where the addition of lambda: fixes the problem.The other common issue is if the original code has thinsg chatgpt doesn't like (misspell, slightly wrong formatting) it will fix it automatically, or if he really think you should have added a particular field you didn't add.
But what I don't understand is how this wasn't caught after the first failure? Does this company not have any logging? Shouldn't the fact the backend is attempting to reuse UUIDs be immediately obvious from observing the error?
UUIDs are just 128-bit values. They might be conventionally encoded for humans as hex, but storing them as 36-byte (plus a few more for length) strings is a pointless waste of both space and performance.
I don't think 128 bits vs 36 byte performance it's a main concern right now
36B vs 16B today, tomorrow you need an array of it, and now it isn't cache aligned, and more than twice the overhead.
Most likely instead of manipulating a 36B fixed length string, it is handled as a dynamic string, for extra runtime memory allocations, most likely consuming at least 64B per allocation. Etc etc.
Do this all over the codebase and now you know why all the moderne software is a sloth on what was a supercomputer 30y ago.
It’s also not just the size itself. Despite being fixed-size in practice, these are variable-sized strings in application code which now means gajillions of pointless allocations and indirection for everything. There are a ton of knock-on performance consequences here, all on the most heavily-used columns in your data model.
Worst of all should they actually succeed, this is going to be absolutely excruciating to fix.
But in either case – MySQL or Postgres – they’ve still made the classic mistake of using a UUID as a PK, which will tank performance one way or another. They’ll notice right around when it starts to matter.
I am also very confused about the apparent lack of logging or recourse to logging. It's been a while, but if I recall correctly ECS should automatically propagate the resulting Duplicate Key exceptions which were presumably occurring to CloudWatch without a bunch of additional configuration - was that not happening? If it was happening, did no one think to go check what types of Exceptions were happening overnight?
I guarantee you that they _will_ have another production bug like this sometime in the future (every fast paced project will). You'd hope this next one wont take 5 days to identify.
Specifically asking why did it take so long to detect and why did it take so long to diagnose is useful in these situations.
Type 1 tries to find the error message and figure out what it really, really means by breaking down the error message and system.
Type 2 does trial and error on random related things until the problem goes away.
I hate to say that I've seen way more type 2s engineers than type 1, but maybe I’m working at the wrong companies.
Here we are talking 1.65 MILLION CAD $ backed YC company
I felt the blog post failed to articulate the root cause of the issue and went straight to blaming ChatGPT.
When you rush and make large or non peer code reviewed commits to main it is going to happen.
The real issue was when you rush, take shortcuts and don’t adequately test and peer code review then errors will occur.
I would have imagined that a test that tried a few different signup options would have found the issue immediately.
To be truly reflective, OP needs to dive into the real reason their code had this issue. It wasn’t using GPT, it was not having the controls in place.
However, this engineer can type infinitely fast, which means it might be useful if used very carefully.
Anyway, letting such a person near financially important code would lead to similar issues, and in both cases, I’d question the judgment of the person that decided to deploy the code at all, let alone without much testing.
/s
(I hope)
I'm surprised there was no lint rule for this case.
This is the eye opener for me, how is a startup justifying a re-write when they don't even have customers?
Perhaps also the tooling because any remotely decent IDE should show an error there, let alone the potential warnings of some code analysis software.
(This is one thing that baffles me about the “let’s use LLMs to code” movement; a lot of the proponents don’t seem to be just adding it as a tool (I don’t think it’s a terribly _useful_ tool, but whatever, tastes differ), but using it as the only tool, discarding 50 years worth of progress.)
In my case (with a real project I'm working on now), it'd be due to realizing that C# is a great language and has a good runtime and web frameworks, but at the same time drags down development velocity and has some pain points which just keep mounting, such as needing to create bunches of different DTO objects yet AutoMapper refusing to work with my particular versions of everything and project configuration, as well as both Entity Framework and the JSON serializer/deserializer giving me more trouble than it's worth.
Could the pain points be addressed through gradual work, which oftentimes involves various hacks and deep dives in the docs, as well as upgrading a bunch of packages and rewriting configuration along the way? Sure. But I'm human and the human desire is to grab a metaphorical can of gasoline, burn everything down and make the second system better (of course, it might not actually be better, just have different pain points, while not even doing everything the first system did, nor do it correctly).
Then again, even in my professional career, I get the same feeling whenever I look at any "legacy" or just cumbersome system and it does take an active, persistent effort on my part to not give in to the part of my brain that is screaming for a rewrite. Sometimes rewrites actually go great (or architectural changes, such as introducing containers), more often than not everything goes down in a ball of flames and/or endless amounts of work.
I'm glad that I don't give in, outside of the cases where I know with a high degree of confidence that it would improve things for people, either how the system runs, or the developer experience for others.
Hell, most simple applications could do with just a single layer - schema registration in EF Core is mapping, or at most two, one for DB and one for response contracts.
Just do it the simplest way you can. I understand that culture in some companies might be a problem, and it's been historically an issue plaguing .NET, spilling over, originally, from Java enterprise world. But I promise you there are teams which do not do this kind of nonsense.
Things really have improved since .NET Framework days, EF Core productivity wise, while similar in its strong areas, is pretty much an entirely new solution everywhere else.
The thing is, that you'll probably have entities mapped against the database schema with data that must only conditionally be shown to the users. For example, when an admin user requests OrderDetails then you'll most likely want to show all of the fields, but when an external user makes that request, you'll only want to show some of the fields (and not leak that those fields even exist).
DTOs have always felt like the right way to do that, however this also means that for every distinct type of user you might have more than one object per DB table. Furthermore, if you generate the EF entity mappings from the schema (say, if you handle migrations with a separate tool that has SQL scripts in it), then you won't make separate entities for the same table either. Ergo, it must be handled downstream somewhere.
Plus, sometimes you can't return the EF entities for serialization into JSON anyways, since you might need to introduce some additional parsing logic, to get them into a shape that the front end wants (e.g. if you have a status display field or something, the current value of which is calculated based on 5-10 database fields or other stuff). Unless it's a DB view that you select things from as-is, though if you don't select data based on that criteria, you can get away by doing it in the back end.
Not to say that some of those can't be worked around, but I can't easily handwave those use cases away either. In Java, MapStruct works and does so pretty well: https://mapstruct.org/ I'd rather do something like that, than ask ChatGPT to transpose stuff from DDL or whatever, or waste time manually doing that.
I'll probably look into Mapperly next, thanks! The actual .NET runtime is good and tools like Rider make it quite pleasant.
I'm sure MapStruct would also require you to handle differences in data presentation, in a similar way you would have to do with Automapper (or Mapperly). .NET generally puts more emphasis on "boilerplate-free happy path + predictable behavior" so you don't have autowire, but also don't have to troubleshoot autowire issues, and M.E.DI is straightforward to use, as an example. In terms of JSON (with System.Text.Json), you can annotate schema (if it's code-first) with attributes and nullability, so that the request for OrderDetails returns only what is available per given access rights scope. In either case different scopes of access to the same data and presentation of such is a complex topic.
Single-layer case might be a bit extreme - I did use it in a microservice-based architecture as PTSD coping strategy after being hard burned by a poor team environment that insisted on misusing DDD and heavy layering for logic that fits into a single Program.cs, doing a huge disservice to the platform.
Another popular mapping library is Mapster: https://github.com/MapsterMapper/Mapster, it is more focused on convenience compared to Mapperly at some performance tradeoff, but is still decently fast (unlike AutoMapper which is just terrible).
For fast DTO declaration, you can also use positional records e.g. record User(string Name, DateOnly DoB); but you may already be aware of those, noting this for completeness mostly.
Overall, it's a tradeoff between following suboptimal practices of the past and taking much harder stance on enforcing simplicity that may clash with the culture in a specific team.
> Our project was originally full stack NextJS but we wanted to first migrate everything to Python/FastAPI.
> What happened was that as part of our backend migration, we were translating database models from Prisma/Typescript into Python/SQLAlchemy. This was really tedious. We found that ChatGPT did a pretty exceptional job doing this translation and so we used it for almost the entire migration.
ChatGPT wasn't a net positive if they wouldn't have tried to do this migration up-front without it.
Possibly they had better error logging in the other stack, possibly they didn't, possibly they needed it less because they were actually writing the code for it themselves and knew how it worked.
("Write all the code a second time before turning on monetization" is itself an interesting decision, of course.)
https://grook.ai/share?id=e269e88a7b1a71eff4f176c864b30161&x...
1. Forward thinking
2. Maturity
3. Temperance
All important traits in engineers and founders
instead of `bug: fix blah`, it's `:bug:: fix blah`, which, honestly actually seems clearer and easier to parse at a glance
edit: hacker news doesn't support unicode emojis
These "constraints" are why I'm terrified of subscribing to software
The previous world where you buying per user seat licenses for hundreds and hundreds of dollars wasn't great.
We had race conditions where we would charge users twice
This has made me paranoid that any time I see timeout or error related to money I assume it went through and come back later.
It took me several threatening emails to make them understand they had already taken my money and I wasn't going to try again until I got a refund. Now I'm paranoid any time I purchase online at mediocre shops.
What the actual fuck hahahaha
This is made worse when they edit saying the reason the codes crap is because of time constraints, but spent their time refactoring across languages and spinning up a distributed system FOR NO REASON. That is self imposed harm juggling features and ridiculous technical complexity. What were they thinking.
Edit: a YC summer ‘23 company who’s product is still behind a waitlist summer ‘24, presumably because of a rewrite to Rust
They literally had 1 instance of the backend per $1 of revenue, and the reason the bug wasn't seen straight away was because they had 40 backend instances each with a single uuid that could be used for users before it broke with non-unique id errors.
Fully acknowledging the irony I am about to invoke - this is why I hate startup culture. Not startups, but this ridiculous culture of "well the VC gave us a million bucks and that bought us $100,000 in AWS credits, so let's just use it."
As someone who has built my company fully on my own dime (and the dimes of two colleagues), it's easy enough to burn piles of money in AWS (or any other cloud) when you're making an attempt at being judicious. Spinning up eight backends (edit: running five instances each, no less!) just because you have money, despite the fact that you know you don't need that much compute, is just insane. If for no other reason than you're just throwing your own credits away.
But then the whole startup culture is generally speaking, a culture of waste, pride and vanity.
Python has this misfeature whereby it didn't correctly crib Common Lisp's evaluation strategy for the expressions that give default values to optional function arguments.
When you have an argument like foo=obj.whatever() the obj.whatever() is evaluated (would you believe it!) at the time the definition of the function is being processed, not at the time when the function is being called and the value is needed.
Is suspect this was done on purpose, for efficiency. Python has another misfeature: it has no literal syntax for certain common objects like lists. There is [1, 2, 3], but that is a constructor and not a literal: it has to create a new list every time it is evaluated and stuff it with 1, 2, 3. (Unless a clever compiler can prove that this can be optimized away without harm.)
The designer didn't want a parameter like list=[] to have to construct a new, empty list object each time the argument is omitted. In Lisp '(1 2 3) and '() are true literals. Whenever they are referenced, they denote the same object. The programmer has a choice here: they can use (list 1 2 3) as the default value expression or '(1 2 3). The former is like [1, 2, 3]: it yields a new object each time that is mutable; the other will (almost certainly) yield the same object and cannot be reliably, portably modified.
Hey, modern popular languages have most of the features of Lisp, so you're not missing anything.
This can't be correct, surely? What if .whatever() relies on internal state that changes after obj is initialized (or after the function surrounding foo is declared, not sure what you're saying)?
basically, having a default argument value in a function definition means to evaluate it during definition time of that function, not when the function is invoked.
This is a foot gun.
The easiest way to see this is by running something like this and seeing what gets printed out and when:
print("1. start")
def function(arg=print("2. func definition")):
print("4. func call")
print("3. after definition")
function()
function()
function()
You should see that the print statement in the default position is called once, when the function definition itself is being evaluated, not when the function gets called. def func(arg=[]):
arg.append(1)
print(arg)
func()
func()
func()
shoving a print into a function definition is weird and not something you'd do normally. But someone who doesn't know this footgun is going to write a function that defaults to an empty list, and then tear their hair out when things are broken.Except every python 101 text seems to go over it, and people seem to have suddenly forgotten about it
ChatGPT driven development maybe?
Thing is, there's no obviously correct behavior here, and there are valid arguments to be made either way. Which is why many languages dodge the bullet by only allowing for compile-time constants as defaults in that context (and if you want to evaluate something at runtime, you can always use an optional and do an explicit check inside the body, or provide an overload).
My eyes were opened one relaxing morning when sipping my coffee and pondering how to tidy up a database column by migrating from string to enum.
I asked ChatGPT for its thoughts and its response seemed perfunctory and on point, until at one particular line, tucked in the otherwise sensible migration file [1], it casually recommended deleting all users whose value for that attribute wasn't among those specified by the enum. I spat my coffee out and learned a very valuable lesson that morning!
Idioms are, very conveniently, questions about what the most common shape of X among the broader community, so it tends to do quite well with that.
I appreciate when it spits out example code, but I never copy/paste from it, I always rewrite anything myself to ensure I don’t slip up and accidentally… well, what it tried to sneak in to you…
That’s quite the scary anecdote
On the other hand, I don’t think advertising the fact that the company introduced a major bug from copy and pasting ChatGPT code around and that they spent a week being unable to even debug why it was failing.
I don’t know much about this startup, but this blog post had the opposite effect of all of the other high quality post-mortem posts I’ve read lately. Normally the goal of these posts is to demonstrate your team’s rigor and engineering skills, not reveal that ChatGPT is writing your code and your engineers can’t/won’t debug it for a very long time despite knowing that it’s costing them signups.
Except, have you met startup devs? This is by and large the "move fast then unbreak things" approach.
Forget knowing anything, just come up with a nice pitch deck and let the LLM write the stack.
Not wholly surprised these people are YC backed. I’ve got the impression YC don’t place much weight on technical competence, assuming you can just hire the talent if you know how to sell.
Well, now replace “hire some talent” with “get a GPT subscription and YOLO”, and you get the foundation these companies of tomorrow are going to be built on.
Which hey, maybe that’s right and they know something I don’t.
My early rant against this mentality: https://news.ycombinator.com/item?id=19214749
People working in "big tech" aren't fundamentally better at building reliable tools and systems; the time and resource constraints are entirely different.
5 days to find out you have "duplicate key" errors in the db is the opposite of fast
If anything, I think this says something about how dangerous ChatGPT and similar tools are: reading code is harder than writing code, and when you use ChatGPT, your role stops being that of a programmer and becomes that of a code reviewer. Worse, LLMs are excellent at producing plausible output (I mean that's literally all they do), which means the bugs will look like plausibly correct code as well.
I don't think this is indicative of people who don't know what they're doing. I think this is indicative of people using "AI" tools to help with programming at all.
I think using AI tools to write production code is probably indicative of people who don't really know what they are doing.
The best way not to have subtle bugs is to think deeply about your code, not subcontract it out -- whether that is to people far away who both cannot afford to think as deeply about your code and aren't as invested in it, or to an AI that is often right and doesn't know the difference between correct and incorrect.
It's just a profound abrogation of good development principles to behave this way. And where is the benefit in doing this repeatedly? You're just going to end up with a codebase nobody really owns on a cognitive level.
At least when you look at a StackOverflow answer you see the discussion around it from other real people offering critiques!
ETA in advance: and yes, I understand all the comparison points about using third party libraries, and all the left-pad stuff (don't get me started on NPM). But the point stands: the best way not to have bugs is to own your code. To my mind, anyone who is using ChatGPT in this way -- to write whole pieces of business logic, not just to get inspiration -- is failing at their one job. If it's to be yours, it has to come from the brain of someone who is yours too. This is an embarrassing and damaging admission and there is no way around it.
ETA (2): code review, as a practice, only works when you and the people who wrote the code have a shared understanding of the context and the goal of the code and are roughly equally invested in getting code through review. Because all the niche cases are illuminated by those discussions and avoided in advance. The less time you've spent on this preamble, the less effective the code review will be. It's a matter of trust and culture as much as it's a matter of comparing requirements with finished code.
You could say the same about the output of a compiler. No one owns that at a cognitive level. They own it at a higher level - the source code.
Same thing here. You own the output of the AI at a cognitive level, because you own the prompts that created it.
Except, for starters, that you're not using the LLM to replace a compiler.
You're using it to replace a teammate.
Notwithstanding the fact that compilers did not fall out of the sky and very much have people that own them at the cognitive level, I think this is still a different situation.
With a compiler you can expect a more or less one to one translation between source code and the operation of the resulting binary with some optimizations. When some compiler optimization causes undesired behavior, this too is a very difficult problem to solve.
Intentionally 10xing this type of problem by introducing a fuzzy translation between human language and source code then 1000xing it by repeating it all over the codebase just seems like a bad decision.
But at least it's supposed to be deterministic. And there's a chance someone else will be able to explain the inner workings in a way I can repeatably test.
(1) Compilers are reproducible (or at least repeatable), so you can share your problem with other, and they can help.
(2) For common languages, there are multiple compilers and multiple optimization options, which (and that's _very important_) produce identically-behaving programs - so you can try compiling same program with different settings, and if they differ, you know compiler is bad.
(3) The compilers are very reliable, and bugs when compiler succeeds, but generates invalid code are even rarer - in many years of my career, I've only seen a handful of them.
Compare to LLMs, which are non-reproducible, each one is giving a different answer (and that's by design) and finally have huge appear-to-succeed-but-produce-bad-output error rate, with value way more than 1%. If you had a compiler that bad, you'd throw it away in disgust and write in assembly language.
> I think using AI tools to write production code is probably indicative of people who don't really know what they are doing.
People said the same to me for using Microsoft IntelliSense 20 years ago. AI tools for programming are absolutely the future.Colour me cynical but I don't feel like pretending the future is here only to have to have to fix its blind incompetence.
It is possible to use libraries correctly.
It is not possible to use AI correctly. It is only possible to correct its inevitable mistakes.
But our existing tools are already built to help us avoid this.
Back in the day, I used a tool from a group in Google called “error-prone”. It was great at catching things like this (and Lorne goal NPE in Java). It would examine code before compiling to find common errors like this. I wish we had more “quick” check tools for more languages.
It's not in quotes. It's a function call.
The issue is that the function call happens once, when you define the class, rather than happening each time you instantiate the class.
I say that not to brag because (a) default args is a known python footgun area already and (b) I'd hope most developers with any real Django or SQLAlchemy experience would have caught this pretty quick. I guess I'm just suggesting that maybe domain experience is actually worth something?
Also, where were their tests in the first place? Or am I expecting too much there?
Not only does it goof up less frequently on small focused snippets like this, it also requires me to pick the example apart and pay close enough attention to it that goofups don’t slip by as easily and it gets committed to memory more readily than with copypasting or LLM-backed autocomplete.
It's also the case that "code review" covers a lot of things, from quickly skimming the code and saying eh, it's probably fine, to deeply reading and ensuring that you fully understand the behavior in all possible cases of every line of code. The latter is much more effective, but probably not nearly as common as it ought to be.
In my experience, hardly anyone in software does know what they're doing, for sufficiently rigorous values of "know what you're doing." We all read about other people's stupid mistakes, and think "haha, I would never have done that, because I know about XYZ!" And then we go off and happily make some equally stupid mistake because we don't know about ABC.
I turn down a lot of jobs I don't feel confident with; maybe more than I should.
An LLM never will.
Now at a new company i have caught several people copy pasting gpt code that is just horrendous.
It seems like this is where the industry is headed. The only thing i have found gpt to be good at is solving interview questions although it still uses phantom functions about 50% of the time. The future is bumming me out.
However, I don't think it's really bad for the technical industries long term. It probably does mean that some companies with loose internal quality control and enough shiftless employees pasting enough GPT spew without oversight will go to the wall because their software became unmaintainable and not useful, but this already happens. It's probably not hugely worse than the flood of bootcamp victims who wrote fizzbuzz in Python, get recruited by a credulous, cheap or desperate company and proceed to trash the place if not adequately supervised. If you can't detect bad work being committed, that's mostly on the company, not ChatGPT. Yes, it may make it harder, a bit, but it was oversight you should already have been prepared to have, doubly so if you can't trust employee output. It also probably implies strong QA, which is already a prerequisite of a solid product company.
Normal interest rates coming back will cut away companies that waste everyone's time and money by overloading themselves on endlessly compounding technical debt.
Yes, it makes the barrier higher even for good products and helps entrench incumbents, but short of a transnational revolution, the macroeconomic system is what is it and you can only chose to find the good things in it or give up entirely.
Yeah I've seen this and i hate it. If i wanted to know what chatgpt said I'd just ask it myself.
Everyone using C++20 compilers: side-glancing monkey.
Then no one knows what they are doing. I really don't know any company that doesn't make what could be considered rookie mistakes by some armchair "developer" here on HN.
It's actually a tricky bug, because usual tests wouldn't catch it (db wiped for good isolation) and many ways of manual testing would restart the service (and reset the value) on any change and prevent you from seeing it. Ideally there would be a lint that catches this situation for you.
...who didn't know how the ORM they were using worked. That's what makes them look so bad here: nobody knew how it worked, not even at the surface level of knowing what the SQL actually generated by the tool looks like.
The takeaway here is that they weren't mature enough to realize they were, in fact, doing something "weird". I.e. Using UUIDs for PKs, because hey "Netflix does it so we have to too! Oh and we need an engineering blog to advertise it".
Edit. More clarity about why the UUID is my point of blame: If they had used a surrogate and sequential Integer PK for their tables, they would never have to tell SQLAlchemy what the default would be, it's implied because it's the standard and non-weird behavior that doesn't include a footgun.
That is, I'm struggling to understand how a dive into the logs wouldn't show that all of these inserts were failing with duplicate key constraint violations. At that point at least I'd think you'd be able to narrow down the bug to a problem with key generation, at which point you're 90% of the way there.
I also don't agree that "usual tests wouldn't catch it (db wiped for good isolation)". I'd think that you'd have at least one test case that inserted multiple users within that single test.
But it took about 5 minutes, and 4 of those minutes were waiting for Kibana to load.
1. First, if you look at the code they posted, they had the same bug on line 45 where they create new Stripe customers.
2. The issue is not multiple subscriptions per user (again, if you look at the code, you'll see each Subscription has one foreign key user_id column). The problem is if you had multiple subscriptions (each from different users) created from the same backend instance then they'd get the same PK.
Your second point is true, but I don't see what it changes. Most automated unit/integration testing would just wipe the database between tests and needing two subscribed users in a single test is not that likely.
Apparently not.
create subscriptions with and without overlapping effective windows
Those seem like very basic tests that would have highlighted the underlying issue
I'm not talking about 100% branch coverage, but 100% coverage of all happy paths and all unhappy paths that a user might reasonably bump into.
OK maybe not 100% of the scenarios the entire system expresses, but pretty darn close for the business critical flows (signups, orders, checkouts, whatever).
I wouldn't pillory someone if they left out a test case like this, but neither would I assert that a test case like this is for some reason unthinkable or some outlandish edge case.
Volume/load testing (or really, any decent acceptance testing) would catch it.
Why? In fact, not having good isolation would have caught this bug. Generate random emails for each test. Why would you test on a completely new db as if that is what will happen in the real world?
Leaking information between test runs can actually make things pass by accident.
This is not an error that would be difficult to spot in an error aggregator, it would throw some sort of constraint error with a reasonable error message.
Our project was originally full stack NextJS but we wanted to first migrate everything to Python/FastAPI.
Seems like this entire mess could've been avoided if they had stuck with their existing codebase, which seemed to have been satisfying their business requirements.Perhaps they are doing some AI thing and want to have python everywhere.
There is well-hidden vendor-lock when using NextJS, at least.
You will find many issues from GitHub which are not considered because they would make the framework ”better” or easier to use on other clouds.
I can imagine worse, too! They haven't even really started turning that knob yet.
Madden has a monopoly license for NFL content. For a decade the biggest complaint was how they gate kept rosters behind the yearly re-release. Eventually they allowed roster sharing but they put it behind the most god awful inept UI you could possibly imagine such that practically casual gamers wouldn't bother with it.
Then Madden came out with Madden Ultimate Team (like trading cards MTX) and have been neglecting non-MUT modes ever since. They don't explicitly regress the rest of their game, they just commit resources to that effect.
Its like malicious compliance. They don't embrace, extend, extinguish, but they get a similar effect just with resourcing, layoffs, whatever.
Do you mind naming some NextJS alternatives without potential vendor lock in? Mulling a change in my fe.
The SST team actually has an open-next project[1] that does a ton of work to shim most of the Vercel-specific features into an AWS deployment. I don't think it has full parity and it's a third party adapter from a competing host. The fact that it's needed at all is a sign of now closely tied Next and Vercel are.
It is not an issue if you host in Vercel.
Implementing the requested feature would make the framework much better and easier to use when self-hosted elsewhere. But there is neglection to resolve the issue. This is just one case.
The lock-in here is the added developer time and complexity vs. just paying premium.
See CF docs for what the url should be https://developers.cloudflare.com/images/transform-images/tr... Next.js docs should tell you how to write the loader
- They were under large time constraints, but decided a full rewrite to a completely different stack was a good idea.
- They copy-pasted a whole bunch of code, tested it manually once locally, once in production, and called it a day.
- The debugging procedure for this issue so significant it made them dread waking up involved... testing it once and moving on. Every day.
The bug is pretty easy to miss, but should also be trivial to diagnose the moment you look at the error message, and trivial to reproduce if you just try more than once.
Because I can see "create a new subscription" in the manual test plan, but not "create 5x new subscription".
May be a wise thing to do anyway. I would even advocate for the extreme, one instance in sandbox. I've seen quite a variety of bugs that would be easier to detect and debug with one instance — from values accidentally persisted between requests (OP is one example of this very broad class of bugs), to a thundering herd bug catched in staging due to less instances and more worker threads per instance. But can't remember even one "distributed" bug caught in staging.
> they'd also have a dedicated QA person
This is a new feature, and it's user-facing, business-critical and inherently fragile (external service integration, though it broke for another reason). I hope multiple people manually run such things end to end before deploying to prod, even if none of them are dedicated QA?
A lot more than once: they had 40 instances of their app, and the bug was only triggered by getting two requests on the same instance.
A bunch of developers including me once spent a whole weekend trying to reproduce a bug that was affecting production and/or guess from the logs where to look for it. Monday morning, team lead called a meeting, asked for everything we could find out, and… Opened the app in six tabs simultaneously and pressed the button in question in one of the tabs. And it froze! Knowing how to reproduce on our computers, we found and fixed the bug in the next 30 minutes.
This is an entirely forgivable error but should have been found the first time they got an email about it:
"Oh, look, the error logs have a duplicate key exception for the primary key, how do we generate primary keys.... (facepalm)"
Funnily enough, I saw the error in their snippet as soon as I read it but dismissed it thinking there was some new-fangled python feature which allowed that to work like the function def defines that default= accepts only functions so the function gets passed? -- I haven't kept up with modern python and that sounds cool and I figured the bug couldn't be THAT simple.
There's value in having your backtrace surfaced to end users rather than swallowing an exception and displaying "didn't work".
> There's value in having your backtrace surfaced to end users
If they weren't, that should be the first thing you fix.
For something like this where you’re generating a unique id and probably need it in every model, it’s better to write a new Base model that includes things like your unique id, created/changed at timestamps, etc. and then subclass it for every model. Basically, write your model boilerplate once and inherit it. This way you can only fuck this up once and you’re bound to catch it early before you make this mistake in a later addition like subscription management.
This mistake would have happened even if they did not use ChatGPT.
Some linters like pyright can identify dangerous defaults in function call, like `def hello(x=[]): pass` (mutable values shouldn't be a default). Linter plugins for widely-used and critical libraries like SQLAlchemy are nice to have.
The intent is to pass a callable, not to call a function and populate an argument with what it returns.
The time people spend learning the quirks of an ORM is much better put into learning SQL.
I agree that this is probably to their disadvantage, but I would much rather have people admitting their faults than hiding them. If everyone did this the world would be better.
Of course the best solution is to not have faults but that is like saying that the solution to being poor is to have lots of money. It's much easier to say than do.
> Yes we should have done more testing. Yes we shouldn't have copy pasted code. Yes we shouldn't have pushed directly to main. Regardless, I don't regret the experience.
None of those are unconditionally bad! Every project I've worked on could use more testing; we all copy-pasted code at least occasionally, and pushing to main is fine in some circumstances.
The real problem is that they went live, but their tooling (or knowledge how to use it) was so bad it took 5 days to solve the simple issue; and meanwhile, they kept pushing new code ("10-20 commits/day") while their customers were suffering. This is what really causes the reputation hit.
Otherwise, I think this comment thread is a classic example why company engineering blogs choose to be boring. Better ten articles that have some useful information, than a single article that allows the commentariat to pile on and ruin your reputation.
The AI angle is probably why people are piling on. There’s a latent fear that AI will take our jobs, and this is a great way to skewer home that we’re still needed. For now.
The one thing I will say is that it probably wouldn’t take me days to track it down. But that’s only because I have lots of experience dealing with bugs the way that The Wolf deals with backs of cars. When you’re trying to run a startup on top of everything else, it can be easy to miss.
I’m happy they gave us a glimpse of early stage growing pains, and I don’t think this was a PR fumble. It shows that lots of people want what they’re making, which is roughly the only thing that matters.
On the one hand it does seem like a fairly inexperienced organization with some pretty undercooked release and testing processes, but on the other hand all that stuff is ultimately fixable. This is a relatively harmless way of learning that lesson. Admitting a problem is the first step toward fixing it.
A culture of ass-covering is much harder to fix, and will definitely get in the way of addressing these types of issues
Pile-on aside, the problem with this blog article is that it doesn't really have much of a useful takeaway.
They didn't even really talk the offending line in detail. They didn't really talk about what did to fix their engineering pipelines. It was just a story about how they let ChatGPT write some code, the code was buggy, and the bug was hard to spot because they relied on customers e-mailing them about it in a way that only happened when they were sleeping.
It's not really a postmortem, it's a story about fast and loose startup times. Which could be interesting in itself, except it's being presented more as an engineering postmortem blog minus the actionable lessons.
That's why everyone is confused about why this company posted this as a lesson: The lesson is obvious and, frankly, better left as a quiet story for the founders to chuckle about to their friends.
Also the CEO: "remember to be defensive on reddit comments saying how we are a small 1 million dollar backed startup and how it's normal do to this king of rookies mistake to be fast."
Go to logs. Filter by errors. Oh, errors in insert subscription. Seems relevant.
I could understand if the errors were somewhere else.
Even if logs didn't exist. Problematic endpoint generating 50 emails per day? I would have immediately thrown a try catch and rendered the error to the user if logging was impossible. Then your very next bug report solves it.
Assuming that they had the error (guid collision) - it's not as easy to spot as some commentators are making out. But surely after reading the code s few times.
Ironically they should have asked ChatGPT for help debugging
The code snippet has a subtle but significant issue in the default value of the id column. Here is the problematic part: ... In this line, default=str(uuid.uuid4()) is evaluated only once at the time of the class definition, not each time a new StripeCustomer instance is created....
Ensure that the revenue generating codepaths have proper logging.
This failure had very little to do with having an LLM write it.
I expect ChatGPT wouldn't have been able to solve the issue given the entire codebase.
Why are you assuming they didn’t? This is an AI company using AI to build their product and trusting it without proper code review, testing, or guardrails. Clearly they’re all in on AI hype. This took them days to solve, so not only would I bet they asked ChatGPT, I’d wager multiple people tried it multiple times.
Bookmarked for the next time we're told ChatGPT's code error rate is acceptable because we review its code just like an intern's.
I remember making the exact same mistake (accidentally using a single function call in a schema) back in 2010. No LLMs required.
The bigger culprit is probably a lack of testing / debugging. This error would immediately get caught if you simply registered twice on a test instance.
Testing and debugging isn't just a matter of stepping through code, it's an exercise of seeing where your mental model of the codebase is faulty versus the current reality of it.
I've encountered similar problems and they'd be fixed in a matter of hours, not days.
I wonder why this victim didn't ask ChatGPT to identify and fix the problem...
Because if humans had written and reviewed the code, multiple team members would have had to have learned Python, and SQLAlchemy specifically, which, even if the mistake was initially made as many times as ChatGPT did, there would have been multiple independent opportunities for it to be caught and questioned and the relevant knowledge shared during development.
ChatGPT may be able to crank out immense volumes of superficially functional code, but if its your only “teammate” that understands the libraries used and touches the code, its a huge single point of failure.
Friendly reminder: check if your codebase is actually testing this!
One of the interesting consequence of running unit tests with a fresh database everytime is that problems related to unique constraints seldom get caught by unit tests.
So /that's/ where ChatGPT learned it! :)
I had to fix some intern code once, and… well, I can't give too many details, but I will say that an FAQ shouldn't consist entirely of quotations from a TV show from a different country in a language the app doesn't support.
If HN is to be believed you should use ChatGPT to generate your data insertion code, to port it to a different language, and to ask for which error it made when you asked before. What to do if an error slips through in this version is always an exercise left for the reader.
Seems like YC in line with the rest of the world, slap AI on anything and boom there are VCs and cash.
Side story: The company I work for even change the company domain from .com to .ai. Cannot wait for the AI bubble to burst.
but then perhaps coding the entire thing with ChatGPT saved the company more than 10K, so they came out well ahead
or maybe tons of other bugs lurk that will cost the company well over 10K over the long run
As a rookie programmer you might not notice this pattern, but after getting hit a few times you’ll immediately see that. And default to either a factory or to None and then set the value inside the function.
ChatGPT code is only safe to use if you understand it. If you don’t, there’s always the risk that it will bite you.
* you have two competing subscription id columns.
* a uuid is not a string, it is a 128 bit integer. If your database limits to 64 bit integers then use a 64 bit integer for the id instead of a string, or use an array of 128 bytes.
Also, is there a way to set up foreign key constraints on `userId` with this ORM? That seems like another oversight.
Sure you should have tests, sure you shouldn't copy paste code you don't understand and you shouldn't push directly to production.
But, regardless of all that, the main issue of all this incident is not the rookie mistake itself, is how they didn't have logs or alerts and it took them 5 days of customer complaining to find out they had "duplication errors" in the db.
That's the thing that should have been fixed first and extensivly mentioned in the post-mortem
That's kind of rational, id is their own internal id. subscription_id is the id on stripe. A better name would be stripe_id.
Many ORMs represent UUID as a string.
"We're too lazy to write our own code, or even test that it works, please give us money VCs/Users."
Chat GPT isn't to blame here, chat GPT is like a tireless intern. It isn't to blame if the CTO pushes it's code straight to prod.
This is giving me a start up idea, code review as a service!
Anyway I believe the product in question is https://agentgpt.reworkd.ai
I also eventually landed on reworkd.ai after some googling. The blog is called "asim" and the OP's username is "asim-shrestha". That lead me to this: https://www.ycombinator.com/companies/reworkd They are S23, which is mention in the blog.
from https://docs.reworkd.ai/introduction
Whereas the blogpost clearly demonstrates that AI agents cannot be left totally "autonomous", their output might seem reasonable for those not well versed in particular domain but might have disastrous consequences.
VC bros are clearly gambling big on Linear Algebra.
Yes, that's a very fair reasoning. YC did the right thing by investing in this company. Fits very well with the rest of their portfolio.
Also, did anyone click on the double ^ at the bottom, hoping to either go back up to the top or find something about the author's company, only to find that they had upvoted the post by mistake? I'm wondering if (by this count) 167 other people might have done that.
default = python code evaluates the default value as necessary for each new record
server_default = the initial CREATE TABLE uses this computed (from python) value, thus the hardcoded UUID. They also could have done server_default=text("uuid_generate_v4()") if they had that corresponding module installed on postgres.
Separate table definition (DDL) from row-insertion (DML) definitions along with their corresponding defaults.
Does there, though? You can always set `Column(default = lambda: 1234)`
To be explicit: You're not seeing a function called later in python. You're setting an attribute on a database table.
It seems quite unfair to place the blame on SQLAlchemy here, or even Python.
Even a statically typed language wouldn't prevent this kind of issue - the author of the code is the only person who can decide when they mean "use this exact string each time" or "use this function to give me a new string each time".
I suppose a column description API could follow the dataclass style definition, with different argument names for default and default_func. That would (I think) prevent this from happening.
I feel like "terrible ORM" is a tautology.
This issue highlighted in one respect a fundamental problem with all ORMs and why I despise them so much - they're the absolute worst of a leaky abstraction, and as I like to say "they make the easiest stuff a bit easier, and they make the harder stuff way harder". E.g. in this specific case they don't prevent you from needing to know the details of default column value initialization, but apparently they must have also obscured a simple duplicate key constraint to some level that it took 5 days to find the root cause of this bug.
Just learn SQL. You'll need to know it anyway even if you do use an ORM, the basics aren't hard, the skill is much more transferable than the esoteric details of some ORM, and there are lots of good libraries that make dealing directly with SQL easy and safe (slonik for postgres is a favorite of mine, but there are other similar ones).
Yes, but that line is only evaluated once when the class is declared (aka when the application starts), it's not evaluated for every instance.
id = Column(default=str(uuid.uuid4()))
As written, a UUID is generated once and used to set a class-level attribute. Each Python process would generate a unique value, so it wouldn't be immediately obvious. Most of the time Python's ability to run code as a file is loaded is helpful, but this is one the well known gotchas.Although I'm not a SQL Alchemy user, I assume the fix is essentially the same as it would be for Django. So the correct code would have been essentially:
id = Column(default=uuid.uuid4)
Instead of executing `uuid4()` and caching a single UUID value, it would have executed `uuid4()` each time a new object was created.Sure,
```
UUID4 = uuid.uuid4()
def default_id():
return UUID4
```Tbh, if they were writing something for prod they should have factored out the uuid generation and passed values into a pydantic dataclass w/ validators for all the fields. Just sayin.
On the other hand, this could and probably should, have been caught by an automated test that tried to create multiple subscriptions on a single server. Or for that matter , manual testing of creating subscriptions against a local copy. I'm not saying that to be dismissive, but one takeaway you should get from it is the value of testing before putting code in production.
Edit: Another takeaway should probably be that if you have a a major bug like this, and you can't easily reproduce it, you should look harder. I bet there were some logs for errors about constraint violations in the database if you had looked for them.
As to the extra parentheses: I bet that's a force-of-habit thing to prevent potential issues. For example, it seems Sqlite requires them for exactly this kind of default definition[1]. It could also read to nasty bugs when the lack of parentheses in the resulting SQL could result in a different parse than expected[2]. Adding them just-to-be-safe isn't the worst thing to do.
[0]: https://docs.sqlalchemy.org/en/13/core/metadata.html
That's... scary, to put it mildly. I wonder how many of those are fixes to things broken by previous commits. Then again, I work on software where the average is far less than one commit per day, although it's a mature product. Nonetheless, "slow down and think" is probably good advice in this case.
lol "worst".
I think it should be "How we used chatGPT to make a 10k mistake." -- at least then it's honest about the party at fault, that being the startup that didn't vet generative code.
Relatedly i've been throwing AIs at the problem of ordering mods and dependencies for game engines; it's pretty astonishing the error rates you see involving medium-sized text lists.
A good experiment: take a 100 line text file of whatever, ask an AI to sort the lines by some criteria and output a text file, you'll get files back with less than 100 lines routinely.
These kind of things really limit my faith in those systems to work without a heavily leashed supervisor along side.
Keep. Asking. Why.
ChatGPT didn't fail, your system allowed ChatGPT to fail. Answering why is the interesting thing to discuss and blog about.
The issue is tiny and slipery for a very big table. But I am still curious of why test can not find it.
>During the work day, this was fine. We probably committed 10-20 times a day (directly to main of course) which would cause new backend deployments to occur, giving us 40 new IDs for customers to potentially use.
They just use the test env for prod? When to push code, the CICD should be run and some examples should be run too here. And every time, the env should be clean. Here the database does not change from test to production.
That’s always been the case, but there is so much more surface area for human and tool interactions now that we have tools that are so generalized.
Good for them for sharing the story, countless others have them but not sharing them.
id = Column(String, primary_key=True, default=str(uuid.uuid4()), unique=True, nullable=False)
Can be fixed by making the default a function: id = Column(String, primary_key=True, default=lambda: str(uuid.uuid4()), unique=True, nullable=False)
Or, if you choose to use native UUID types in Python and SQL: id = Column(Uuid, primary_key=True, default=uuid.uuid4, unique=True, nullable=False)1. Bunch of tests that simulate exactly the scenario of signups. Hundreds of them actually inserting db records with maybe some kind of dummy stripe code.
2. Logs of the actual uuid for each person.
The second would never have been used since the tests would have caught this bug. But are important anyway.
Seems a bit rich to blame the absence of these two things on chatgpt. That’s just immature engineering practices.
Nothing here really sounds like GPTs fault to me. The issue is something that could easily have been done by a human and missed in PR.
I have no idea how to validate people's anecdotes, though. (To be clear I don't doubt this story at all. But if I set up a site where people could submit stories I wouldn't trust any submissions at face value.)
I'm writing to share a painful lesson learned firsthand about the risks of integrating LLM (Large Language Model) generated code into production systems. Recently, my team and I experienced a catastrophic failure due to an error in code generated by an LLM, which resulted in our site being offline for a staggering 12 hours.
The fallout from this incident was devastating. Not only did we lose valuable revenue and user trust, but the company's stock plummeted by 35% on the second day of trading following our IPO. It's a nightmare scenario no developer ever wants to face.
Here's what happened: in our rush to meet deadlines and optimize processes, we turned to LLM-generated code to expedite development. While it seemed like a shortcut at the time, we failed to thoroughly vet the code for potential flaws and dependencies. Consequently, when an overlooked error surfaced, it triggered a cascading failure that crippled our entire system.
The repercussions of this oversight extend far beyond our organization. It serves as a stark reminder to the entire development community about the inherent risks of relying on AI-generated code in critical production environments. While LLMs are undoubtedly powerful tools, they're not foolproof, and blindly trusting their output can have dire consequences.
In hindsight, I deeply regret the decision to incorporate LLM-generated code without adequate scrutiny. I hope by sharing our experience, others can learn from our mistake and approach the use of AI-generated code with caution.
Let this be a warning to all: while LLMs can be valuable assets in certain contexts, proceed with caution when considering their implementation in production systems. The allure of efficiency must never compromise the integrity and reliability of our codebase.
<mid 2000s product specialist> the dot-com is available!
User generated content in general these days is completely poisoned.
Once I asked the imports are not the issue, It correctly pointed out, and explained the problemetic code at me...
I whish they could ask another question to LLM and have an issue pointed out..
Tell me you had no business being invested in without telling me.
I’m going to be harsh here but I honestly have no clue how else to respond. You wrote your backend in Node/Typescript and then decided to change it to Python. What in the world would make that a good idea? No seriously, there is absolutely nothing sane about that decision. In top of that, you used ChatGPT to do the conversion for some of your DB models, was that just for speed or because you didn’t know what you were doing (new language/framework?).
Also, you say you had credits to burn (god this industry is so messed up sometimes) so why rewrite? Clearly not for cost/performance and Node to Python seems like a very lateral move all things considered.
I’m completely flabbergasted as to why you would rewrite your backend like this.
Check out their comment history to see who invested in them.
Salespeople, executives, engineers, it doesn’t matter.
Every day on HN reminds me a little more of 1998.
Adding monitoring is the last thing you think about when pushing out a prototype, and it's easy to forget that a "prototype that no one will probably use" could cost thousands with accidental infinite loops and bugs like these.
Always set your spending limits!
‘$ EXPORT MAX_REQUESTS=1 gunicorn bear:app’
… which would have in fact fixed your subscription problem, but gave you a new problem :)
Great share!
However, luckily in my case, it was caught immediately in the staging env since collisions caused exceptions.
Realizing when an expression is evaluated is pretty easy to miss. That code is probably live somewhere else right now surreptitiously causing issues.
Though in practice in decent languages it's much less likely you'd write your own `any -> any, any`-typed library for whatever (in this case DB interactions), and use a strongly typed one in which this would at least have been a much more explicit mistake to make.
Generally, static languages will just culturally be less likely to have this kind of invisible "T | (() -> T)" overload.
[1] https://doc.rust-lang.org/std/option/enum.Option.html#method...)
At least you got away "easy". I'm waiting for the "...cost us $100k..." post.
Oh.. this is not good. How did you see it worked fine? You did not try inserting new customers?
Very believable.
Having insufficient testing and 0 monitoring does not improve it.
Y'all need to hire some senior backend Devs.
> This problem became really well hidden because of our backend setup. We had eight ECS tasks on AWS, all running five instances of our backend (overkill, yes we know, but to be fair we had AWS credits).
I mean sure, you acknowledge that it's overkill, but my word is that OVERKILL. You're servicing customers numbering in the double digits and you're using more cloud resources than could run entire established businesses. I feel like a lot of developers today have totally lost sight of what computers are capable of, and just over-provision (and overcomplicate) as a default approach. This is scary.
I use ChatGPT constantly and it is not the type of error it would make. It is such a common pattern.
And if you ask GPT-4o whether the code is correct, it is able to spot the issue.
They were having it translate NextJS code to Python, so the prompt probably included their NextJS code (actually, since they’d never turned on the feature that led to them realizing the problem in NextJS, and maybe didn't have enough volume to hit it on the other pathways that the Python code had it on, it’s not implausible the same bug existed in their NextJS code but was never triggered, and ChatGPT just translated the bug. But in any case, their prompt would include their proprietary code to translate.)
Edit: the entire subdomain gives a 404 actually.
One thing is to healthly discuss how dangerous can be assuming a GPT-generated code is safe or not. Or how unit tests could identify this (could really in this specific case?). Or why you need oncall shifts and good alerts. But, come on.
I spotted the issue in the code at first sight, but that doesn't make me morally superior, nor smart enough to blame someone to publicly talk about their mistake. It only means I'm currently reading that kind of code a lot, and I know where to look at. Pass me some clever ARM code and I'll be unable to spot even the most superfluous mistake.
It seems HN is crowded by the most smart guys on the planet, who never had dumb mistakes and are SO "quality inclined" they need to blame someone for theirs.
edit: of course, the decision to make it public is questionable, but that's topic for another thread, IMHO.
Edit: nothing works at all.
404 ʕノ•ᴥ•ʔノ ︵ ┻━┻ It looks like this page doesn't exist. Let's get you back home.
though a little logging would also have gone a long way. I don't really get how this could have taken 5 days to find, since they knew exactly where the problem was
I have no idea why people are letting chatGPT do anything without pouring over everything it says first, and at that point why bother.
I understand time constraints, but there should be a law or something forbidding using whatever any GPT vomits. In fact, many does it blatantly, so my employer banned using GPT codes and only gave us access to it for /entertainment/ usage.
Or more specially, given the context: "We were in a rush to translate a bunch of code and ChatGPT was doing such an impressive job helping that we became complacent and forgot that it just parrots back text it has seen before with something that looks like intelligence but without actual comprehension. So when it copied a common bug, we weren't paying enough attention to catch it."
Meanwhile, I go through the tedious process of understanding ChatGPT's code letter-by-letter, also reading the docs, searching StackOverflow, even offering and rewarding bounties on StackOverflow, all to see if the code makes a shred of sense.
Yeah just a secure shell what ever could go wrong with code you don't understand? Probably nothing much right?
the tricky part , really , is the syntax for ssh config but again that's easy to understand even if it's very hard to write.
I agree, the headline is... misleading. Yes, ChatGPT made a mistake, but the issue is multiple. Perhaps it would be better if the headline was something like "A ChatGPT mistake taught us a $10k lesson".
I love using ChatGPT as much as the next person, but I also have run into more than a few circumstances on simple code where it's simply... made shit up. Like AWS (Boto3) functions that don't exist at all. So any code that comes out of it gets tested and understood by me. I'll ask it to explain and dig into the docs when it does things I don't understand.
That being said, the valuable lesson is in QA, debugging, logging and alerting. It's something that isn't a surprise a small (couple person, few months) startup would have done well. Often the developers of these projects are DEVELOPERS and not DevOps/SysAdmins/DBAs. The code gets written like developers do and not instrumented like a DevOps engineer would. Most get away with this for a long time (honestly, most companies get away with this for far too long).
So great write up, good lesson.
No tool can do everything
I'm surely not using it to ask for just the date yet still it should know such a simple thing as current date if asked..ppl expect it to like Siri and Alexa can tell u it..users expect the same UX and way better!
That is because it is only trained on info up to a certain date
It's good to know what tools are good and bad at
User experience is utmost important without it Apple wouldnt be apple and chatGPT as it's current trajectory is going looks to become as big as the current biggies. Thus everyone including parents/grandparents/etc will be using it and those users dont care about LLM talk and etc.. they have no idea what it is. chatGPT just works for them and somewhat magically so.. it even knows the current date as mom, dad or grandma be like this thing doesnt know the current date Siri does.
You must not be a UX professional ;-)
Chatgpt can do many things that siri and alexa cannot: draft emails, write code, write unit tests, help me brainstorm, give me cooking recipes, or even talk about relationship issues. Every tool has strengths and weaknesses