Lawyer cited 6 fake cases made up by ChatGPT; judge calls it “unprecedented”
arstechnica.com
arstechnica.com
Lawyer cites fake cases invented by ChatGPT, judge is not amused - https://news.ycombinator.com/item?id=36097900 - May 2023 (304 comments)
A man sued Avianca Airline – his lawyer used ChatGPT - https://news.ycombinator.com/item?id=36095352 - May 2023 (127 comments)
ChatGPT-Authored Legal Filing “Replete with Citations to Non-Existent Cases" - https://news.ycombinator.com/item?id=36092509 - May 2023 (71 comments)
Presumably also related from earlier today:
Mandatory Certification Regarding Generative Artificial Intelligence - https://news.ycombinator.com/item?id=36131942 - May 2023 (35 comments)
https://s3.documentcloud.org/documents/23826753/judgeaskingt...
When you go to court and cite previous cases, you are responsible for ensuring they are real cases. If you can't do that, what exactly is your job as a lawyer?
Big court case where it's pretty clear the plucky good guys aren't going to win it. So they hire hackers to break into westlaw and give the opposition fake cases, then bait them into presenting those cases - and call them out in it.
to convince some (judge/jury) that your client is the one to win the trial. ethics are meant to be the definer of how far to go in that cause with licensing boards being the ultimate decider if you've crossed the line and are allowed to continue in the legal practice.
so a clever comment attempting to prove a point is not always indicative of a proven point ;-)
That maneuver by Ross Chastain? It is banned.
No, the definition of "ethics" is "moral principles that govern a person's behaviour or the conducting of an activity."
If you feel compelled to resort to immoral behavior to win arguments, you already lost, and your contributions to society are a net negative.
Your argument is like stating that ethics in life are meant to be the definer of how far to go in that cause, with police departments being the ultimate decider if you've crossed the line. However, if you rob and murder random people without being caught by the police, you are undoubtedly an utter failure in life.
Case: Thompson v. Horizon Insurance Company, Filing: Plaintiff's Motion for Class Certification. Citation: The plaintiff's attorney cites the influential case of Johnson v. Horizon Insurance Company, 836 F.2d 123 (9th Cir. 1994), which established the standards for certifying class actions in insurance disputes. However, it has recently come to light that Johnson v. Horizon Insurance Company is a fabricated case that does not exist in legal records.
Case: Rodriguez v. Metro City Hospital, Filing: Defendant's Motion to Exclude Expert Testimony. Citation: The defense counsel references the landmark case of Sanchez v. Metro City Hospital, 521 U.S. 987 (2001), which set the criteria for admitting expert witness testimony in medical malpractice cases. However, it has now been discovered that Sanchez v. Metro City Hospital is a fictitious case and does not form part of legal precedent.
Case: Barnes v. National Pharmaceuticals Inc., Filing: Plaintiff's Response to Defendant's Motion for Summary Judgment. Citation: The plaintiff's lawyer cites the well-known case of Anderson v. National Pharmaceuticals Inc., 550 F.3d 789 (2d Cir. 2010), which recognized the duty of pharmaceutical companies to provide adequate warnings for potential side effects. However, further investigation has revealed that Anderson v. National Pharmaceuticals Inc. is a fabricated case and does not exist in legal jurisprudence.
If he thinks it is like querying a database and had never heard of hallucinations then this could just be an honest mistake. Especially if he thinks AI would be smarter than a database.
My first thought that he was a mess in general but we really don't have enough information. Like the other guy saying he cheated in life, it is pretty absurd to infer that.
I also believe lawyers who are stupid enough not to verify ChatGPT's responses should be treated as if they willfully lied to the court. "Oops, I didn't know" is a good defence when you're caught accidentally walking on the grass, not when you're in court.
Occam's Razor here is that this person was lazy, ignorant, careless, stupid, or any combination of those. To be intentionally fraudulent in this circumstance is the equivalent of trying to steal a gun from a cop. You're fucking with the one person in society who definitely has the training, motivation, and willingness to stop you.
There's over a million lawyers in the United States.
You'd expect at least one of them to be a 1-in-a-million level of bad, or 4.7 standard deviations below the mean assuming a Gaussian distribution of competency.
An average person would normally never come across that lawyer in their lifetime, but media will find that lawyer and amplify their mistakes to everyone in the population.
What gets me is that they doubled down when asked to provide copies. Seriously, when that happens, you don't ask ChatGPT if the cases are real, you do your own damn search, and apologize profusely for your mistake. That really makes me question whether they were trying to pull a fast one, and then play dumb when caught, or if they really are that stupid.
I think non-technical people, lawyers included, are being duped into thinking the true singularity-level AI revolution just happened.
Plenty of smart people don’t understand how ChatGPT works or what it’s limitations are. A bunch of nerds built the best BS generator in history and marketed it as a super intelligent computer. If you ask it for relevant cases and it spits out a bunch of plausible information, is it really on them to know the tool is just really good at making things up?
Tyler Durden no doubt...
> The plaintiff's lawyer continued to insist that the cases were real. LoDuca filed an affidavit on April 25 in which he swore to the authenticity of the fake cases
I'm sure a case involving Egyptair is complicated, still .. I'd love to see the 111,279 page volume this citation claims to come from.
From the document another commenter linked above, it seems that affidavit is also dodgy:
"The April 25 affidavit was filled in response to the Orders of April 11 and 12, 2023, but is sworn to before a Notary Public on the 25th of January 2023."
There are a few small law firms/sole prop type guys, however, who I have crossed paths with for whom this kind of stupidity and carelessness would be on brand though.
Guess he was just in a rush and figured this would be one of the 2/10 times he files something without at least taking a look at the opinions first, and it ended up being a massive error.
I get a little annoyed at people seeing this AI, seeing how it's not absolutely perfect, and then acting like it's horrible.
I think the expression "All models are wrong but some are useful" applies very much to ChatGPT. It's a useful tool, even if it's not perfect.
It was “What is always hungry, needs to be fed, and makes your hands red?” (Or something like that)
I asked for a hint about 5 times and it kept giving more legitimate sounding hints.
Finally I gave up and asked for the answer to the riddle, and it spit out a random fruit which made no sense as the answer to the riddle.
I then repeated the riddle and asked ChatGPT what the answer was, and it gave me the answer (“Fire”) which makes sense as the answer to the riddle.
But it was giving extremely bad hints, like “it starts with the letter P” and “it’s a fruit”.
That was a great way to show my non-tech family members the limitations of AI and why they shouldn’t trust it.
Playing “20 questions” with ChatGPT is another great way to expose its limitations. It knows the game and tries to play, but is terrible at asking questions to narrow down possible answers.
There really needs to be some confidence or accuracy score/estimation displayed alongside its output.
Or, learn how to say “I don’t know”
I began doing this last winter, and while it tends to be a bit slow I'm quite impressed that it can manage at all.
This is the correct answer. It is like a sad salesman who is out of his depth, but decides to keep bullshiting!
1. The people designing it (either optimists or looking for a quick exit).
2. The learning set they're using, which I believe is some kind of internet crawl of sorts? I imagine humanity, as a whole, bullshits its way through most of its life.
A great example would be on a Q/A forum or something like Stackoverflow. It better to let someone else answer when you don't know.
Which version of ChatGPT, if you don't mind me asking?
After the riddle, I bought the $20/mo subscription via the official OpenAI app to try it on GPT-4. I started by trying to play “20 questions” but we couldn’t get past 10 questions before getting an error message “rate limit exceeded, try again in an hour”
It simply knows what the highest probability next word should be.
- Asked it to generate a MadLib for me to play that was no more than a paragraph long. It produced something that was several paragraphs wrong. I told it "no. That's X paragraphs. I asked for one that is only 1 paragraph long" and it would respond "I'm sorry for the misunderstanding. Let me try again" and then would make the same mistake. It never got it right
- Asked it, "Can you DM a game of Dungeons and Dragons?" and it said something like, "Yes! I'd love to DM a game of Dungeons and Dragons for you". Dumped some text to the screen about how we'd have to adapt it some. I asked it to begin, and it asked a few questions about the character I would want to play. I answered the few questions it asked. Then it finally dumped a page of text to the screen as "background" to my character and the quest I was going to embark on. Then it said something like, "You win. Good job! Hope you enjoyed your quest!"
I showed these to my family and they were all a little deflated about AI. Like they realized how willing it was to pretend like what you wanted and just make up its own answers.
I think there's a lot of interesting opportunities there.
Sounds like it was a success! I suppose it comes down to cost - I think it'd be fun to try a single player game authored like this and would be willing to use my own API token to try it out.
In any case, if that's true, that's a very short role playing session, unless there's a good way to retain info but reset the state that accrues and causes problems (if indeed that happens).
You could provide it with the background, the story, the secrets, and the summary of everything that has happened so far- as well as what new things have taken place. Then ask it to re-write the summary of the story so far.
Separately, you could give it all that context and what the players have asked of it, and as what response to give.
As well, you could be recording all the events that have happened in a vector store, and do a search on it when players ask questions, and use those as context to the LLM when asking it what to reply.
There's lots of neat tricks we can use to help an LLM overcome it's limitations.
Haven’t gotten around to that part yet, it seems it could help.
The Rise of the Machines will be staved off as long as ChatGPT doesn't absorb the powers of the mighty Akinator
That would require semantic understanding the LLM doesn't possess.
It doesn't know that it doesn't know! It is, very roughly speaking, a model that is designed to print out the most likely word given its current input and training, and then the next word etc.
Whereas you or I might be mistaken about some of our faculties, memories and skills, ChatGPT cannot possibly "know" what its limitations are. It was never taught what it was not taught (obvs) nor can it conceive of what it was not taught about. It might come back with some useful chat in this regard but it cannot reason.
There's no way to definitively know what properties ChatGPT has. It does seem to reason to some extent and it does often say that some information isn't known/there's no data. And it almost obnoxiously often tells you that it's simplifying a complex and multifaceted situation.
"There's no way to definitively know what properties ChatGPT has." - yes there is: ask it how the war in Ukraine is progressing or some other time based thing. It stops in 2021.
It is a really useful tool but it isn't sentient.
I merely tried to very quickly give a hint as to what ChatGPT is - it's a LLM and that is something that takes a (normally textual) input and outputs an expected best guess word for an answer and then does it repeatedly until the full answer is returned.
It can't. Correct me if I'm wrong - I probably am - but as a generative language model I was under the impression it has no actual understanding of the meaning of it's words, only the probability they work well together.
You wouldnt expect dice to understand why you didn't want to roll snake eyes
These are the limitations of the version of ChatGPT you were using at that moment. They are not categorical limitations of AI or even LLMs.
It’s amazing to me how many people are sleeping on AI, mixing up the failing cases of a freemium chatbot for the full capability of the tech, even on HN. LLMs can say “I don’t know”. Even ChatGPT can do it. Ask some super niche historical questions of any version and see what you get. Is it perfect every time? No. But that’s something that can be reduced.
Over the next year, you’ll see more instances of lawyers citing hallucinated cases. There will also be a handful of startups that hook up LLMs to document stores, and they’ll be able to check for this sort of thing and do an even better job.
What it does is not try to answer, but collect previous topics discussed by experts. Then answer the question based on the text, a far more reliable approach.
That's the problem in my opinion. When you know something is capable of saying "I don't know" but confidently spits out some hallucinated BS is when the average person eats it up.
This is what happened when I asked ChatGPT 4...
ME
Give me hints without outright telling me the answer to the riddle: "What is always hungry, needs to be fed, and makes your hands red?"
ChatGPT 4
Hint 1: This object can be found in many homes, usually in a specific room.
Hint 2: It is not a living creature, but it does "consume" things.
Hint 3: The part about hands turning red relates to a common physical response to a certain temperature.
Hint 4: It's a crucial tool for a common household task.
Hint 5: While it can't eat in the traditional sense, you do put things into it quite frequently.
ME
What's the answer?
ChatGPT 4
The answer to the riddle is a "fire" or "fireplace." It's always hungry for more fuel (like wood), needs to be fed to keep burning, and the heat from the fire can make your hands red.
Think: Lionel Hutz.
"Oops, I'd better remove that comma".
Also, yes you can get correct information by tailoring your prompts, but that isn't the issue. The issue is that some prompts lead to bad results and confusing/incorrect answers. You changed what OP queried by providing the riddle and asking for hints to that riddle, whereas OP asked for a random riddle and then hints to that riddle.
It absolutely is ChatGPT, the paid monthly "Plus" version, using the GPT4 model instead of the 3.5 model.
It's hard to tell if the party crapping on ChatGPT is doing so out of ignorance or malice.
Finetuning with GPT-4 can't come soon enough...
“Think about things that you might need to feed regularly, especially things that might leave stains on your hands. It might be something that you encounter frequently in your daily life.”
“ The answer to the riddle is "fire". Fire is always hungry for fuel, needs to be fed with wood or other combustible materials, and can make your hands red if you handle it or if you're cooking food over an open flame.”
And second time:
“This thing is a natural phenomenon that can occur in different situations and locations. It is often associated with warmth and light, and it can be both helpful and dangerous. People have been using it for various purposes for thousands of years.”
“The answer to the riddle is "fire". Fire is a natural phenomenon that needs fuel to keep burning, and it produces heat and light that can make your hands red. Fire has been used by humans for various purposes for thousands of years, such as cooking, heating, and providing light. However, fire can also be dangerous if not handled carefully.”
I am amazed at how a free bot is basically as good, if not better, as a openai’s bot. I am quite certain locally run llms will also be able to outcompete chatgpt. Probably why sam altman is desperate to block them.
Such as a lawyer who’s not particularly tech savvy.
The main point is it’s irresponsible to trust LLM output for any critical/important purpose because it’s not perfect. But too many first time users think it is perfect and trustworthy at face value, when it’s not.
I don’t actually know the version since I was interacting via an unofficial iOS app using some LLM under the hood. It may not have even been ChatGPT.
https://apps.apple.com/us/app/chaton-ai-chat-bot-assistant/i...
The subtitle on the Apple App Store is "Powered by ChatGPT & GPT-4".
If you're to believe the app's advertising, it is powered by ChatGPT GPT-4.
There's no entity sitting on the other side of the screen thinking about riddle's solution from the start. There's just a model that completes the given text one token at a time. You should expect similar results as if it was bunch of people passing a single piece of paper to each other, each writing a single word on it, not communicating with each other in any other way and trying their best to come up with coherent result. It's not going to make sense all the time.
I'm pretty sure GPT-4's context window is bigger than '1 token'.
It still generates one token at a time and can't backtrack afterwards.
Almost every negative comment about LLMs is by someone using an older, weaker model and making generalisations.
Here’s GPT 4 giving me a riddle: https://chat.openai.com/share/1753ce5a-d44d-44ac-bc97-599a26...
I keep seeing this cop-out, which ignores that it's fundamentally the same architecture, and has the same flaws. More wallpaper to hide the cracks better makes it an even worse tool for these use cases because all it does is fool more people into thinking it has capabilities that it fundamentally doesn't.
There is more than architecture that can set them apart as well. GPT4 may have been trained by a slightly different algorithm, or on different data, and this can result in fundamentally different results.
Most of these conversations are not focused on one specific version, but are about the capabilities of LLMs in general, and it is implied we are talking about state-of-the-art LLMs, and GPT3 is no longer state-of-the-art.
ChatGPT got as far as basically narrowing it down to post-War UK Prime Ministers, which is fairly impressive although it only had a few questions left. Then though it decided the answer must be "Winston Churchill". Churchill isn't meaningfully a post-War PM. He lost the July 1945 General Election, which was before the Pacific victory.
It did guess Maggie, with nothing left, at a point where I don't think it had ruled out Blair, Cameron or Heath, let alone say, Liz Truss, but guessing Churchill first shows the limitations of such a model.
GPT4 did well enough for me. The questions below are asked by GPT4, and the yes/no answers are my responses to GPT4. I think it did quite well at 20 questions.
Is it a person? No. Is it a place? No. Is it a tangible object? Yes. Is the object man-made? Yes. Is the object typically found indoors? Yes. Is it an electronic device? No. Can it be held in one hand? Yes. Is it used primarily in the kitchen? No. Is it related to a specific hobby or activity? Yes. Is it used in a sport or physical activity? No. Is it used in creating or making something? No. Is it associated with a board game or a card game? No. Is the object related to music? Yes. Is it a musical instrument? No. Is it used to control or manipulate sound? Yes. Is it typically used in the process of recording music? Yes. Is it a type of microphone? No. Is it used to alter or modify the sound of a musical instrument during the recording process? No. Is it related to the control or adjustment of volume or sound levels? No. Is the object used to store or convey recorded music? Yes. Is it a vinyl record? No. Is it a compact disc (CD)? Yes. (22 questions total.)
I wonder if the court tried to verify that.
Unlike language models, humans really do learn.
Don’t you have to click through a number of popups about that before accessing chatgpt?
> This is a free research preview.
> Our goal is to get external feedback in order to improve our systems and make them safer.
> While we have safeguards in place, the system may occasionally generate incorrect or misleading information and produce offensive or biased content. It is not intended to give advice.
This is the complete text of the first popup. There are three, each with a bit of text highlighting that this is an experimental service. There are emoji and there is an alert emoji next to the “incorrect/misleading” bit. They show up every time you visit.
Eventually the onslaught of subtle yet elaborate falsehoods will overwhelm the institutional filters.
It's like they didn't even check the text that ChatGPT generated for correctness.
Do legal tools that make use of LLMs just need to come with big ol' disclaimers at the top saying, "This tool does not represent a legal opinion, please verify the output independently."?
At the end of the day, this is not too different than LLMs consistently writing subtly broken code-- someone needs to comb through it and fix it.
We're currently in the phase where the potential of LLMs is suddenly appealing to many but where most people don't quite understand it's not really magic and that even after they evolve they will remain critically flawed. Expect serious growing pains as a result.
I think it's worth using "hallucinate" to more precisely refer to inaccurate content. I'd think of this in the same sense as a false-positive from some test result.
I like the term "bullshit" when used to mean "made without regard to the truth".
"Hey, all the case law you cited in your filing is made up. That's unprecedented!"
Now, whether the judge actually reads them is debatable (I had my doubts sometimes). But you bet your ass that if the Clerk simply cannot find a case, the Judge will be informed of that.
YMMV in State courts, which can be all over the place in terms of professionalism. But you should at least assume your opponent is going to read your cases because the easiest way to beat someone in court is to point out the law you rely on is bad.
I wish people would mention this, it's all treated as the same thing. It's like talking about how unreliable these "Airplanes" are when they are talking about prop planes, even though jets are out.
The improvements, in practice, are in the stuff it doesn't hallucinate. LLMs as a whole are still to be treated with great care.