Better Call GPT: Comparing large language models against lawyers [pdf]
arxiv.org
arxiv.org
In terms of contract review, what I've found is that GPT is better at analysis of the document than generating the document, which is what this paper supports. However, I have used several startups options of AI document review and they all fall apart with any sort of prodding for specific answers. This paper looks like it just had to locate the section not necessarily have the back and forth conversation about the contract that a lawyer and client would have.
There is also no legal liability for GPT for giving the wrong answer. So It works well for someone smart who is doing their own research. Just like if you are smart you could use google before to do your own research.
My feelings on contract generation is that for the majority of cases, people are better served if there were simply better boilerplate contracts available. Laywers hoard their contracts and it was very difficult in our journey to find lawyers who would be willing to write contracts we would turn into templates because they are essentially putting themselves and their professional community out of income streams in the future. But people don't need a unique contract generated on the fly from GPT every time when a template of a well written and well reviewed contract does just fine. It cost hundreds of millions to train GPT4. If $10m was just spent building a repository of well reviewed contracts, it would be a more useful than spending the equivalent money training a GPT to generate them.
People ask pretty wide range of questions about what they want to do with their documents and GPT didn't do a great job with it, so for the near future, it looks like lawyers still have a job.
I’ve been trying to pair system design with ChatGPT and it feels just like talking with a person who’s confident and regurgitates trivia, but doesn’t really understand. No sense of self-contradiction, doubt, curiosity.
I’m very, very impressed with the language abilities and the regurgitation can be handy, but is there a single novel discovery by LLMs? Even a (semantic) simplification of a complicated theory would be valuable.
I was still able to rewrite the result into something that more suited me, but for a service with a $150 price tag I kind of hoped it would do more.
Our solution like you point out is more rigid than having a lawyer write it, but for the majority of people having something that is accessible and free is worth it and then having services layer on top makes the most sense. It is easier to have a well written contract that you can "turn on and off" features or sections of the contract than to try to have GPT write a custom contract for you.
The idea of "legalzoom-style" businesses has always seemed like a bamboozle to me. You pay hundreds of dollars for essentially the form documents to fill in, and you don't get any of the flexibility that an actual lawyer gives you.
As another example, Northwest Registered Agent gives you your corporate form docs for free with their registered agent services.
The difference is that when Westlaw presents me with a decision point and I choose the wrong option for my client, my client sues my insurer and is made whole. (My premiums increase accordingly).
If you make the wrong choice in your choose-your-own-legal-adventure, you lose.
(For some contracts, this is probably the right approach).
Pet trusts [1]! My lawyer literally used their existence, which I find adorable, to motivate me to read my paperwork.
[1] https://www.aspca.org/pet-care/pet-planning/pet-trust-primer
Knowing someone who works in Trusts & Estates, that is terrible. I've often heard complaints about drafting by percentages of anything but straight financial assets which have an easily determined value, because that requires an appraisal(s). Yes, there are mechanisms to work it out in the end, but it is definitely better to be able to say $X to Alice, $Y to Bob and the remainder to Claire.
You have to think of not only what you want, but how the executors will need to handle it. We all love complex formulae, but we should use our ability to handle complexity to simplify things for the heirs - it's a real gift in a bad time.
I guess there's an understanding the being an executor is a best-effort role, but maybe you could specifically codify that +/-5% on the individual shares is fine, just to take off some of the burden of needing it to be perfect, particularly if there are payouts occurring at different times and therefore some NPV stuff going on.
Law in general is interpretation. The most "lawyerese" answer you can expect is "It depends". Technically in the US everything is legal unless it is restricted and then there are interpretations about what those restrictions are.
If you ask a lawyer if you can do something novel, chances are they will give a risk assessment as opposed to a yes or no answer. Their answer typically depends on how well they think they can defend it in the court of law.
I have received answers from lawyers before that were essentially "Well, its a gray area. However if you get sued we have high confidence that we will prevail in court".
So outside of the more obvious cases, the actual function of law is less binary but more a function of a gradient of defensibility and the confidence of the individual lawyer.
So much of contract law boils down to confidence in winning a case, or it's a business issue that just looks like a legal issue because of legalese.
If this is dangerous with normal English, how much more so with legal text.
At least if a lawyer drafts the text, there is at least one human with some sort of intentionality and some idea of what they're trying to say when they draft the text. With LLMs there isn't.
(And as I say in the linked post, I don't think that is fundamental to AI. It is only fundamental to LLMs, which despite the frenzy, are not the sum totality of AI. I expect "LLMs can generate legal documents on their own!" to be one of those things the future looks back on our era and finds simply laughable.)
It was my understanding that there is also no legal liability for a lawyer for giving the wrong answer. In extreme cases there might be ethical issues that result in sanctions by the bar, but in most cases the only consequences would be reputational.
Are there cricumstances where you can hold an attorney legally liable for a badly written contract?
Nit: Malpractice insurance is (a species of) E&O insurance.
There is plenty of legal, ethical and professional liability for a lawyer giving the wrong answer, we don't often see the outcome of these things because like everything in the courts they take a long time to get resolved and also some answers are not wrong just "less right" or "not really that wrong."
Look at your most recent engagement letter with an attorney. I’d bet that you agreed to arbitrate all fee disputes, and depending on your state you might have also agreed to arbitrate malpractice claims.
Wrong Word in Contract Leads to $2M Malpractice Suit[1].
[1]https://lawyersinsurer.com/legal-malpractice/legal-malpracti...
About the only case where this works in practice is someone going pro se and using their own toolset to gin up a legal AI model. There's arguably a case for acting as an accelerator for attorneys, but the problem is that if you've got an AI doing, say, doc review, you still need lawyers to review not just the output for correctness, but also go through the source docs to make sure nothing was missed, so you're not saving much in the way of bodies or time.
I think you will find that this is because they "outsource" the AI contract document review "final check" to real lawyers based in Utah ... so, it's actually a person, not really a wholy-AI based solution (which is what the company I am thinking of suggests in their marketing material)
Which company is that? I don't see any point in obfuscating the name on a forum like this.
I mean, i get your point but lets be real: I cannot count the number of times I sat in a meeting and looked back at a contract and wished some element had a different structure to it. In law there are a lot of "wrong answers" someone could foolishly provide, but its way more often something more variable as to how "wrong" the answer is, than it is a binary bad/good piece of advice.
I personally feel the ability to have more discussion about a clause is extremely helpful, v's getting the a hopefully "right answer" from a lawyer, and counting the clock / $ as you try wrap your head around the advice you're being given. If you have deep pockets, you invite your lawyer to a lot of meetings, they have context and away you go....but for a lot of people, you're just involving the lawyer briefly and trying to avoid billable hours. That's been me at the early stage of everything, and it's a very tricky balance.
If you're a startup trying to use GPT, i say do it, but also use a lawyer. Augmenting the lawyer with GPT to save billable hours so you can turn up to a meeting with your lawyer and extract the most value in the shortest time period seems like the best play to me.
> I cannot count the number of times I sat in a meeting and looked back at a contract and wished some element had a different structure to it.
The only way to have something "bullet proof" is to have experience in ways in which something can go wrong. Its just like writing a program in which the "happy path" is rather obvious but then you have to come up with all the different attack vectors and use cases in which the program can fail.
The same is with lawyers. Lawyers at big firms have the experience of the firm to guide them on what to do and what they should include in a contract. A small town family lawyer might have no experience in what you ask them to do.
Which is why I advocate for more standardized agreements as opposed to one off generated agreements (with GPT or a lawyer). Think of the YCombinator SAFE, it made a huge impact on financing because it was standardized and there were really no terms to negotiate compared to the world before which the terms Notes were complex had to be scrutinized and negotiated.
> Augmenting the lawyer with GPT to save billable hours so you can turn up to a meeting with your lawyer and extract the most value in the shortest time period seems like the best play to me.
The issue is that a lot of lawyers have a conflict of interest and a "Not invented here" way of doing business. If you have a Trust for instance written by one lawyer and bring it to another lawyer, the majority of lawyers we talked to actually prefer to throw out the document and use their own. This method works well if you are a smart savvy person, but for the population at large, people have some crazy and weird ideas about how the law works and need to be talked out of what they want into something more sane.
Another common lawyer response besides "It depends" is "Well you can, but why would you want to?" So many people of a skewed view on what they want and part of a lawyers job is interpreting what they really want and guiding them on the path of that.
So the hybrid method really only works if you find a lawyer that accepts whatever crazy terms you came up with and are willing to work with what you generated.
When i suggest going down a hybrid path, I mostly mean use GPT on your own (disclose this to your lawyer at your own risk) as a means to understand what they're proposing. I've spent so many hours asking questions clarifying why something is done a certain way, and most of that is about understanding the language and trying to rationalize the perspective the lawyer has taken. I feel I could probably have done a lot of that on my own time, just as fast, if GPT had been around during these moments. And then of course, confirmed my understanding aligns with the lawyer at the end.
I need to be upfront, I really don't know I'm right here....its just a hunch and gut reaction to how I'd behave in the present moment, but I find myself using AI more and more to get myself up to speed on issues that are beyond my current skill level. This makes me think law is probably another good way to augment my own disadvantages in that I have a very limited understanding of the rules and exceptional scenarios that might come up. I also find myself often on the edge of new issues, trying to structure solutions that don't exist or are intentionally different to present solutions...so that means a lot of explaining to the lawyer and subsequently a fair bit of back and forward on the best way to progress.
It's a fun time to be in tech, I'm hoping things like GPT turn out to be a long term efficiency driver, but I'm fearful about the future monetization path and how it'll change the way we live/work.
I notice same things in other professions, especially where it requires a huge upfront investment in education.
For example (at least where I live), there was a time about 20 years ago where architects also didn't want to produce designs that would be then sold to multiple people for cheap. The thinking was that this reduces market for architecture output. But of course it is easy to see that most people do not really need a unique design.
So the problem solved itself because the market does not really care and the moment somebody is able to compile a small library of usable designs and a usable business model, as an architect, you can either cooperate to salvage what you can or lose.
I believe the same comes for lawyers. Lawyers will live through some harsh time while their easiest and most lucrative work gets automated and the market for their services is going to shrink and whatever work is left for them will be of the more complex kind that the automation can't handle.
When you can hire less lawyers and get more work done and cheaper and at the same (or better) quality, you are going to upend the market for lawyering services.
And this does not require to replace lawyers. It is just enough to equip a lawyer with a set of tools to help them quickly do the typical tasks they are doing.
I work a lot with lawyers and a lot of what they are doing is going to be stupidly easily optimised with AI tools.
One of the top comments on this thread says that LLMs are going to better at summarizing contracts than generating them. I've heard this in legal tech product demos as well. I can see some utility to that--for example, automatically generating abstracts of key terms (like term, expiration, etc.) for high-level visibility. That said, I've been told by legal tech providers that LLMs don't do a great job with some basic things like total contract value.
I question how the document summarizing capabilities of LLMs will impact the way lawyers serve business organizations. Smart businesspeople already know how to read contracts. They don't need lawyers to identify / highlight basic terms. They come to lawyers for advice on close calls--situations where the contract is unclear or contradictory, or where there is a need for guidance on applying the contract in a real-world scenario and assessing risk.
Overall I'm less enthusiastic about the potential for LLMs in the legal space than I was six months ago. But I continue to keep an eye on developments and experiment with new tools. I'd love to get some feedback from others on this board who are knowledgeable.
As a side note, I'm curious if anyone knows about the impact of context window on contract interpretation a lot of contracts are quite long and have sections that are separated by a lot of text that nonetheless interact with each other for purposes of a correct interpretation.
I would add, that sometimes being a newcomer is a benefit. Many times a particular industry is stuck in groupthink, having shared understanding on what is and what isn't possible. And it sometimes requires a person with a fresh perspective to be able to upend it. See example of Elon upending multiple industries by essentially doing exactly that.
It’s just like programmers and artists. As the tools improve, you’ll need fewer, smarter humans.
AI has outperformed radiologists for a while now, and I don't care how much better AI performs, radiologists aren't going away.
Radiologists get to decide the laws of the field essentially.
Why would they vote to kick themselves out of the highest paying job in the world?
Like the medical industry, the legal industry is designed around costing you as much as possible - not really anything related to your benefits.
I just can't see disruption here. The industry will wail against it tooth and nail at every chance.
Refer me to the evidence where AI is outperforming radiologists in the entire body of cross-sectional imaging. Are you seriously claiming that there is AI technology today that can take a brain MRI or CT A/P and produce a full findings and impression without human intervention? You have a reference for that?
Yet this is not an option for patients, and likely won't be in the next 10 years.
Happening in Austria since 2018-ish: https://contextflow.com/
As for Chest Xrays https://pubs.rsna.org/doi/10.1148/radiol.231236 AI has not even demonstrated superiority outside of "routine".
Maybe revisit this in a decade. But your entire argument was that we still have radiologists because of protectionism (and the lawyer case isn’t any more load bearing either). Seems a bit uninformed and premature.
Then another business that wants to compete with you will no longer have an option, they will have to do this or more to be able to stay in the business at all.
We wanted to make sure there would be no cross pollination between legal advisory services and other professional services in most jurisdictions, but the only thing that division did was significantly restrict our ability to widen our service offerings to provide more value.
The end result is that we protected our little nest egg while our share of the professional services pie has been getting eaten by consulting and multi-service accounting firms for the past 20 years.
Doctors, for instance. You hear no end of stories about how incredibly high pressure medicine is with insane hours and stress, but will they increase university placements so they can actually hire enough trained staff to meet the workload? Absolutely fkn not, that would impact salaries.
I think this overlooks a big part of how the legal market works. Our easiest work is only lucrative because we use it to train new lawyers, who bill at a lower rate. To the extent the easy stuff gets automated, 1) it’s going to be impossible to find work as a junior associate and 2) senior attorneys will do the same stuff they did last year. If there’s a decrease in prices for a while, great, but a generation from now it’s going to be a lot harder to find someone knowledgeable because the training pathway will have been destroyed.
More that lawyers will hoard contracts but there will be financial incentives to defect early (the idea being to sell your contracts while the price is high).
Then that lawyer's contracts will become the mass-produced "good enough" boilerplate and be used widely driving down the value of hoarding contracts at all.
Unfortunately people stop at step #1, they use Google and that is their research. I don't think ChatGPT is going to be treated any different. It will be used as an oracle, whether that's wise or not doesn't matter. That's the price of marketing something as artificial intelligence: the general public believes it.
That's a concise explanation that also applies to GPTs and software engineering. GPT4 boosts my productivity as a software engineer because it helps me travel the path quicker. Most generated code snippets need a lot of work because I'm prodding it for specific use cases and it fails. It's perfect as an assistant though.
How is that good for the end user? Malpractice claims are often all that is left for a client after the attorney messes up their case. If you use a GPT, you wouldn't have that option.
In our research, we found out that most everyone has the same questions: (1) "what does my contract say?", (2) "is that standard?", and (3) "is there anything I can/should negotiate here?"
Most people don't want an intense, detailed negotiation over a lease, or a SaaS agreement, or an employment contract... they just want a normal contract that says normal things, and maybe it would be nice if 1 or 2 of the common levers were pulled in their direction.
Between the structure of the document and the overlap in language between iterations of the same document (i.e. literal copy/pasting for 99% of the document), contracts are almost an ideal use-case for LLMs! (The exception is directionality - LLMs are great at learning correlations like "company, employee, paid biweekly," but bad at discerning that it's super weird if the _employee_ is paying the _company_)
What's your objection to Nolo Press? They seem to have already done that.
They’re going to be great for a lot of stuff. But when it comes to things like the law the other 20% is not optional.
I'm working on Hotseat - a legal Q&A service where we put regulations in a hot seat and let people ask sophisticated questions. My experience aligns with your comment that vanilla GPT often performs poorly when answering questions about documents. However, if you combine focused effort on squeezing GPT's performance with product design, you can go pretty far.
I wonder if you have written about specific failure modes you've seen in answering qs from documents? I'd love to check whether Hotseat is handling them well.
If you'r curious, I've written about some of the design choices we've made on our way to creating a compelling product experience: https://gkk.dev/posts/the-anatomy-of-hotseats-ai/
If your focus is narrow enough the vanilla gpt can still provide good enough results. We narrow down the scope for the gpt and ask it to answer binary questions. With that we get good results.
Your approach is better for supporting broader questions. We support that as well and there the results aren’t as good.
Specific failure modes can be something as simple as extraction of beneficiary information from a Trust document. Sometimes it works, but a lot of times it doesn't even with startups with AI products specific to extracting information from documents. For example it will have an incomplete list of beneficiaries, or if there are contingent beneficiaries, it won't know what to do. Not even a hard question about the contingency. Just making a simple list with percentages of if no-one dies what is the distribution.
Further trying to get an AI to describe the contingency is a crap shoot.
While I expect these options to get better and better, I have fun trying them out and seeing what basic thing will break. :)
If the example is representative, I see two problems: a simple extraction of information that is laid out bare (list of beneficiaries), and reasoning to interpret the section of contingent beneficiaries and connect it facts from other parts. Is that correct?
If that's the case, then Hotseat is miles ahead when it comes to analyzing regulations (from the civil law tradition, which is different from the US), and dealing with the categories of problems you mentioned.
The legal liability is an issue in several countries but contract generation can also be. If you are providing whatever is defined as legal services and are not a law firm, you will have issues.
[1]legalreview.ai
> If you are providing whatever is defined as legal services and are not a law firm, you will have issues.
that is a big reason why we haven't integrated AI tools into our product yet. Currently our business essentially works as a free product that is the equivalent of a "stationary store" of you are filling out a blank template and it is your responsibility what happens. This has a long history of precedence since for decades people could buy these templates off the shelf and fill them out themselves.
Giving a tool to our users to answer legal questions opens a can of works like you say. We decided that the stationary store templates are a commodity and should be free (even though our competitors charge hundreds for them) so we make money providing services on top of it.
Do you have any views on whether context window limits the ability of LLMs to provide sound contractual interpretations of longer contracts that have interdependent sections that are far apart in the document?
Has your level of optimism for the capabilities of LLMs in the legal space changed at all over the past year?
You mentioned that lawyers hoard templates. Most organizations you would have as clients (law firms or businesses) have a ton of contracts that could be used to fine tune LLMs. There are also a ton of freely available contracts on the SEC's website. There are also companies like PLC, Matthew Boender, etc., that create form contracts and license access to them as a business. Presumably some sort of commercial arrangement could be worked out with them. I assume you are aware of all of these potential training sources, and am curious why they were unsatisfactory.
Thanks for any response you can offer.
To answer some of your questions:
- contract review works very well for high volume low risk contract types . Think slas, SaaS… these are contracts comercial legal teams need to review for compliance reasons but hate it.
- it’s less good for custom contracts
- what law firms would benefit from is just natural language search on their own contracts.
- it also works well for due diligence. Normally lawyers can’t review all contracts a company has. With a contract review tool they can extract all the key data/risks
- LLM doesn’t need to provide advice. LLM can just identify if x or y is in the contract. This improving the process of review.
- context windows keep increasing but you don’t need to send the whole contract to the LLM . You can just identify the right paragraphs and send that.
- things changes a lot in the past year. It would cost us $2 to review a contract now it’s $0.2 . Responses are more accurate and faster
- I don’t do contract generation but have explored this. I think the biggest benefit isn’t generating the whole contract but to help the lawyer rewrite a clause for a specific need. The standard CLM already have contract templates that can be easily filled in. However after the template is filled the lawyer needs to add one or two clauses . Having a model trained on the companies documents would be enough.
Hope this helps
Do you think LLMs have meaningfully greater capabilities than existing tools (like Kira)?
I take your point on low stakes contracts vs. sophisticated work. There has been automation at the "low end" of the legal totem pole for a while. I recall even ten years ago banks were able to replace attorneys with automations for standard form contracts. Perhaps this is the next step on that front.
I agree that rewriting existing contracts is more useful than generating new ones--that is what most attorneys do. That said, I haven't been very impressed by the drafting capabilities of the LLM legal tools I have seen. They tend to replicate instructions almost word for word (plain English) rather than draw upon precedent to produce quality legal language. That might be enough if the provisions in question are term/termination, governing law, etc. But it's inadequate for more sophisticiated revisions.
For rewriting contracts keep in mind that you don't have to actually use the LLM to generate the text completely. It is helpful if you can feed all the contracts of that law firm into a vector db and help them find the right clause from previous contracts. Then you can add the LLM to rewrite the template based on what was found in the vector db.
Many lawyers still just use folders to organize their files.
That's a trap: If you don't have prior expertise then you can't distinguish plausible-sounding fact from fiction. If you think you are "smart", then afaik research shows you are easier to fool because you are more likely to think you know.
Google finds lots of mis/disinformation. GPTs are automated traps: they generate the most statistically likely text from ... the Internet! Not exactly encouraging. (Of course, it really depends on the training data.)
>>>"Laywers hoard their contracts and it was very difficult in our journey to find lawyers who would be willing to write contracts we would turn into templates because they are essentially putting themselves and their professional community out of income streams in the future.
This is why Lawyers must die.
If "letter of law" is a thing - then "opinions" in legal parlance, should be BS.
We should have every SC decision eviscerated by GPTs
Any any single person saying "You need a lawyer to figure this out" is a grifter and a charlatan
--
Anyway - I'd like to know more about your startup. I'd like to know what you o for lkegal DMS, ala hummingbird for biotech, but there are so many real applications of GPT to legal troves, such as an auto index/summary/ parties/contacts/dates/ blah blah that GPTs make legal searching all that more powerful and ANY legal-person in any position telling you anything bad about computers keeping them in check is a fraud.
The legal industry is the most low hanging fruit for GPTs.
A wave of (better) legally informed common-person is coming, and I couldn't be more excited!
I've already used some LLMs to ask questions about licenses and legal consequences for software related matters, and it gave me a base, without having to involve a very expensive professional into it for what are mostly questions for hobby things I'm doing.
If there was a significant amount of money involved in the decision, though, I will of course use the services of a professional. These are the kinds of topics you can't be "mostly right".
They'd be sued out of existence.
"In terms of being convincingly wrong, it's not like lawyers never make mistakes..."
They have malpractice insurance, they can potentially defend their position if later sued, and most importantly they have the benefit of appeal to authority image/perception.
I guess you'd have to have some way of knowing that the "malpractice insurance ID" that the GPT gave you at the start of the session was in fact valid, and with an insurance company that had the resources to actually cover if needed...
In the tests they are shown to be pretty close. The point I made wasn't about more mistakes, but about other factors influencing liability and how it would be worse for AI than humans at this point.
This is the key point. Even if assume the AI won't get better, the liability and insurance premiums will likely become similar in very near future. There is a clear business opportunity that's there in insuring AI lawyer.
If a LLM can pass the bar, and has a corpus of legal work instantly accessible, what prevents the deployment of the LLM (or other AI structure) to provide legitimate legal services?
If the AI is providing legal services, how do we assign responsibility for the work (to the AI, or to its owner)? How to insure the work for Errors and Omissions?
More practically, if willing to take on responsibility for yourself, is the use of AI going to save you money?
The law, which you can bet will be used with full force to prevent such systems from upsetting the (obscenely profitable) status quo.
Basically, add a "validate" step. So, you'd first chat with the LLM, create conclusions, then vet those conclusions with an expert specially trained to be skeptical of LLM generated content.
I would be shocked if there aren't law agencies that aren't already doing something exactly like this.
When your attorney is wrong, you get to point at the attorney and show a good faith effort was made.
Hacks are fun, just keep in mind the domain you're operating in.
And possibly sue their insurance to correct their mistakes.
Maybe that’s where legal AI will find the most demand.
https://www.forbes.com/sites/mollybohannon/2023/06/08/lawyer...
So many people continually use arguments that revolve around 'I used it once and it wasn't the best and/or me things up', and imply that this will always be the case.
There are many solutions already for knowledge editing, there are many solutions for improving performance, and there will very likely continue to be many improvements across the board for this.
It took ~5 years from when people in the NLP literature noticed BERT and knew the powerful applications that were coming, until the public at large was aware of the developments via ChatGPT. It may take another 5 before the public sees the developments happening now in the literature hit something in a companies web UI.
We've already seen a fair bit of stagnation in the past year as ChatGPT gets progressively worse as the company is more focusing on neutering results to limit its exposure to legal liability.
https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboar..., In blinded human comparisons, newer models perform better than older ones.
Edit - the website finally loaded for me and while their methodology is listed, the actual prompts they use are not. The only example prompt is "correct grammar: I are happy". Which doesn't do anything at all to assess what we're talking about, which is ChatGPT's inability to deal with subjects which are "risky" (where "risky" is defined as "Americans think it's icky to talk about").
Worse overall? You can use chatgpt 4 and 3.5 side by side and see an obvious difference.
Your specific example seems fairly reasonable. Is there liability in saying x bolt can handle y torque if that ended up not being true? I don't know. What is that bolt causes an accident and someone dies? I'm sure a lawyer could argue that case if ChatGPT gave a bad answer.
If windows 11 is far worse in many metrics than windows XP or Linux, does that mean that technology is useless?
It's one instance of something with a very particular vision being imposed. Windows 11 being slow due to reporting several GB of user data in the first few minutes of interaction with the system does not mean that all new OS are slow. Similarly, some older tech in a web UI (ChatGPT) for genAI producing non-physical data does not mean that all multimodal models will produce data unsupported by physics. Many works have already shown a good portion of the problems in GPTs can be fixed with different methods stemming from rome, rl-sr, sheavNNs, etc.
My point isn't even that certain capabilities may get better in the future, but rather that they already are better now, just not integrated into certain models.
It also may take 10, 20, 50, or 100 years. Or it may never actually happen. Or it may happen next month.
The issue with predicting technological advances is that no one knows how long it'll take to solve a problem until it's actually solved. The tech world is full of seemingly promising technologies that never actually materialized.
Which isn't to say that generative AI won't improve. It probably will. But until those improvements actually arrive, we don't know what those improvements will be, or how long it'll take. Which ultimately means that we can only judge generative AI based on what's actually available. Anything else is just guesswork.
My concern is we're going to get to a place where we think the machines can just take over all important professions, but they're not quite there yet, however people don't bother learning those professions because they're a career dead end and then we just end up with a skill shortage and mediocre services, when something goes wrong, you just have to trust "the machine" was correct.
How do we avoid this? Almost like we need government funded "career insurance" or something like this.
I feel LLMs are great at suggestions that you follow up yourself (if only for sanity checking, but nothing you wouldn't do with a human too).
In my experience, this does not get you close to what the top-level comment is describing. But it gets around the "nerfing" you describe
Our core business is legal document generation (rule based logic, no AI). Since we already have the users' legal documents available to us as a result of our core business, we are perfectly positioned to build supplementary AI chat features related to legal documents.
We recently deployed a product recommendation AI to prod (partially rule based, but personalized recommendation texts generated by GPT-4). We are currently building AI chat features to help users understand different legal documents and our services. We're intending to replace the first level of customer support with this AI chat (and before you get upset, know that the first level of customer support is currently a very bad rule-based AI).
Main website in Finnish: https://aatos.app (also some services for SE and DK, plus we recently opened UK with just a e-sign service)
Given your ownership in a company and real estate, a lasting power of attorney is a prudent step. This allows you to appoint PARTNER_NAME or another trusted individual to manage your business and property affairs in the event of incapacitation. Additionally, it can also provide tax benefits by allowing tax-free gifts to your children, helping to avoid unnecessary inheritance taxes and secure the financial future of your large family.
Uhh... What are the privacy implications here?!
In any case, all startups today are created on top of a mountain of cloud services. Any one of those services can leak private user data as a result of outsider hack or insider attack or accident. OpenAI is just one more cloud service on top of the mountain.
If the current pricing would be $500 an hour for a real lawyer, and at some point your costs are just keeping services up and running, how big cut will you take? Because it is enough if you are only a little cheaper than the real lawyer to win customers.
There is an upcoming monopoly problem, if the users get the best information from the service after they submit all their documents. And soon the normal lawyer might be competitive enough. I fear that the future is in the parent commenter’s open platfrom with open models and the businesses should extract money from some other use cases, while for a while, you get money momentarily based on the typical ”I am first, I have the user base” situation. It is interesting to see what will happen to lawyers.
Zero. We're providing the AI chat for free (or free for customers who purchase something from us, or some mix of those 2 choices). Our core business is generating documents for people, and the AI chat is supplementary to the core business.
It sounds like you're approaching the topic with the mindset that lawyers might be entirely replaced by automation. That's not what we're trying to do. We can roughly divide legal work into 3 categories:
1. Difficult legal work which requires a human lawyer to spend time on a case by case basis (at least for now).
2. Cookie cutter legal work that is often done by a human in practice, but can be automated by products like ours.
3. Low value legal issues that people have and would like to resolve, but are not worth paying a lawyer for.
We're trying to supply markets 2 and 3. We're not trying to supply market 1.
For example, you might want a lawyer to explain to you what is the difference between a joint will and an individual will in a particular circumstance. But it might not be worth it to pay a lawyer to talk it through. This is exactly the type of scenario where an AI chat can resolve your legal question which might otherwise go unanswered.
That is the cynical future, however, and based on the evolution speed of the last year, it is not too far away. We humans are just interfaces for information and logic. If the chatbot has the same capabilities (both information and logic, and natural language), then they will provide full automation.
The natural language aspect of AI is the revolutionary point, less about the actual information they provide. Quoting Bill Gates here, like the GUI was revolutionary. When everyone can interact and use something, it will remove all the experts that you needed before as middle man.
I uploaded all of my bloodwork tests and my 23andme data to Chat GPT and it was better at analyzing it than my doctor was.
Did you do anything special to achieve this? What were the results like?
LLMs don't have to compete against the cutting edge of human professional knowledge. They only have to compete against the disinterested, arrogant, greedy, and overworked professionals that are actually available to people in practice. No wonder they're winning.
The more common, for individual purchasers of legal services, lawyering is going to be family law matters, criminal law matters, and small claims court matters. I can not see a time in the near future where an LLM can handle the fact specific and circumstantial analysis required to handle felony criminal litigation, and I see nothing that would imply LLMs can even approach the individualized, case specific and convoluted family dynamics required for custody cases or contested divorces.
I’m not unwilling to accept LLMs as a tool an attorney can use, but outside of more rote legal proof reading I don’t think the technology is at all ready for adoption in actual practice.
Humans are pretty bad at this. Based on the results, it seems the judges' personal views and emotions are a large part of these cases. I'm not sure what they would look like without emotion, personal views, and the case law built off of those.
At least until the LLMs surpass humans at being charismatic, but that would seem to be its own nightmare scenario.
Look into "virtual influencers". Sounds like you should find it interesting.
That's a completely separate question. We're talking about automating lawyers, not judges. (to be a good lawyer in such a situation, you would need to model the judge's emotions and use them to your advantage. Probably AIs can do this eventually but it's not easy or likely to happen soon)
Language models are great at digesting legalese and analyzing what's going on. However, many legal applications involve around pretty important decisions that you don't want to get wrong ("am I contractually covered with my current insurance program?"). Because of that, we've built LLM products in the legal space with the following principles in mind:
- Human-in-the-loop tooling -- The product should be built around an expert using it whenever possible, so decision support as opposed to automation. You still see massive time savings with that in place
- Transparency / citations -- With a human-in-the-loop tool, you need mechanisms to build trust. Whether that's highlighting clauses in the document that the LLM referred to or explaining why a part of the analysis wasn't provided, citing your work is important
- Tuned for precision instead of recall -- False positives (and hallucinations) are especially bad in many of these legal use cases, so tuning the model or prompts for precision helps with mitigation.
When you say “tuned for precision” is this your prompt engineering or are you actually fine-tuning GPT-4?
Appreciate the insights.
We're not fine-tuning today. Instead, "tuning for precision" is done through prompt chains. A simple example would be returning "I don't know" early on if the document isn't a contract or doesn't have clear insurance requirements in it. We've had success with various guardrail prompts.
The citation work we did at Google used model internals to highlight text (path integrated gradients). It's also easier to finetune for precision when you have control over the model itself.
If you're interested in the capabilities and limitations, I suggest these informative, but still light reads as well: https://kirasystems.com/science/ https://zuva.ai/blog/ https://www.atticusprojectai.org/cuad
I kind of assumed they were in the same space as government documents.
Ah yes, the story of bad people not wanting their livelihoods taken from them by good tech giants. Seriously, is there no room for empathy in all of this ? If you went through law school and likely got yourself in debt in the process then you're not protecting any monopoly but your means to exist. There are people like that out there you know.
Are you joking?
Do you not empathize with the far, far larger number of people who can't afford adequate legal representation and have no legal recourse?
There are people like that out there you know!!!!
That's because even lawyers understand that that's an ineffective argument. Lawyers running government have allowed buggy repairmen, secretaries, and telephone operators to be automated out of jobs in the past, and now we're seeing cashiers, call center support staff, and writers have their numbers reduced due to automation.
If lawyers use "but think of the poor lawyers" reasoning to suddenly take a stand and pass laws further guaranteeing their legal monopoly, they would rightfully be called out as selfish hypocrites. I think lawyers know they're better off with the "ineffective counsel", "but think of the children", and similar types of arguments.
January 2023 my self-proclaimed smartypants lawyerbro tried to bully me into accepting that ChatGPT wasn't anything special.
I still maintain that he is incorrect.
So despite the early news about lawyers submitting 'fake' cases, it is only a matter of time before the legal profession is up-ended. There are a ton of paralegals, etc.. doing 'grunt' work in firms, that an LLM can do. These are considered white color, and will be gone.
It will progress in a similar fashion to coding.
It will be like having a junior partner that you have to double check, or that can do some boiler plate for you.
You can't trust completely, but you don't trust your junior devs do you, but it gets you 80% there.
I kind of agree with this, but this is why I am confused that I only ever see people (at least on HN) talk about AI up-ending the legal profession and putting droves of lawyers out of work--I never see the same talk about the coding industry being transformed in this way.
Maybe HN is full of coders that still think themselves 'special' and can't be replaced.
Or maybe, the law profession has a lot more boilerplate than the coding profession?
So legal profession has more that can be replaced?
Coders will be replaced, but maybe not at same rate as paralegals.
What if you had a $99/mo pro-se legal service that does two things, 1) teaches you how to move all of your assets into secure vehicles. 2) At the same time it lets you conduct your own defense pro-se-- but the point is not to win, it's just to jam the system. If you signal to the opposing party that you're legally bankrupt and then you just file motion after motion and make it as excruciating as possible for them they might just say nevermind when they realize it's gonna take them 5 years to get through appeals process.
It's true lawyers don't want to give up their legal documents for a template service-- but honestly just going to the court house and ingesting tons of filings might do the trick. With that strategy in mind you don't really need GREAT documents or legal theory anyway. Just docs that comply with court filing requirements. Yeah we're def gonna need to deposition your housekeepers daughters at $400/h and if you have a problem with that I would be happy to have a hearing about it. If enough people did this is would basically bring the legal system to a standstill and give power back to the people.
RIP Aaron Swartz who died fighting for these issues :'(
It's called a Strategic Lawsuit Against Public Participation.
It's also illegal in most states.
> A normal Tuesday in Amerika.
We already understood your derogatory outlook without it needing to be literal.
Anyways, that's why the founders gave us the First, Fourth, and Fifth amendments. It turns out, they recognized, the "law" is literally the worst mechanism for discovering truth and managing outcomes.
> Yes, the banality of tuesday terrorism.
It's terrorism because it has no logical conclusion nor any ability to positively benefit anyone's life, in particular, the person who would wield it. Your equivocation ignores this.
In all criminal prosecutions, the accused shall enjoy the right ... to have the Assistance of Counsel for his defense -- 6th amendment to US Constitution
When an LLM is more competent than an average human counsel, does this amendment still require assistance of a human counsel?People who attended law school/university/have a law degree are not necessarily practising law as lawyers/attorney or whatever they call them in the given country, as a member of the government. But of course, once a marine, always a marine...
On the other hand, the parent was trying to refer to criminal defense (Amendment 6 of US Constitution). That's not something active politicians can participate in (conflict of interest, etc.).
If one has the means to pay for their criminal defense (or family law, anything that truly matters for the average citizen), they tend to be willing to pay a lot and expect some quality of service in exchange. They will certainly not choose ChatGPT, just because that provides the answer faster for long documents.
But a large number of people don't have this kind of money. As long as governments don't have a proper budget for legal aid, they will be inclined to cover the vanishing legal aid budget with false claims that they fulfill "appointment of counsel" with LLM services - while they don't really care about the outcome of the case. (And of course, successful politicians have rarely worked as pro bono counsels before their current job, not that it has any more relevance here.)
Non-Western governments are also interested in having this replacement service - the cost of lawyers is also money there. It's much better for all governments to have chatbots they have reliably trained than less controllable humans.
That's a hypothetical threat for now, but I hope you understand why there is something more here beyond "guildthink" and protecting the livelihood of lawyers at all costs behind the Sixth Amendment. There are lots of countries (or even within the US, like in Utah) where you can review contracts without being a practising lawyer, like LPOs mentioned in the paper. So far, it failed to make the world a better place.
But replacing including LLMs as "counsel" under Sixth Amendment would mean a very different situation.
A lawyer can handle the trial for you and things like that. The LLM can help you with issues of fact, etc. And could even make stronger privacy guarantees than a lawyer if setup right. (But I doubt that will ever happen.)
This isn't worth the pdf it wasn't printed on.
They sometimes are.
If you have an issue with the methodology of paper then all well and good but "conflict of interest" is pretty weak.
Yes, Google and Microsoft et al regularly publish papers describing Sota performance they sometimes use internally and even sell. I didn't have to think much before wavenet came to mind.
Besides, the best performing models here are all Open AI.
Not that this isn't exactly what all the big "tech innovation" of the last decade were either. It's depressing and everyone involved should be ashamed of themselves.
Many people don't even try to read these because they're too long and you wouldn't necessarily understand what it means even if you did read it.
What if, before you signed something, you could have an LLM review it, summarize the key points, and flag anything unusual about it compared to similar kinds of documents? That seems better than not reading it at all.
We're not building for consumers today, because I think it's vanishingly unlikely that you'll, like, pick a different car rental company once you read their contract :) but leases, employment contracts, options agreements, SaaS agreements... all common, all boilerplate with 5-10 areas to focus on, all ready for LLMs!
You never know how consumer behavior may change when something that was either impossible or impractical becomes very easy b
If you want to share experience feel free to reach out [1] legalreview.ai
I'd love to hear that LLMs can be used to trim and simplify complexity, but I don't believe it. They generate content, not reduce it.
It's not lawyers' job to publish the laws online. Lawyers are the ones who would benefit from more easily searchable online laws, as they are the ones whose job is actually to read the laws. That is why there are various commercial tools that provide this functionality, that lawyers pay for. You need to ask your government why public online legal databases are so poor, not your lawyer.
Also helped lawyers looking for a CLM, and they rejected something if it caused any inconvenience.
But this is what happens when industries spend a decade brain-draining academia with offers researchers would be insane to refuse
In my experience current practice (unchanged for the decades I've been using lawyers) is that associates start with an existing contract that's pretty similar to what's needed and just update it as necessary.
Also in my experience, a contract of any length ends up with overlooked bugs (changed sections II(a) and IV(c) for the new terms but forgot to update IX(h)) and I doubt this would be any better with a machine-generated first draft.
1. ChatGPT can't be held responsible, it has no body, like summoning the genie from the lamp, and about as sneaky with its hard to detect errors
2. ChatGPT is not autonomous, not even a little, these models can't recover from error on their own. And their input buffers are limited in size, and don't work all that well when stretched at maximum
Especially the autonomy part is hard. Very hard in general. LLMs need to become agents to collect experience and learn, just training on human text is not good enough, it's our experience not theirs. They need to make their own mistakes to learn how to correct themselves.
We use open ai but they only get segments of a contract. Not the full one and can’t connect them.
You get the review via email and after you can delete the document and keep the review.
[1] legalreview.ai
Also how do you reconcile several logical arguments without a solver? Like "If all tweedles must tweed", "If X is a tweedle, therefore it must tweed unless it can meet conditions in para 12". How can it possibly learn to solve many such conjunctions that are staple in legal language?
https://www.lexisnexis.com/en-us/products/lexis-plus-ai.page
I haven't had a chance to test it out as anyone should be a bit weary to add more paid features to an already insanely expensive software product!
If so, the abstract and title feels misleading.
I’d be more interested in a study done on thousands of contracts of different types. I also have my doubts it would perform well on novel clauses or contracts.
This is not about drawing up a benchmark for an industry, let alone validating any useful method in general.
Ten docs reviewed based a single review 'playbook' (I think it's not in the paper, but probably max 20 questions per contract?) and compared across 3 different providers/roles + LLMs...
This seems like a weirdly arbitrary and forced statistic, when "100% reduction" would be just as valid of a statement.
But, if the stochastic analysis was ... wrong ... who would be left to correct it?
"Onit Announces Generative AI-Powered Virtual Legal Operations Assistant for In-house Counsel"
[1] [PDF] https://www.supremecourt.gov/publicinfo/year-end/2023year-en...