I tried ChatGPT 4 for a specific TX county's laws. This was after I had spent hours doing traditional search. ChatGPT got everything correct, and even gave me some more direction.
My anecdote with technical information is exactly the opposite, unfortunately.
I routinely use it to help guide my research - generally about topics I know something about but want to know more.
It often will provide some response that is entirely wrong but looks very good. When called out, it apologizes, then proceeds to produce more information that's also not entirely correct. There's an artform to teasing out the right answer... but of course you have to be knowledgeable enough to know what is right in the first place.
The subtleties of correct or wrong can be very difficult to determine for a layperson. Without deep knowledge, one might be inclined to believe the first response it provides.
That's pretty terrifying when you think of the possible implications revolving around research, law, etc.
The flow I've been going through:
* Do this task with this constraint > GPT4 Does task
* Analyse all of the items and whether they adhere to the constraint > GPT4 "Apologies, I got some things wrong, here's what I did incorrect"
* Fix the things you did wrong, remember the constraint is of upmost importance > GPT4 "I have completed the task and fixed the constraint
* Analyse all of the items and whether they adhere to the constraint > GPT4 "Apologies, I got some things wrong, here's what I did incorrect"
...
At least it can analyse the results, but even then if you don't ask in the right way it will blatantly lie to you and tell you it's all correct when it's not.
What is the difference here? No hallucinations, everything correct… is it just random chance?
There is no logic, no reasoning, no nothing. Yet people complain it's 'giving me wrong answers'. Well, it doesn't know what's it's giving you either way. It only knows the statistics or odds about the sentences it creates being similar to what others have said before.
As an aside, the claims that people are willing to make about language models are quite astounding considering that they never seem to realize that most of those claims apply to humans also...
But on programming language and other logical question areas it does not matter much, as you can verify with logic if something is correct, and then it is incredible useful
The estate was tiny, and the actual legal advice was locked behind a paywall ($7k minimum in legal fees) which would have otherwise taken the majority of the estate.
Still a huge win for $20/month.
Some months back, I tried asking it detailed questions about Australian drug laws. Initially it responded accurately, but then told me some absolute crazy nonsense (that LSD was a completely legal drug in Australia). I just tried it again and it isn't doing that any more. I still wouldn't trust anything it says about legal topics.
IBM mainframe assembly language remains a topic on which ChatGPT (the GPT-3.5 version at least) rather consistently hallucinates. For example, I just asked it to explain the difference between SVC (Supervisor Call) and PC (Program Call) instructions. It wrongly claimed PC is used to make calls within the current program. On the contrary, the PC instruction is basically an LPC/IPC mechanism, it is used to make a call to another process running in a different address space.
https://reason.com/volokh/2023/08/21/another-note-from-a-jud...
https://reason.com/volokh/2023/06/22/sanctions-issued-in-cas...
In law, being correct 99% of the time is not really good enough.
I recently vectorized a PDF of material that contains a lot of complex scheduling language. It covers legal regulations as well as contractual constraints. Using langchain’s QA function I queried the vector DB. I then tested it against a sample of questions from a Facebook group on the subject. It did shockingly well.
It seems so much has to do with feeding it the right chunks of source material. When I experimented with just copy and paste small chunks on my own it rarely provided useful feedback.
I love the law idea. I just bought a house and would like to do an addition. Feeding local code and regulation for QA could be a god send in navigating this stuff for a non-expert.
The safest approach currently is to use GPT-4 and tell it something like:
"I will give you a block of text and a question. Please answer the question using only the information in the provided text. Don't make any guesses and admit if you are not sure about your answer. If you don't know the answer, say you don't know the answer.
Question: X
Text: Y"
Trouble is that you cannot put the entire books of law into Y because the context is limited. So currently you need to do something else to sieve through the law material and find only the passages most relevant to the question through embeddings or full text search or some other method. It's not very reliable.
I don't recall which one it was (or if it might have been another one) that had a rental contract as the example document. The interesting part (to me) was the ability to ask it questions like "can you bring a dog to a party" and have it answer and show the passages that restricted pets and parties.