AI-Generated Voice Evidence Poses Dangers in Court
lawfaremedia.org
lawfaremedia.org
Generally speaking, I think evidence tampering is not a new problem, and even though it's easy in some cases, I don't think it's _that_ widespread. Just like it's possible to lie on the stand, but people usually think twice before they do it, because _if_ they are found to have lied, they're in trouble.
My main concern is rather that legit evidence can now easily be called into question. That seems to me like a much higher risk than fake evidence, considering the overall dynamics.
But ultimately: Humanity has coped without photo, audio or video evidence for most of its existence. I suppose it will cope again.
AI images of real life still don’t pass intense scrutiny.
Evidence of provenance is already important, to be sure, but the the ability to have some degree of validation of the contents has itself provided some evidence of provenance; lose that and there is a real challenge.
But then you have an inside job where the perpetrators work for the store and have doctored the footage before the police come to pick it up, or a corrupt cop who wants to convict someone without proving their case or is accepting bribes to convict the wrong person and now has easy access to forgeries. Chain of custody can't help you in either of these cases, and both of those things definitely happen in real life, so how do you determine when they happen or don't?
The store manager is in the chain of custody but isn't a suspect, the accused is the kid. The kid doesn't even know who actually committed the crime. How is the kid supposed to prove this?
There are some experimental specifications that exist to provide attestation as to the authenticity of media. But most of what I’ve seen so far is a “perjury based” approach that just requires a human to say that something is authentic.
There are two major problems with this.
First, is all footage from existing surveillance systems going to be thrown out because it doesn't use this technology? Answer: No, because it would be impractical. But then nobody cares to adopt the technology because using it isn't required. How's that IPv6 transition going?
Second, that sort of thing doesn't actually work anyway. Surveillance cameras are made by the lowest bidder. Their security record is appalling. They're going to publish their private keys on github and expose buffer overflows to the public internet and leave a telnet server running on the camera that gives you a root shell with no password. Does it sound like hyperbole? Those are all things that have actually happened.
There is only one known way to prevent this from happening: Do not allow the hardware vendor to write the software. Any of the software. Instead, demand hardware documentation so that the firmware can be written by open source software people instead of lowest bidder hardware companies. This is incompatible with using the hardware vendor as the root of trust, which is a natural consequence because the hardware vendors are completely untrustworthy.
But let's suppose we find some way to do it. We'll pass a law imposing a $100 fine on any company that has a security vulnerability. Then there will never be a security vulnerability again because security vulnerabilities will be illegal; I'm assured this is how laws work. At that point the forger takes the camera and points it a a high resolution playback of the forgery, and the camera records and signs the forgery.
I kind of wish people would stop suggesting this. It's completely useless but it creates the false impression that it can be solved this way and then people stop trying to find a real solution.
Imagine if you will, that the NVR (recording system) has a unique private key flashed in during manufacturing, with the corresponding public key printed on it's nameplate. The device can sign a video clip and its related meta-data before exporting. Now, any decent hacker could see possible holes in this system, but it could be made tamper-resistant enough that any non-expert wouldn't be able to fabricate a signed video. Then the evidence becomes the signed video and the NVR's serial number and public key. Not perfect, but probably good enough.
As long as chain of custody ca be discarded because 'good faith' whenever it becomes inconvenient it is not a real thing.
We've been there for at least two years.
https://arstechnica.com/tech-policy/2023/04/judge-slams-tesl...
In any case, you can probably do this with a fairly simple prompt with any LLM.
When it comes to generative AI, I personally don't see a lot of good applications, but a plethora of bad ones. The only solution I could imagine would be regulation to the degree that using or distributing models with certain capabilities is just illegal. Judging from the war on file sharing a few decades ago, probably very difficult to enforce, even if it is perhaps still worth doing.
But I don't see any governments line up to do it. Given that, this particular (semi) new development that generative AI is effective for evidence tampering, I think we'll manage to deal with it.
Most problems can be overlooked because they’re not that prolific. There’s a chain of events that really hinges on that assumption.
For example, piracy is really not a problem. In fact I’m pro-piracy most of the time. But if everyone could pirate with no thought and no down time, I don’t think our economy will survive. Just because something is okay in small doses doesn’t mean it scales well. We understand this intuitively with substances, but not with technology like LLMs.
In fact, I remember a drug dealer I was helping with his defense. Hidden mic on undercover was taped under his armpit. He arrived to defendant's hotel room for a deal and defendant made him undress. The recording is hilarious because defendant is like "You fucking snitch, what's that under your armpit?" and the undercover says "It's my .. er .. MP3 player?" LOL
Same argument for electricity, the internet, sanitation, democracy... doesn't seem like a great test for stuff that we just didn't have it before and survived.
If the cops spend the first month going after some poor Mark that had nothing to do with it, by the time they realize they've been had it will likely be much more difficult to catch the actual perp.
Also, that's saying nothing about the court of public opinion.
which allowed to burn witches based on testimony of citizens in good standing.
And that leads us to using Neuralink and similar tech and the next gen lie detectors (like say the defendant's fMRI (most probably interpreted by AI too:)) to look into the brain and extract confession. No need for evidence, deposition and all that expensive time consuming stuff standing in the way of truth especially given that it can't be trusted anymore in the face of AI capabilities.
I don't have data, but I suspect a LOT of people lie on the stand. This is mostly based on what I see on reality court shows and true crime type shows, so admittedly not a great sample, but I figure once something gets to trial things are going to be contentious.
Why would the judge be better qualified to determine whether the voice was authentic, as opposed to the witness? And why should the judge effectively determine the witness's credibility or ability to discern, when that's what juries are for?
All that said, emulated voices do pose big problems for litigation.
This also reminds me of one of Norm Macdonald's bits. He says (very seriously) that if he were on a jury he would not convict someone on the basis of DNA evidence. "What do I know about DNA?" he says.
- perhaps the judge make less experts interviews than the corporations, leading to less experience in that but also takes each of them more seriously.
- one way to remove/add some credit to someone claim is to ask some of their peers opinion and see if there’s a strong majority.
- the judge personal expertise may help him forge a precise opinion but that wouldn’t clear him of making a mistake. For important matters it’s always a good idea to ask for peers reviews. Academics knows that too.
While certainly Fox News headlines would not reach the jury in most instances, that is on account of hearsay, lack of qualification, materiality, relevance, and similar rules. It is not a prior credibility or weight determination by the judge, as I understand TFA to be advocating. So: did the witness hear a voice that he believed to be the one in question? If so, jury gets to decide (unless unfairly prejudicial or some other overriding rule comes into play).
Nor should obviously fake evidence reach the jury. They can judge for themselves whether testimony is credible, but this is far different than admitting faked evidence. And if you can't see the difference, I'm not sure I'm qualified to explain it to you.
It's pretty simple. If the "evidence" can be created by software, it's not evidence. Perhaps some specially designed recording hardware might digitally sign a voice recording, and one might be inclined to trust that it was a real recording if the design of that hardware/software system was vetted.
But just for a recording? There's no point. Allowing a jury to decide that this recording is bad, but this other one is okay when they have no expertise to be able to determine if it is fake or not and none of the technical details either is just asking for prejudicial and even superstitious deliberation.
We are at the point where audio (and probably video) recordings no longer count as evidence of real world events. It's been this way for photographs for a long time already.
Or gun matching (ballistics) that are no longer considered conclusive but subjective? https://www.mdcourts.gov/data/opinions/coa/2023/10a22.pdf
Hair and Fiber expert analysis that was wrong? https://innocenceproject.org/fbi-agents-gave-erroneous-testi...
Or do you mean bite mark analysis that was again wrong?
Many of these forensic methods were used for decades and presented/treated as conclusive evidence before being challenged leading to wrongful convictions. But yes, let's have 'certified' voice experts whose living is based on being hired by the government and giving testimony to convict people. Surely this time it will be altruistic and scientific.
It is just as easy to fake many paper documents, and we have accepted documents as evidence for centuries.
Photos can be faked, video can be edited or faked, witnesses lie or misremember.
Is this just about telling lawyers that unvetted audio recordings can be unreliable? Because that shouldn't be news.
Edit: this is a good faith question. I'm legitimately just curious. Splicing and editing have been around since recording was invented, I was legitimately curious why voice recordings would have been given extra evidential weight when manipulating recordings is a known possibility.
If it takes an FX house to generate a plausible recording of me saying something I didn't say, that's a risky enterprise with a lot of witnesses.
If my enemy can do it in their basement with an hour of research, the exposure risk goes way down, and consequently the expectation you'll see it in real life goes way up.
This seems more like people losing their minds over 3d printed guns, when hobbyists with a drill press have been making guns in the garage for decades.
Yeah, its easier now to fake voice, but its not as if what this article warns against wasn't possible before the latest AI hype cycle. And it is also worth noting that voice cloning/changing technology is not particularly new either (I've been able to sound like Morgan Freeman using a phone app for at least half a decade).
I agree that courts should be cautious around accepting voice recording evidence, I just don't think that the ability to do this is new.
Sure, and https://en.wikipedia.org/wiki/David_Hahn managed to get quite a bit of nuclear material.
Being able to do it at scale, convincingly, in real-time, for any arbitrary text, with just 30s or so of someone's voice as a sample, changes the calculus a lot.
Cutting up what recordings, exactly? If you're mixing-and-matching, you need a pretty broad corpus of source material with fairly consistent recording quality (background noise, etc), and even still you're limited to reproducing words that are already present. You can't cut-and-splice together a recording of me saying 'Yes, dghlsakjg, I embezzled $1 million and blew it on the ponies' because there will be no recording of me saying your username, 'embezzled', or 'ponies.'
The problem comes with the voice-cloning technology that can construct entirely new sentences based on relatively short voice profile samples.
> I've been able to sound like Morgan Freeman using a phone app for at least half a decade
You've been able to sound like Morgan Freeman because of specific, hard tuning work put into the voice changer. Now, you can sound like your boss, or your neighbour, or your ex.
I edited the audio of a few math videos to upload to YouTube and I had to fix "typos" like "then one plus zero is zero". The problem is that there are no spaces in the voice, so it sounds like "thenonepluszeroiszero."
So I had to look for a "one" that is in a similar context, after an "s" and before a ".", or at least a similar one. And hope the speed matches, and adjust the volume.
They were free videos, so it was not necesary to get it perfect, but at least reduce the distraction caused by mistakes.
My broader point is that audio evidence has never been perfect, just like every other medium that evidence can exist in.
Generative audio technology is not the first time that audio has been corruptible as evidence, and any lawyer who believed otherwise was naive at best.
Then the cost of long distance started to drop to zero and the commercial calls started coming in. After it was zero, the outright illegal scams showed up. Now, probably one out of every fifty calls to me is an actual legitimate call. My phone goes to automatic voicemail if you're not already in my contacts.
Were talking about both quantitative changes and qualitative changes. Things like this can have large impacts on society.
There's ample evidence that this is already happening, eg. recent headlines about kids being radicalized at increasingly younger ages, groups like No Lives Matter that embrace violent nihilism, increased domestic terrorism, record high gun ownership across both sides of the political spectrum, authority figures that just do whatever they want and ignore any form of law or accountability, etc.
We're already at a place where most people don't care what other people think of them.
Issue is, as long as the government has the big guns, what the government thinks of you will still matter in a major way.
In such an environment, most people are going to choose to have some kind of way to prove to the government what they did and what they didn't do. Not because they care what other people think as you're implying, but rather because they very much care that the government not get the wrong idea about them. Because the government getting the wrong idea about you can be fatal.
I suspect that we're saying the same thing but with reversed causality. Both of us agree that non-deterministic enforcement breaks down the incentives needed for pro-social behavior. You're saying that this will cause people to demand ways to improve the governments enforcement abilities. I'm saying that this will cause people to adapt their behavior to the new, lessened enforcement abilities. In defense of my point, I'd point out that changes to government are a coordination problem while adaptation of behavior is an individual-only response, and it is much easier to effect changes to your own behavior than it is to convince 300 million people to agree on a solution and implement it, particularly when the root problem is a lack of enforcement ability.
The government coming up with one way to track everyone is not really necessary to get people to submit to tracking. In fact, most of the law enforcement "watcher" types wouldn't want that anyway. Not only is having a myriad number of ways to track and surveil people is far preferable to law enforcement, but it also allows people to individually choose to set up all the Ring doorbells and security cameras and GPS trackers and smart glasses based body cams etc etc all on their own.
And they'll happily choose to buy all that stuff voluntarily and without regard to what everyone else is doing.
Imagine that the populace supports widespread surveillance techniques, and so cameras are setup everywhere. Some hacker group figures out how to hack into the cameras and insert deepfakes in them. Now members of that hacker group have a government-proof alibi whenever they want it, and can commit crimes at will, and get it blamed on others. Justice goes out the window.
This is how the "watcher" types are trained. Information is only valid if they can get it from multiple independent sources. So they love when new tracking and surveillance channels are released. (Social media apps, or smart glasses, or doorbell cameras or what have you.)
The bar would be much higher than compromising a single channel. With multiplying channels, the task of compromising them all approaches impossibility.
Yea. Russia is going to collapse any day now........... um, I'm still waiting.
1. A projector displays a challenge pattern (Perlin noise derived from of a hash) 2. A camera captures this projection 3. The system hashes the captured image concatenated with the previous hash and uses it to derive the next projection 4. This chain demonstrates true temporal sequentiality that's difficult to forge
By incorporating random noise derived from Byzantine Fault Tolerant networks and using these networks as timestamping servers, the proofs inherit the network's decentralization properties. ML then confirms that the feature distributions in projection-photograph pairs match expected patterns from the training dataset.
Demo video and GitHub repo available here: https://www.reddit.com/r/PoliePals/comments/1j8qm2j/truth_be...
As they say, once you have a signature, you have a most of a cryptosystem. I've been experimenting with those and other applications of non-linear functions in projector-camera systems.
* Proof that this starts after a given time. Traditionally this has used methods like "this is the headline of a major newspaper today", which is limited to 1-day granularity and has problems if you can just generate a large number of expected headlines and use them in parallel. But with crypto, we can just query any random-number-timestamp-signing server, and a network of such servers can mutually sign each other's previous packets so it's very reliable both against downtime and against attacks.
* Proof of sequencing. This is trivial with a chain of hashes, though it does prevent recompression.
* Proof that this ends before a given time. This requires actively submitting your signature data to a timestamp server for additional signing, which is a much more complicated task than the initial half. It is still possible to eliminate the single source of vulnerability, but much more work.
"Camera looks at monitor" is going to be a much cheaper way to make this air-gapped than adding a projector. And this doesn't strictly need to be continuous; most things are tolerant of one-day granularity and almost everything of 15-minute granularity.
"Camera looking at a monitor": While that might be simpler in some setups, it doesn’t really solve my main issue. I want the signal to permeate the entire scene, not just appear in the corner of a display or overlaid on the video. By projecting the challenge onto all visible surfaces, we create a physical environment that’s difficult to fake (since you’d have to convincingly generate or remove those patterns in real time). Air-gapping isn’t really the goal right now.
Finally, we're need much finer granularity than 15 minutes! The point is to lower the generation time below what is achievable via generative model.
Thank you for the comment, and I hope these clarifications are useful. It's a new concept, so please forgive the clumsiness with which I may be communicating it.
"Overlay the entire scene" doesn't actually appear to add any information-theoretic value compared to simply bounding the timestamp at which the video was made. Nothing either of us is talking about will actually prevent fakes (before the camera signs it), only constrain the time at which the fakes are made.
Slower-than-real-time generation of fakes is still significantly inhibited by the fact that the hash sequence can be checked for continuity across long lengths of time whose bounds are verified using the other steps.
By projecting across the entire scene, we significantly increase the complexity of any real-time forgery attempt. A single display that barely interacts with the surroundings is comparatively easy to spoof and not particularly convenient to deploy. Of course, if you think that simpler setup is worth exploring, go ahead and experiment!
Keep in mind, this method differs from merely signing an image. It creates an optical puzzle that reality solves faster than a simulation can. The physical interaction with the environment adds complexity that’s difficult to replicate artificially. I tried placing a display in front of the camera about a decade ago, but found that approach too “analytically solvable,” so it lost its appeal for robust authentication.
1. Cryptographically hash each piece of media when it's recorded.
2. Submit the hash to a "trusted" authority.
3. It will add a timestamp and sign the result.
4. Now, as long as you keep the original, without re-compressing, and you trust the authority, you have some evidence that the media existed at a timestamp. On or before.
This doesn't prove authenticity, but in many cases, establishing a timestamp would be enough. Forgeries probably wouldn't be created until later, after the shit hit the fan.
Or maybe this doesn't work at all.
Also, a lot of surveillance systems are purposely kept offline to prevent them from being compromised, but your system doesn't allow that because they would need external connectivity to get signatures.
The issue here is that the skillset required to produce the forgery has fallen to the level that anybody can do it and then that's enough people that somebody actually does.
Every x seconds, collect all the pieces of data that need to be notarized as existing in the world prior to some time t. It can be an arbitrary amount of data.
Construct a Merkle Tree[1] of all the hashes (or HMACs) of all the data to be notarized. Compute the Merkle Root. Make sure that everyone gets their Merkle Proof (path from leaf to root) or publish the Merkle Tree publicly.
Embed the Merkle Root into a one or more cryptocurrency blockchains to exploit their immutability guarantees. By either including it as "additional data" to some transaction or just straight up as a fictional cryptocurrency address.
Every piece of data processed in this way will have a Merkle Proof (cryptographic path from hash of the data to the Merkle Root) that proves it existed prior to the creation of the Merkle Root. The Merkle Root will have its creation time bounded by the proof of work conducted on the cryptocurrency blockchain.
I don't mean "date modified" in the file metadata when it's uploaded (of course that could easily be spoofed before upload) but the actual "date uploaded". You know something must have been made before a certain time.
The tools aren't perfect yet, so it's not too late to stop. Stop the ridiculous image and audio generation tools before it's too late. Nothing of value is lost when these models are made private again, and research is simply halted.
Personally, I'd rather we all know this tech is out there and develop defense mechanisms rather than thinking hiding it away will prevent harm.
The cyber security industry exists because of all the privacy and security issues posed by all the tech we already have had for the past several decades.
I'm confident the same will happen for AI simply because it is a business opportunity and other businesses and institutions are already talking about these issues.
has always worked with guns so it'll always work with anything else :) and guns will kill more people that AI-generated shit for foreseeable future
Human communication tools are responsible for, say, the holocaust. At its core, you can distill the holocaust to that. Or choose any tragedy, really.
I think what we’re dealing with is much worse than we are giving it credit. I’m envisioning the death of trust as a concept itself. Such has never before been faced by humanity, and I don’t know if even our biological social circuits can handle it.
We’re already half way there. It’s trivial to convince large amount of people of something that isn’t true, and then get them to act on that belief. It just requires some time and some money. Remove the time and the money requirement and I don’t know where we land.
I am tired of reading so much stuff which could well be spoken to me. Like this comment here, which you now need to read.
It's only very recently that living beings on this world have learned to read and write, it's not normal. It's normal to communicate through sound.
It's normal for humans to communicate via sight and sound combined.
That being said, we invest a lot of calories and brain power into vision. It's our primary mode of interacting with the planet. This is what we have evolved to rely on.
I think, therefore, you're wrong. We rely heavily on eyesight, naturally. It's why our eyes are focused on the front, allowing precise binocular vision, whereas our ears are on the side, allowing for broader coverage, but less specific information received. Reading and writing (or at least some form of image communication) is probably more natural for communicating ideas than talking, for humans.
Even if you are still trying to average us out across other life on earth (which doesn't really make a bit of sense), I think vision is the way to go. Colors and visual attractants and other visual forms of communication are common even in plants, as well as animals.
To list a few uses I see for voice-generation/cloning tools:
* Real-time translation of a user's voice, maintaining emotion and intonation
* Professional-quality audio from cheap microphone setups (for video tutorials, indie games, etc.)
* Allowing those with speech impairments to communicate using their natural voice again, or:
* Allowing those uncomfortable with their natural voice to communicate closer to how they wish to be perceived
* Customization of voice assistants, such as to use a native accent/dialect
* Movies, podcasts, audiobooks, news broadcasts, etc. made available in a huge range of languages
* And of course: memes, satire, and parody
If it exists but isn't widely accessible, it's likely in the hands of Musk/Zuck and and various state actors. To me that seems possibly the worst alternative - to have the public generally unaware that it's possible and receiving few of the benefits listed, yet still having it available as a tool for competent disinformation.
This is some serious level of cope for HN.
Simply put, it's not going to happen.
Of course in a darkroom someone skilled could always make a fake photo - but the bar is a lot lower with AI.
Or they just point a camera at a high resolution playback of a forged video.
This also assumes you can trust all the camera makers, because by doing this you're implicitly giving them each authorization to produce forged videos. Recall that many of these companies are based in China.
It seems like it would just be a mechanism to lend undeserved credibility to something that still isn't trustworthy.
Exactly, it's just like DRM or cheating in video games. Even if everything in software is blocked, there's always the analogue hole.
Rather than cryptography which could be difficult to grok for a non-technical jury, a physical slide of film would be the source of truth.
It can still be faked by photographing a manufactured/generated scene though.
BTW, I've always wondered about the validity of email evidence. Email is text on a computer. You can edit it to say whatever you want.
Note: pretty sure I got the concept right and misused a few terms here.