VASA-1: Lifelike audio-driven talking faces generated in real time
microsoft.com
microsoft.com
(via https://news.ycombinator.com/item?id=40088826, but we merged that thread hither)
Meanwhile, yesterday my credit card company asked me if I wanted to use voice authentication for verifying my identity "more securely" on the phone. Surely the company spent many millions of dollars to enable this new security-theater feature.
It begs the question: Is every single executive and manager at my credit card company completely unaware that right now anyone can clone anyone else's voice by obtaining a short sample audio clip taken from any social network? If anyone is aware, why is the company acting like this?
Corporate America is so far behind the times it's not even funny.
---
[a] With apologies to Daft Punk.
Your mistake is assuming the company cares. The "company" is a hundred different disjointed departments that only care about not getting caught Equifax-style (or filing for bankruptcy if caught). If the marketing director sees a shiny new thing that might boost some random KPI they may not really care about security.
However in the rare chance that your bank is actually half decent, I'd suggest contacting their IT/Security teams about your concerns. Maybe you'll save some folks from getting scammed?
Corporations are ultimately no better than governments and likely worse depending on what their regulatory environment looks like.
Find an exec that needs a project to advance their career. Make your software that project.
Suck in as many other execs into the project so their careers become coupled to getting your software rolled out.
If you treat it as OR instead of AND, then your security is only as good as the worst link in the chain.
"My voice is my passport. Verify me." [2]
Asking customers questions that they don't remember and that fraudsters have in front of them isn't working and the time it takes for agents to authenticate is very expensive.
While there is no doubt that companies will screw up with security, you are making wild accusations without reference to any evidence.
Maybe one day someone will successfully argue that adding easily defeated checks lowers security, by adding friction for no reason or instilling false confidence in users at both ends.
This one has too much fake looking body movement and looks eerie/robotic/uncanny valley. The lips don't sync properly in many places. Eye movement and over all head and body movement is not very natural at all.
While EMO looks just perfect mostly. The very first two videos on EMO page are perfect example of that. See the rap near the end to see how good EMO is at lip sync.
With modern secure hardware keys it may yet be possible. The difficulty is that any kind of photo/video manipulation would break the signature (and there are practical reasons to want to be able to edit videos obviously).
In the ideal world, any mutation to the original source content could be traceable to the original source content. But that's not an easy problem to solve.
In a few years time when (if) faking realistic footage becomes trivial, I suspect this kind of video will have a much, much higher level of scrutiny or only be accepted from certain sources such as government owned traffic cameras.
We still need a way to establish truth. It's important for security cameras, for politics, and for public figures. Here are some things we could start looking into.
* Cameras that sign their output. Yes, this camera caught this video, and it hasn't been modified. This is a must for recordings being used in court evidence IMO. Otherwise framing a crime is as easy as a few deep fakes and planting some DNA or fingerprints at the scene of the crime.
* People digitally signing pictures/audio/videos of them. Even if they digitally modified the data it shows that they consent to having their image associated with that message. It reduces the strength of the attack vector of deep fake videos for reputation sabotage.
* Malicious content source detection and flagging. Think email spam filter type tagging of fake content. Community notes on X would be another good example.
* Digital manipulation detection. I'm less than hopeful this will be the way in the long term, but could be used to disprove some fraud.
I’ve always had a suspicion that governments and large companies would prefer a world without hard cryptographic proofs. After wikileaks they noticed DKIM can cause them major blowback. Somehow general public isn’t aware all the emails were proven authentic with DKIM signatures and even in fairly educated circles people believe the “emails were fake” but it’s not actually possible.
You say this as if it were not a big deal, but losing a century's worth of authentication infrastructure/practises is a Bad Thing which will have large negative externalities.
People will expect or require that chain of custody only if all or at least the vast majority of the content they want would have that chain of custody.
Photo/video content will have that chain of custody only if all or almost all of devices recording that content will support it - including all the cheapest mass-produced devices in reasonably widespread use anywhere in the world.
And that chain of custody provides the benefit only if literally 100% of these manufacturers have their private keys secure 100% of the time, which is simply not happening; at least one such key will leak, if not unintentionally then intentionally for some intelligence agency who wants to fake content.
And what do you do once you see a leak of the private keys used for signing the certificates for the private keys securely embedded in (for example) all of 2029 Huawei smartphones, which could be like 200 million phones? The users won't replace their phones just because of that, and you'll have all these users making content - so everyone will have to choose to either auto-block and discard everything from all those 200 million users, or permit content with a potentially fake chain of custody; and I'm totally certain that most people will prefer the latter.
Also, for the potential creators of political fakes, such a multisig won't change things - getting a manufacturer's key may take some effort, but getting (and 'burning') keys of a dozen random people is relatively trivial in many ways - e.g. buying off of poor people, stealing from compromised random machines, or simply issuing fake identities for state-backed actors.
I can just take a (crypto-signed) photo of another photo.
You are correct that I as a viewer can't just rely on a crypto-signature like a watermark, I'd have to verify the chain of custody, but if I wanted to do that, it is available to do so.
Am I going to have to do AuthN and AuthZ on every phone call and zoom now?
Over the past century and a half, we've moved into vast, anonymous spaces, where I'm as likely to know and get along with my neighbour as I am to win the lottery.
And this is important. No, it's not just a matter of putting on an effort to learn who my neighbour is -- my neighbour is literally someone whose life experiences are wildly different, whose social outcomes will be wildly different, whose beliefs and values are wildly different, and, for all I know, goes to conferences about how to eliminate me and my kind.
(This last part is not speculation; I'm trans; see: CPAC)
And these are my reasons. My neighbour is probably equivalently terrified of me, or what I represent, or the media I consume, or the conferences that I go to.
Generalizing, you can't take a bunch of random people whose only bond is that they share meatspace-proximity, draw a circle around them, and declare them a community; those communities are _gone_, and you can no more bring them back than you can revive a corpse. (This would also probably not be a good idea, even if it were possible: they were also incredibly uncomfortable places for anyone who didn't fit in, and we have generations of fiction about people risking everything to leave for those big anonymous cities we created in step 1.)
So, here we are, dependent on technology to stay in touch with far-flung friends and lovers and family, all of us, scattered like spiderwebs across the globe, and now into the strands drips a poison.
Daniel Dennett was right. Counterfeit people are an enormous danger to civilization. Research like this should stop immediately.
Believing that everything you eat is poisoned is no way to live. Believing that everything you see is a lie is also no way to live.
Why would you expect this to happen? Lots of people are gullible, if it were otherwise a lot of well-known politicians would be out of a job or would never have been elected to begin with.
It's fascinating how research can take on a life of its own and will be pushed, by someone, to its own conclusion. Even for immensely destructive technologies (e.g., atomic weapons, viruses), the impact of a technology is its own attractor (could you say that's risk-seeking behavior?)
> Am I going to have to do AuthN and AuthZ on every phone call and zoom now?
"Alexa, I need an alibi for yesterday at noon."
Because even this precise tech has legitimate use cases?
> The only purpose of this technology I can think of is getting spies to abuse others.
Can you really not think of any other use cases?
I think it's mostly "because it can be done". These types of impressive demos have become relatively low hanging fruit in terms of how modern machine learning can be applied.
One could imagine commercial applications (VR, virtual "try before you buy", etc), but things like this can also be a flex by the AI labs, or a PhD student wanting to write a paper.
Subdermal X.509 maybe with some sort of neurolink adapter so you can confirm the request for identity. Though, first versions might be just a small button you need to press during the handshake.
Teams started rolling out Avatars https://techcommunity.microsoft.com/t5/microsoft-teams-blog/..., this would be a step up. I'm not really a fan but that doesn't mean I can excuse the use case.
- Advertising. Where I am, there is a pervasive commercial with (very realistic) talking goats. Why not do the same thing with people? No need for actors, when you can just tell the computer what you want.
- Cinema and television. Especially for bit parts and extras, just create the characters, instead of rounding up a bunch of extras.
- Video games and alternate realities. They've been getting more and more realistic - this is just the next step.
- Pornography. Again, why trouble yourself with real actors? Sell premium videos customized to each customer.
- Politics. Give "live" speeches in different venues, without all the bother of travelling. It's a very small step to answering questions live - just train up an LLM with the responses you want it to give.
Lots more applications as well - those are just a few that come to mind.
Is there something equivalent but MIT or Apache?
I feel like diffusion transformers are key now.
I wonder if OpenAI implemented their SORA stuff from scratch or if they built on the Facebook Research diffusion transformers library. That would be interesting if they violated the non-commercial part.
Hm. Found one: https://github.com/milmor/diffusion-transformer-keras
I thought deepfakes were still quite a bit away but after this I will have to be way more careful online. It's not far from behind something that can show up in your "YouTube shorts" feed and trick you if you didn't already know it was AI.
This one has too much movement and looks eerie/robotic/uncanny valley. While EMO looks just perfect.
Both are, BTW, AMAZING!! Pretty crazy.
But that's nitpicking. It's good enough to fool someone not watching too closely. And the fact that the result is this good with a single photo is truly astonishing, we used to have to train models on thousands of photos for days only to end up with a worse result!
Technically videos could've been faked before but it would require a ton of effort and skill that no average person would have.
Courts already wouldn't generally approve random footage without clear provenance.
There’s likely also a an unsaid statement. This is for us only and we’ll be the only ones making money from it with our definition of “safety” and “positive”.
With AI stuff, I have learned to be very skeptical until and unless a relatively publicly accessible demo with user specified inputs is available.
It is way too easy for humans to cherry pick the nice outputs, or to take advantage of biases in the training data to generate nice outputs, and is not at all reflective of how it holds up in the real world.
Part of the reason why ChatGPT, Stable Diffusion, Dall-E had such an impact is the people could try and see for themselves without being told how awesome it was by the people making it.
Of course, I'm sure that whoever put these demos together also invested time in getting them right. Still: so would anyone seeking to put words in another person's mouth.
Imagine your favorite (or least favorite) politician coming out with a speech promoting something awful. Even if the video were immediately debunked, people would remember it. And many people would never believe the debunking.
This is the world we now live in: you literally cannot trust anything you see online.
Still, apart from the teeth this looks extremely convincing!
Weird implications for various regulations though.
Tech like this has the potential to bring us back to the days of "on the Internet, nobody know's you're a dog" https://en.wikipedia.org/wiki/On_the_Internet,_nobody_knows_...
Every second day HN has some post about some new amazing AI system. Never available to download run and use.
Why the vast investment and no startup selling consumer downloadable software to do it?
And replacing a person which spreads lies, as can be seen in most TV or glossy cover ads, shouldn't trigger some new legal action. The only difference is that now the actor is also a lie.
And countries which use actors or news anchors for spreading propaganda surely won't see an issue with replacing them with AI characters.
People who then get to read that their most favorite, stunningly beautiful Instagram or TikTok influencer is nothing but a fat, chips-eating ugly person using AI, may try to raise some legal issues to soothe their disappointment. They then might raise a point which sounds reasonable, but which would then force politicians to also tackle the lies which are spread in TV/Magazines ads.
Maybe clearly labeling any use of this tech, maybe even with a QR code linking to who is the owner of the AI, similar to QR codes on meat packaging which allow you to track the origin of the meat, would be something what laws could be helpful with, in the spirit of transparency.
Today a big ML model can do this and it's somewhat regulate-able, tomorrow people can do this on their contact-lens supercomputers and anyone can generate a video of anything.
Is going back to personally knowing your local representative the only way? How will we vote for national candidates if nobody knows what they think or say?
I've got my popcorn ready.
But you can rest easy. Everyone just votes for the candidate their party picked, anyway.
I wonder if it's possible to digitally sign footage as it's captured? It'd be nice to have some share-able demonstrably true media.
Edit: I'm a centrist and I definitely would lean one way or the other based on who the options are (or who I think they are).
Not that big:
Just as there's no privacy on the internet, how about 'theres very little trust on the internet'. Assume everything not securely signed by a trusted party is false.
Hell I've been guilty of spouting BS before, just because I've heard something from so many people. Then find that when I look it up, it's not true.
It's not really a tech problem, it's more of a human problem imo, like so many others. But there is literally nothing we can do about it.
i’m going to burst your bubble here, but most voters have no idea about policies or candidates. most voters vote based on inertia or minimal cues, not on policies or candidates.
i suggest you look up “The American Voter”, “The Democratic Dilemma: Can Citizens Learn What They Need to Know?” and “American National Election Studies”.
Potentially we'll need slightly tighter regulations on formal press (so that people that care for accurate information have a place they can get it) and definitely we'll want to steer the culture back towards holding them accountable for misinformation, but credulous people have always had easy access to bad information.
I'm much more worried at the potential abuse cases that involve ordinary people that aren't public figures, and have much less ability to defend themselves. Heck, even celebrities are a more vulnerable targets than politicians.
Knowing that retractions rarely get viral exposure, it's not difficult to imagine that a few sufficiently-viral videos could swing enough votes to impact a presidential election. Especially when considering that the average person is not up to speed on the current state of the tech, and so has not been prompted to build up the mindset that's required to fend off this new threat.
[0] https://www.washingtonpost.com/news/the-fix/wp/2016/12/01/do...
The civilization might be fine, sure. Now, democracy, on the other hand...
If a business is showing a demo of this you can be assured that the Government already has this tech and has for a period of time.
> How will we vote for national candidates if nobody knows what they think or say?
You don't know what they think or say now - hopefully this disabuses people of this notion.
That may have been true once upon a time, but it no longer is. And even in the areas it was true it was mostly for niche areas like cryptanalysis.
Governments simply cannot attract or keep the level of talent required to have been far ahead of industry on LLMs and similar tech, especially not with the huge difference in salaries and working conditions.
Democracy is already an illusion of choice anyway; just look at democratic candidates. It's gonna be Biden V Trump _again_. For London mayoral elections Sadiq is pretty much guaranteed to get in _again_. For UK main election it's gonna be the typical Tories V Labour BS _again_, with no new fresh young candidates with new ideas.
Democracy is rotting everywhere it exists thanks to the idea of parties, party politics and the human need to pick a tribe and attack every other tribe.
That's basically "never" then, so we'll see how long they hold out.
Scammers are already using the existing voice/image/video generation apparently fairly successfully. :(
But knowing that this is possible is important to know.
I'm fairly clued in, and am constantly surprised at how fast things are changing.
Who knowing this is possible?
The general elderly person isn't going to know any time soon. The SV IT people probably will.
It's not an even distribution of knowledge. ;/
Like somebody on Ars noted "anybody notice it's an election year?" You don't need to release an API, all online videos are now suspicious authenticity. Somebody make a video of Trump or Biden's eyes following the mouse cursor around. Real videos turned into fake videos.
Can someone explain the commercial need to take someones likeness and generate video content?
If I was an a-list celebrity, I would give permission for coke to make a commercial with my likeness, provided I am allowed final approval of the finished ad?
Do I have an avatar that attends my zoom work calls?
Personal disinformation and propaganda campaigns.
Oh Brave New World, that has such fake people in it!
(but actually, because laziness is the driver of all innovation, I wouldn't be surprised if this happens).
This would be as if we invented and sold nuclear weapons to dig out quarry mines faster. The inconvenience it saves us quickly disappears into the overwhelming shadow of the enormous harm now enabled.
”Project Plowshare was the overall United States program for the development of techniques to use nuclear explosives for peaceful construction purposes.”[0]
Disney has been using digital likeness to maintain characters who's actors/actresses have died. Princess Leia is the most prominent example. Arguably, there is significant realistic value in being able to generate a human-like character that doesn't have to be recast. That character can be any age, at any time, and look exactly like the actor/actress.
For actors/actresses, I suspect many of them will start licensing their image/likeness as they look to wind down their careers. It gives them on-going income with very little effort.
"Such technology holds the promise of enriching digital communication, increasing accessibility for those with communicative impairments, transforming education methods with interactive AI tutoring, and providing therapeutic support and social interaction in healthcare."
A convincing simulacrum of empathy could plausibly be the most profitable product since oil.
Even though most of those things are illegal you could just have foreign cat's paw firms do it. Maybe you fire them for "going to far" after the damage is done, assuming some even manages to connect the dots.
If they can do it, so can someone else and hiding it makes it worse. If it's widely available people will quickly realise that the talking head on YT spouting racist BS is AI. This process needs to happen faster.
Ofc there will still be people who don't care or understand, but there will always be people who are for example racist and don't care if the affirmation for their beliefs comes from a human or a machine.
real jurassic park "too preoccupied with whether they could" vibes
This should end well.