Create AI videos by simply typing in text
synthesia.io
synthesia.io
It's a shame they chose that name, since it was such a great play on words for the midi software (synesthesia is sound into colorful visuals, and midi uses synths) whereas this product has basically no relation.
Sure it does. It synthesizes video from text.
Is there anything other than the “synthes” that synthesis shares with synthesia , connecting this product to synthesia in particular?
Thought the piano software was named after that word. Didn’t check the spelling.
Edit: It's worth pointing out that while the words look quite alike, they don't mean the same thing,
synthesis: to put together
aesthesia: ability to feel or perceive sensations
Ahh.. the anchor fm problem.. guess I'll need an open source version.
I started toying with libreBot I think it's called - which allows you to do anything you want with these things if you self-host license for a grand I think it was.
This synthesia didn't even get the first sentence I tried. It also requires a 'business email' and agree to terms that includes "I agree to receive occasional product information as per Synthesia Privacy Policy *"
trying hard to keep the genie in the bottle aren't they.
I may end up hiring some merc help to get it self hosted and running with some customized avatars.. the custom digital avatar stuff can be really complex if you want it to be - which I like - but I know I would spend weeks playing with all that, and someone that knows the digital stage / lighting tools out there can likely slap together what I would need for a small launch project in a day.
(On a side note, I'm not sure I understand the appeal of emotionally bland fake-smile talking heads in general, even when they're real.)
Would it not just have the same issue the existing product has for non-deaf people?
Use cases are education, museums, games, sweet sweet jesus
No I'm not asking if you think you can you use this to make money, I'm asking do you personally want to sit through a video of a robot telling you do things? Are we supposed to believe this is preferable to simply reading this or hearing recorded audio? This is flat out consumer hostility, basically telling your customers to talk to a sock puppet instead of a real person, I hope this fails, I would pay money to make this illegal.
Another one outside marketing is making educational lecture videos is a lot of work even with speech pre-written. Often execution goes through several passes. If we continue in the direction of moocs/online education than making it easier for teachers to make videos is valuable.
Multi language support they have is also big use case.
Honestly, I'm not impressed.
Animations are pretty good. Pronunciation could use some work. There also does not seem to be a way to influence the inflection, which is an absolutely crucial component for sales pitches. It's not so much what you say, but how you say it. Also, the right people have to sell the right things. Words coming from Elon's mouth in regards to cryptocurrency have a far greater effect on market behavior than the exact same words coming from this AI person's mouth.
The incoherent facial expressions actually manage to confuse the message more than the dissociated pronunciation.... "witch is know small feet".
This tech is a neat trick at this stage but is less useful than just leaving the text as text, in fact adding negative value to an already fully functional process.
Fiver is a better option, and I would not recommend that.
For an interesting and highly unethical experiment, someone should raise a thousand infants with this drivel and see what happens...I’m going to posit that the result is not good. Children’s narrations is exactly where this is headed though, I can see this as a multimillion view no effort YouTube babysitter.
Children find a pleasant, smiling female face soothing...so this is going to be another way that the dollar and human laziness will use AI to make the world a slightly worse place.
But yes, if "improving conversion rates" is your main priority, it may be helpful.
Looks like it's still a thing as well: https://www.nawmal.com/
But what’s the point?
If you’re gonna send someone a soulless corporate drone video, is that really better than a soulless corporate email? I thought the goal of doing video was that it’s more personable and human ... an AI video doesn’t quite hit those goals does it?
Can be made even more personable and human and can be customized not just for broad appeal, it can be individually customized for the appeal to the target person based say on the target's profile, browsing history, etc. Similar to how Cambridge Analytics did for text based messages.
The lips, eyes, and facial features move in natural ways, but the head remains frozen in a somewhat unnatural manner. It's just inside the uncanny valley, with barely perceptible creepiness.
I would hope to see improvements to make face/neck movements look more natural, to overcome these issues over time!
I mean, combine it with GPT-3 and you've got something that's nearly science fiction. Really interested to see where this goes.
https://youtube.com/watch?v=DFM5zbekZ7c hour-long dev talk (GDC)
https://share.synthesia.io/2761933d-4ec7-48c7-b67e-85fc9d686...
Reminds me of the movie The Congress.
Obviously this technology has a long way to go, but it seems that that actors should feel less secure about their jobs being resistant to automation.
As more things become intellectual property, the tendency of property to pool becomes severe.
Actors get a % rev share + upfront fee to work with us :)
Yay! At last! And when we've automated away everyone's work, also say goodbye to synthesia and every other automation service, because there's no business left to use it. Woo-hoo, future world, here I come!
What a great sample.
Again, super creepy and not really clear if it would drive engagement.
Anybody can stand blankly in front of a camera without emotion. But this is an impressive start.
Does anyone know what's under the hood for the text to speech?
To answer a few recurring questions in the thread
---> Use case.
Video is a way more effective way to communicate than text. Not for the HN crowd, but if you're a blue collar worker a 2 minute video in your native language is much preferred to a 5 page pdf for training.
Anyone who has tried to record a simple corporate video know the pain of cameras, film crews, 25 takes to get one that works and post production. Cumbersome, slow and multidisciplinary. By the time the video is done the content is out of date.
Synthetic video is not yet at the quality of real video. Eventually it will be. But the mistake many are making here is comparing it to real video; it should be compared with text.
In X years we'll be able to make Hollywood films on a laptop without needing anything but time and imagination. Just like we can digitally compose music in Ableton, create images in Photoshop and type novels on keyboards rather than with pen and paper.
My (obviously biased;)) belief is that synthetic media will eventually become foundational technology that will move media production from cameras/microphones to API's. We'll be able to do all kind of things we couldn't do before.
Eg. personalized and interactive rich media, video-driven chatbots and eventually Hollywood blockbusters made by your favourite YouTuber from his or her bedroom.
---> Uncanny valley
Simulating real video is incredibly hard. We're constantly improving and launching more expressive synthesis soon.
From our tests with some of our largest clients 8/10 people don't realise it's a synthetic video (unless they are asked to look for it).
---> Tech
Has been developed over the last 3 yrs. Origins/team from Stanford/UCL/TUM.
Learning: Going from research to working, scaleable product is hard and takes time. But very rewarding when it works.
[1] https://www.youtube.com/watch?v=ohmajJTcpNk [2] https://www.youtube.com/watch?v=qc5P2bvfl44
---> Bad uses
Bad actors will do bad things with synthetic media. Like with any other technology from smartphones to cars. We're moderating all content and building safeguards and verification + working with FAANG and others on detection and provenance technology.
Recommended read - deepfakes perfectly follow the story arc of any new, powerful technology: https://journals.sagepub.com/doi/full/10.1177/17456916209193...
---> Actors
Real actors getting rev share + upfront free from every video generated with their likeness. Like being a stock photo actor.
It seems to me that this technology could have immediate application to dubbing over curse words in movies (since that's already done in a not so subtle way today).
The next step I see in that progression is full dubbing for translation, which already exists in a very conspicuous form. The old meme about out of sync karate movie dubs comes in mind.
How close do you think this technology is to use for syncing lips in Hollywood tier movie dubs using real voice actors? What are the main obstacles left to achieving that?
Maybe – one of my co-founders is Prof Matthias Niessner who's been behind a large chunk of the seminal and widespread research in this space.
[1] https://www.youtube.com/watch?v=ohmajJTcpNk [2] https://www.youtube.com/watch?v=qc5P2bvfl44