HNHacker News
TopNewBestAskShowJobs

mahmoudfelfel

84 karma · joined October 4, 2014

Cofounder @ play.ht
submissionscomments
mahmoudfelfel··on Play Dialog: A contextual turn-taking TTS model like NotebookLM Playground
The current deployed model is English only, we are rolling out a multilingual version later this week!
mahmoudfelfel··on Play Dialog: A contextual turn-taking TTS model like NotebookLM Playground
PlayAI (fma PlayHT) founder here, this is a native multiturn voice model that is built for conversations like real-time agents or podcasts. Try it through our playground (https://play.ai/playground) or API (https://docs.play.ai/). Feel free to ask anything.
mahmoudfelfel··on Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio
Very soon.
mahmoudfelfel··on Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio
The free accounts have 5k words and then you can upgrade to an API-plan from here https://play.ht/app/api-plans
mahmoudfelfel··on Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio
Society definitely needs to adapt to this new norm; we are trying to roll this out as safely as possible, but others are not as careful, and this technology will just become more ubiquitous over time.
mahmoudfelfel··on Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio
We have seen use cases in audiobooks, podcasts, marketing videos, explainer videos, Commercials, and Gaming, among others.
mahmoudfelfel··on Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio
Whisper is Speech to Text; we are building Text to Speech LLMs.
mahmoudfelfel··on Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio
Yes, we just released that for the UltraRealistic TTS (https://docs.play.ht/reference/api-getting-started), and it will soon be added to our Standard voices as well.
mahmoudfelfel··on Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio
What is in this demo is a very rate-limited, early version of our new model. We have many mitigations in place to increase the safety of our main product (Play.ht); I mentioned some of that here https://news.ycombinator.com/item?id=35331310
mahmoudfelfel··on Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio
We have been seeing some of these genuine use cases: youtube creators, audiobooks, elearning videos, podcasts, commercials, dubbing, and gaming.
mahmoudfelfel··on Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio
We have many mitigations in place to increase the safety of this service, I mentioned some of that here https://news.ycombinator.com/item?id=35331310
mahmoudfelfel··on Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio
You can try the high-fidelity voice cloning here https://play.ht/voice-cloning/
mahmoudfelfel··on Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio
Cofounder here, What you see in the above demo is a very rate-limited demo of our upcoming model. We realize how dangerous this technology can be and have built a lot of mitigations on our main product (Play.ht) to reduce possible abuse: - We strictly moderate the generated text of any sexual, offensive, racist, or threatening content. It automatically gets detected and blocked.

- We built and are offering for free a tool that can identify AI generated vs human-generated audio (https://play.ht/voice-classifier-detect-ai-voices/), we will continue to invest in this tool, and we hope it helps with deploying this technology safely.

- If we get any reports of a cloned voice without consent, we block the user and remove the voice instantly.

- The price of high-fidelity voice cloning is too high for scammers to use at scale; we have been live with it for four months and haven't had any cases of abuse so far.

Like any technology, it has the potential to be abused, and we are working hard to mitigate that and deploy it safely. We will continue to observe the use cases and user feedback and improve the safety of the service accordingly.

Since we launched voice cloning 4 months ago, we have seen enough genuine use cases which motivated us to keep moving forward and figure out safe ways to make the technology useful for all.

mahmoudfelfel··on Joe Rogan Interviews Steve Jobs
Everything, the content and the voices are generated by AI. check my other comments for how we did that.
mahmoudfelfel··on Joe Rogan Interviews Steve Jobs
Yes, a finetuned GPT3 on Jobs' biography
mahmoudfelfel··on Joe Rogan Interviews Steve Jobs
All AI generated.
mahmoudfelfel··on Joe Rogan Interviews Steve Jobs
Thank you :) that is the main point of it, to show people what is possible and inspire them to create.
mahmoudfelfel··on Joe Rogan Interviews Steve Jobs
No, All was GPT3 generated, check the other comment about how we prompted GPT3 to start the conversation
mahmoudfelfel··on Joe Rogan Interviews Steve Jobs
We will never use any cloned voice in any commercial way without consent and compensation, we only wanted to show the community what is possible and what generative AI models can do.
mahmoudfelfel··on Joe Rogan Interviews Steve Jobs
The original model (https://play.ht/blog/introducing-truly-realistic-text-to-spe...) was trained on 50k hours of audio, the above voices were just finetuned on the model, only 4-6 hours each.

We just finetuned another voice recently with only 1hr though... I think eventually (soon) we will only need 15-20 mins with zeroshot not even finetuning.

mahmoudfelfel··on Joe Rogan Interviews Steve Jobs
This was the prompt: " Podcast.AI Great people, great interviews with our host Joe Rogan. Episode 1 - Steve Jobs Summary: " Then GPT3 generated the summary, then we added: " Transcript: Joe: " That is all.
mahmoudfelfel··on Joe Rogan Interviews Steve Jobs
It was really hard to find good quality of Steve Jobs' voice, most of his speeches are keynotes on stage, I honestly was surprised that we managed to get to that quality from such poor quality recordings.
mahmoudfelfel··on Joe Rogan Interviews Steve Jobs
Hi everyone, cofounder of play.ht (the startup behind this podcast) here. let me know if you have any questions.

To give more context, the podcast was totally AI generated, the content itself was generated from a finetuned GPT3 on SteveJobs' biography, the voices were cloned from few hours of both Joe and Steve voices, even though it was tough to get good content for Steve Jobs. And the podcast artwork was generated by SD.

We will be releasing more episodes soon which will be even more mind-blowing!

mahmoudfelfel··on Podcast.ai E01:Joe Rogan Interviews Steve Jobs
AI generated podcast. The future of content creation is generated! in this podcast the voices were generated using a Speech Synthesis LLM finetuned with samples of Joe Rogan and Steve Jobs speech, The content was generated from a GPT3 finetuned with Steve Job's biography, and the artwork was generated using Stable Diffusion.
mahmoudfelfel··on Listen to Paul Graham's essays as a podcast
We have just launched https://play.ht, an app that helps you listening to the best articles from Medium.com and other websites and blogs from around the web. The above link is for Paul Graham's blog, you can find all his essays as audio. Give it a try and let us know your feedback.