HNHacker News
TopNewBestAskShowJobs

gkucsko

193 karma · joined February 24, 2015

submissionscomments
gkucsko··on Suno AI
Our models support full 2+ mins of coherent generation but generating a couple of verses at a time through continue gives good results you can keep picking the continuations that sound best!
gkucsko··on Suno AI
We’re US based but come from pretty much across the globe and my wife is Punjabi, so yeah that’s the origin of the name :)
gkucsko··on Bark – Text-prompted generative audio model
yeah sometimes there are definitely artifacts. technically they can be removed pretty easily with another model (like denoiser from FB) but for now we wanted to keep it simple to learn to control these things better through prompt engineering. Like when using a high quality input prompt it generally continues with high quality
gkucsko··on Bark – Text-prompted generative audio model
history is semantic, coarse and fine. so essentially the same thing thats getting generated just using it as an input before the generation
gkucsko··on Bark – Text-prompted generative audio model
thanks, the model itself is a pretty vanilla gpt model based heavily on karpathy's nanogpt, so should not need too many bells and whistles to get it running on specific architectures. that said i have very little experience with platform specific development, so would looove some help from the community :)
gkucsko··on Bark – Text-prompted generative audio model
Hey, one of the Suno founders/creators of Bark here. Thanks for all the comments, we love seeing how we can improve things in the future. At Suno we work on audio foundation models, creating speech, music, sounds effects etc….

Text to speech was a natural playground for us to share with the community and get some feedback. Given that this model is a full GPT model, the text input is merely a guidance and the model can technically create any audio from scratch even without input text, aka hallucinations or audio continuation.

When used as a TTS model, it’s very different from the awesome high quality TTS models already available. It produces a wider range of audio – that could be a high quality studio recording of an actor or the same text leading to two people shouting in an argument at a noisy bar. Excited to see what the community can build and what we can learn for future products.

Please let us know with any feedback, or if you’re interested in working on this: bark@suno.ai

gkucsko··on Bark – Text-prompted generative audio model
haha https://demo.suno.ai
gkucsko··on Bark – Text-prompted generative audio model
history prompts are just unconditionally generated TTS from the same model. any of those can be used as history, but for convenience 10 are provided for each language (to generate things with consistent voices)
gkucsko··on Bark – Text-prompted generative audio model
it's more meant to show code switching. more examples here: https://suno-ai.notion.site/Bark-Examples-5edae8b02a604b54a4...
gkucsko··on GPT-4-Audio: generative Text-to-Audio using Bark
Thanks :) Let us know if you come across interesting findings!
gkucsko··on GPT-4-Audio: generative Text-to-Audio using Bark
Thanks! The model learns a lot from unsupervised (as well as supervised) audio, so technically low-quality and high-quality audio are both just as likely to the model as music, background sounds or really anything else including echos or bad microphones :) will be interesting to learn how to control for these things, either through prompting or other switches during training/inference
gkucsko··on Open source PowerPoint add-in for broadcasting presentations in a browser
There is no sign-up (microsoft account or similar) necessary, the broadcast link is much shorter and easier to share, and it's open source :)