HNHacker News
TopNewBestAskShowJobs

JonathanFly

824 karma · joined July 9, 2019

submissionscomments
JonathanFly··on Suno AI
I got to 18 minutes. Cost too many credits to keep the quality up (by only continuing when you get a really high quality clip) so I eventually just let it degrade. https://app.suno.ai/song/6f334b5c-c992-446b-8b46-2227c34e730...
JonathanFly··on An update on Twitch in Korea
According to https://news.ycombinator.com/item?id=38539167#38539479 Twitch is the most popular streaming service in Korea, just slightly ahead of Afreeca.
JonathanFly··on An update on Twitch in Korea
Are bandwidth fees just as high for any domestic Korean company creates a Twitch competitor?

Or do the 3 major telecoms operate their own streaming services, so only companies under their umbrella can afford a site like Twitch?

JonathanFly··on An update on Twitch in Korea
>Mentality: SK Terran is a highly aggressive style that is intended to pressure the Zerg with a large number of MnM and Science Vessels. The ideal scenario is to split apart a big MnM ball and engage in guerrilla tactics around the map while irradiating key Zerg units such as Defilers, Ultras, and Lurkers.

I assumed this was a humorous analogy for this Korean policy but and it's going over my head (though I only know SC2) or maybe you meant to link SK Telecom https://liquipedia.net/starcraft/SK_Telecom_T1

>SK Telecom is a mobile telecommunications operator in South Korea that is a part of the SK Group, one of South Korea's largest conglomerates.

JonathanFly··on An update on Twitch in Korea
>Entertainment and the infrastructure that delivers it is an important pillar of soft power.

Might be a miscalculation if the goal is soft power as in cultural export (the so called Korean Wave). At least in the gaming/streaming space where Korea has been punching above it's weight in cultural impact since the StarCraft 2 days.

My guess is this policy results in Korean streamers on Korean sites streaming to Korean-only audiences. Reduces the worldwide crossover reach of Korean cultural export at least near and medium term.

Though YouTube is probably adequate for this if it can stick around...

JonathanFly··on StyleTTS2 – open-source Eleven-Labs-quality Text To Speech
>Are infill and outpainting equivalents possible?

Do you mean outpainting as in you still what words to do, or the model just extends the audio unconditionally the way some image models just expand past an image borders without a specific prompt (in audio like https://twitter.com/jonathanfly/status/1650001584485552130)

JonathanFly··on ElevenLabs Launches Voice Translation Tool to Break Down Language Barriers
It autodetects language by default but you can set to a specific one. Though you'd still have that problem on a multi-lingual input video.
JonathanFly··on ElevenLabs Launches Voice Translation Tool to Break Down Language Barriers
Genuinely impressive used as intended. Spectacular when not - English to English translation.

https://twitter.com/jonathanfly/status/1711607166371561805

JonathanFly··on SoundStorm: Efficient Parallel Audio Generation
>So how long will it be until we will be able to download something of this quality onto a future-gen Raspberry Pi which can do some AI processing, where we make an HTTP call and it starts speaking through the audio out in a perfect voice without relying on the cloud?

5 years? It's probably possible roughly whenever the larger Whisper models can run on it. Probably the next Raspberry Pi, running quantized or optimized versions of some audio model.

It may be almost possible right now if you tried really realy hard, and you used a small model fine-tuned on a single voice, instead of something larger and more general purpose that can do any voice. I think whisper-tiny works on a Pi on real time, right? And that's not leveraging the GPU on the Pi. (https://github.com/ggerganov/whisper.cpp/discussions/166)

Edit: looks like medium is 30x slower on the Pi than tiny model, so I may have been overly optimistic. I didn't realize Whisper tiny was that much faster than medium.

This method works pretty well with Tortoise, letting you use the super fast Tortoise quality settings but get quality similar to the larger models. Fine-tuning the whole thing on just one voice removes a lot of the cool capabilities of course. With Tortoise, that would still be way too slow for a Pi but potentially that same strategy could work with faster models like SoundStorm.

In terms of quality there's still a lot of room to go with long term coherence, like long audio segments. When a real person reads an audiobook the words at the top the page have a pretty big impact on how many words at the bottom the page are read. And there can be some impact at any distance, page 10 to page 300. When you try audiobooks on super high end TTS models and listen carefully you really notice the mismatch. It's like the reader recorded the paragraphs out of order, or a video game voice lines where you can tell the actors recorded all the lines separately, and were not reacting to each other's performance.

You can bump the context windows, a minute, two minutes. That's gonna get you closer and probably good enough for some books. In the short term a human could simply adjust all the all the audio samples and manually tweak things to sound correct. So this will enable fan-created audiobooks where they take the time to get it right. But for fully automated books the mismatch drives me nuts. The performance is just soooo close for certain segments that when you get a tonal mismatch it hurts.

JonathanFly··on SoundStorm: Efficient Parallel Audio Generation
Yeah I often try to think about what might be in a YouTube caption when finding prompts that work in Bark. But pipe character isn't one I remember seeing on YouTube. Maybe it's part of some other audio dataset though. Or maybe it's on YouTube but only in non English videos.
JonathanFly··on SoundStorm: Efficient Parallel Audio Generation
Interesting that SoundStorm was trained to produce dialog between two people using transcripts annotated with '|' marking changes in voice. But the exact same '|' characters seem to mostly work in the Bark model out of the box and also produce a dialog?

Maybe a third or a bit more of Bark outputs are a dialog person talking to themselves -- and it often misses a voice change. But the pipe characters do reliably produce audio that sounds like a dialog in the performance style.

https://twitter.com/jonathanfly/status/1675987073893904386

Is there some text-audio data somewhere in the training data that uses | for voice changes?

Amusingly, Bark tends to render the SoundStorm prompts sarcastically. Not sure if that's a difference in style in the models, or just Google cherry picking the more straightforward line readings as the featured samples.

JonathanFly··on Building Boba AI: Lessons learnt in building an LLM-powered application
I just realized we already chatted briefly about banned tokens (I'm the Grover tongue twister guy) but I somehow completely missed this gist at that time. Total facepalm moment, would have been helpful reference.
JonathanFly··on Building Boba AI: Lessons learnt in building an LLM-powered application
AHhhhhhhhhhhhhhhhhhhhhhhhhhhhh. That's me screaming. I constantly wonder why this stuff does not exist in LLMs. But my technical depth and competence is quite low. Way lower than the people implementing the models and samplers. So I just assume: there must be a good reason, right. Right?

But recently I threw a just a bit of similar-ish stuff as you describe there into a TTS model, barely knowing anything, and yeah it's totally works and is fun and cool. The stuff that doesn't work fails in interesting and strange ways, so it almost STILL works. (Well, it gives people really bizarre speech impediments, at least...)

I was just working on prompt editing actually. Which is weird to imagine in a TTS model. It makes sense for the future tokens of course, for words the model has not said yet. I think it even makes sense for the past right? You can rewrite the past context, and it still changes future output audio model. In bark it's two different things: one is the text prompt, and one is the generated audio tokens/context, which is not the same. (The text and the past audio is concatted in the Bark prompt, so this idea makes sense in Bark but not in other models. You could change either text OR 'what was generated with the text' independently.)

As long as you don't rewrite the time touching the last token, at 0 seconds - if it's like a segment 2 to 4 seconds in the past, it should influence future output but not cause a discontinuity in the audio. I think?

BTW an easy and fun thing - just let generation parameters be dependent variables. Of anything.

A trivial example: why is temperature just a number, why not a function? Like the temp varies according to how far long in the prompt you are. For music, just that is already a fun tool. Now as a music segment starts or ends the style transitions. Or: spike the temperature at regular intervals - like use a sine wave for temp, input is current token position. You can probably imagine that works great in music model.

Even in a TTS model this you can get weird and diverse speech patterns.

The thing is: I really very a low level of competence. Total monkey hitting keys and googling, and even I can make it work, easily. Sampling is just a loop, okay, what if I copy logits from sample A and subtract them from sample B. What if take the last generation, save the tokens, ban then in the next. Really just do anything and you end up in interesting places in the model you didn't know existed and are often cool. (Recently, TTS output overlapping speech, for example.)

Like I recently generated french accents from any voice in the Bark TTS model, with no fine-tuning, no training, actually not even really any AI. Just by counting token frequencies in the french voices, and having the sampler loop go, "Okay let's bump these those logits up a bit, and the others down" and it just somehow works. No Loras, no fine-tuning, not stats, it's like middle school level math, but sounded great.

(I'm in a bit of a stream of consciousness ramble mode from lack of sleep, but I'll keep going on this message anyway so I don't forget to come back to your post when I'm back at normal capacity. And just hope I don't cringe too hard reading this when better rested.)

Oh I'd love to hear your thoughts on negative prompts in LLMs.

1) What does 'working correctly' look like?

For an audio LLM, I'm thinking something like: a negative prompt "I'm screaming and I hate you!!!" makes the model more inclined to generate quieter, friendly speech, in your positive prompt. Something like that?

2) How to make it work.

This is probably very model dependent and fiddly. My first thought was generate two samples in sequence. The first sample is the negative prompt. Save all the logits and tokens. Use them as a negative influence in the second prompt. At least in Bark you can't just like flat subtract them or what you actually get is more like 'the opposite of speech' than 'the opposite of your prompt' but when I did french accents I basically just fiddled with a bunch of constant values and weights and eventually it worked. So I'm hoping the same applies. I can imagine a more complicated versions where you do some more math to figure out what's unique about a text prompt, versus 'a generic sentence from that language' and only push on those logits. I suppose that might be necessary.

JonathanFly··on Show HN: Mofi – Content-aware fill for audio to change a song to any duration
What's the difference between "Search" and "Edit song"? It seems like Search is just about length, and Edit also considers sections you marked? Or does Edit do the same thing but with many edits, rather than a few?
JonathanFly··on Bark: A transformer based text to audio system
Replying to myself in an old thread as a little easter egg. Bark is just so fun I can't resist a teaser.

Suno is seriously underselling the power of the fully operational Bark model. I'm already cranking out "French Obamas" and I didn't know anything about TTS a month ago. Heck, I still barely know anything. (Obama has been annoyingly resistant to gender flipping though.)

https://drive.google.com/file/d/1ZbJYXoH8gmrEyMe1AJ0VdwzZkf_...

JonathanFly··on Bark: A transformer based text to audio system
>The way it tends toward context-sensitive shifts in delivery is kind of amazing

A blessing and a curse! Super cool though.

JonathanFly··on Bark: A transformer based text to audio system
Leverage sure, but you still basically need a model that does the opposite. You can't like take OpenAI Whisper, which turns speech to text, and just run the model backwards and generate audio. For example. I mean you probably could with some work, but not out of the box.
JonathanFly··on Bark: A transformer based text to audio system
Must be some low hanging fruit to optimize in Bark. It would be somewhat close to realtime if it was close to 100% and scaled linearly.
JonathanFly··on Bark: A transformer based text to audio system
Yeah, it's kind of hand crafted. There's more to the story and more results. I would normally just Tweet but I think it's actually so interesting that it deserves more than a tweet, at least a thoughtful writeup or a youtube video. (And I need to catch up on real work this week first, so end of week at best.)
JonathanFly··on Bark: A transformer based text to audio system
Actually I was just checking, and Bark isn't that close to maxing out GPU utilization. Running two instances on a 3090 seems like a throughput increase and the models fit. Update: And getting weird CUDA issues. Hmn...
JonathanFly··on Bark: A transformer based text to audio system
I barely touched that, it's just from the Serp cloning github, but people kept asking so I put it in. Their clone just isn't really doing much though, it's just loading up the coarse model with the encoded wav file as a fake last generation history to the current segment, and that's basically it. But a lot of the voice is in the semantic model so just doing the coarse doesn't get you much. And even if the semantic model didn't matter as much, the coarse model has both semantic and coarse tokens as inputs and your injected coarse tokens aren't going to line up just right with what a true Bark generated pair of tokens would look like. So what you get is like a robot clone that has the most superficial similarity and lacks the depth of cadence that makes Bark awesome. That's in best case, more often get voices full of static or that don't even read the text you give them. (To be fair, any bark voice can do occasionally not read the text, it's a risk.)

I'm sure somebody will train a model that actually maps an input text to the Bark semantic representation, it shouldn't be that hard, it's just been a few weeks. But the existing clone there is just so primitive I don't know how it got so popular.

JonathanFly··on Bark: A transformer based text to audio system
Thanks. I'll consider it but I haven't deployed a model or service like that, so it'd be more of a second project in itself than a funding mechanism, probably. I was just realizing this morning how far behind I am on paying work from getting overly distracted by Bark lately. And it's a lot behind so it was kind of wake-up call. Though some of that is adding new features and trying new ideas (some to be seen later on the public fork) not all is support and stuff. Bark is a wild model, a lot of silly ideas kind of work.
JonathanFly··on Bark: A transformer based text to audio system
Not used at all, at least in my case.
JonathanFly··on Bark: A transformer based text to audio system
>Yeah I was trying to figure out how good it was in Korean. The cadence and flow was pretty good but there was kind of artifacts in the audio. Then I check the samples of the default audio prompts for Korean any my god, they were godawful. Switching it up made a world of difference.

I have a decent amount really clear Korean voices BTW. That and French, from people asking on Discord. But I can't judge the accents only that they are clear speakers.

Korean was interestingly the lone somewhat-coherent one-shot long term music sample I ever managed to get out of Bark.

https://www.youtube.com/watch?v=4pV9d25KqCE

The second music bit in this Youtube was one continuous generation where the last prompt was used as the history for the next, with no cherry picking or assembling the clips, just one solid segment. And it sort of holds together for like almost a minute!

I was excited because I thought maybe Bark could be like a real-time OpenAI Jukebox. But that was literally the only time so far where using a full-feedback held together like that. You can kind of 'cheat' it to by using a very popular song as the input text, and sometimes Bark will produce the appropriate melody. But of course that's not really the point of using your own text. I have some ideas for making it more coherent, but nothing easy. Too bad, Jukebox is just SO SLOW.

Actually with what I know now I should re-render this and clean up the distortion. At the time I couldn't do it. Though I only have the first segment prompt.

I have a decent amount really clear Korean voices BTW. But I can't judge the accents, only that they are clear.

JonathanFly··on Bark: A transformer based text to audio system
I'll link my Bark fork with long audio generation and other features on the root thread, I suppose: https://github.com/JonathanFly/bark

There's going to be a big update this week with some new stuff I haven't talked about. And a bunch of amazing, clear voices, with a huge variety of styles, that blow the default Suno voices out of the water. Arguably even better than Eleven in some ways. I'm excited even though I have nothing to DO with the voices!

Don't get too attached though. I was just playing around and made a Bark fork and it got more popular than expected. And now I'm dreading a future full of hours of unpaid support and maintenance that I definitely can NOT afford, for a software product I don't even really have a personal use case for. I'm not generating my own audiobooks or anything, I won’t be using it long term myself, I was just curious what Bark could do. (Turns out a LOT more than you might think at first glance, as you'll see this week.) So I'm already trying to work out how I can elegantly wind this thing down and transition people somewhere else. But I'll keep it updated for at least a little while.

JonathanFly··on Bark: A transformer based text to audio system
It's just a characteristic of some of the default voices. I made some perfect French voices by having somebody on Discord check for native accents, because otherwise I can't tell. Once somebody does this for every language it should be all set. (Somebody who isn't me though, it takes awhile and I'm just playing around a bit, don't even have a use case for TTS honestly.)
JonathanFly··on Bark: A transformer based text to audio system
It's like 50% realtime on a 3090, not quite real time on a 4090. You can also use smaller models and it's a bit faster.

3080 is same speed, you don't need the extra memory.

JonathanFly··on Bark: A transformer based text to audio system
>I'm curious, how did you generate the David Attenborough voice? The repo says: >> Bark tries to match the tone, pitch, emotion and prosody of a given preset, but does not currently support custom voice cloning

Check back later in the week, I'll have a bit more on that later after I catch up on actual work and can write a bit.

JonathanFly··on Bark: A transformer based text to audio system
>when i tried a similar on elevenlabs, it sounded a lot worse in terms of the metallic s-s at the end.

That's interesting. When I'm judging Bark I'm looking at my own random samples, but for eleven I'm seeing stuff people post on Twitter or YouTube, which I suppose must be cherry-picked. I didn't even realize Eleven did the same thing!

JonathanFly··on Bark: A transformer based text to audio system
I actually think Bark actually beats Eleven right now. Bark tends to add a bit more of a metallic twinge towards the end of longer audio segments (I wonder if this is a bug and not a limitation actually...), but Bark is more expressive.

The trick to exceeding Eleven quality is using multiple speaker prompts, and swapping them in and out over the course of longer texts. Each is a slight variation on the original. If you hand tune the swaps it's super good, but if you just swap in 'beginning of new paragraph' 'continuation speaker' automatically, that alone is a huge boost.

I don't have a clip handy but I will make a YouTube or something this week, because the quality is wild if you do this. There's some passable audio clips on my README here https://github.com/JonathanFly/bark but all those are just using the same speaker for every line. That was all just the first samples I tried, from weeks ago, without really putting any effort into it. You can do a lot better.

The extra expressiveness of Bark does come with it being harder to control and having a bit of a mind of its own at times. (And this can also be very very funny: https://twitter.com/jonathanfly/status/1657658109001596929)

So for a production or real time use case Eleven makes more sense. For example even my best speakers will switch to a new voice mid prompt, once in awhile. (As if the audio clip was from interview segment.)

← PreviousPage 2 of 6Next →