FOSDEM 2023 Transcribed by Whisper
jonatron.github.io
jonatron.github.io
https://jonatron.github.io/fosdem2023whisper/files/fast_data...
Original: https://www.youtube.com/watch?v=JlcI2Vfz_uk
But some mistakes are hilarious!
> 500 megabytes per second, Britain.
Britain. Guess what it actually was :)
- coreutils => corridors/coroutils/cori-teals/curricules/corretails
- Rust => rest/Resi
- glibc => GLC
- Gentoo => Gen2
- Fuchsia => Fushia
- SELinux => AC Linux
Other proper nouns it transcribed well - LLVM, Mozilla, Debian, Ubuntu, clang. These are much better known I suppose.
It finds the speaker's French accent a bit difficult, because I see a few mistakes like "find => fine", "imposter => imposterous", "command => comment".
Overall pretty good. The only thing I'd like to improve readability is a pass with an LLM to break it into paragraphs but that would be way more compute and no longer local (like Whisper).
The result is not that far. There are some errors of course, notably on acronyms and technical terms ("IBM", "GNU", "GitHub"), but a lot of the mismatches are in areas where both models got things wrong.
Neither of the models got "FOSDEM" right, though Whisper is consistent at using the word "fosstem", whatever that is?
In terms of resource usage, for this and many other "emerging" opportunities, is there some thinking about pooling idle CPU's/GPUs for creating open and widely available models and/or datasets based on models?
SETI type crowdsourcing of compute was always a bit niche but maybe now we have some use cases of very broad interest
https://www.newyorker.com/tech/annals-of-technology/whispers...
For them, the key difference (other than it's great performance) was that it was a standalone executable that did not need to access remote resources or "share" sensitive information with such remote resources.
Also, since OpenAI was already used, it could have been cool to provide a summary of the transcript!
There is also a CPU only version discussed a few times: https://news.ycombinator.com/item?id=33877893
I liked this idea, so I opened a PR for exactly that: https://github.com/jonatron/fosdem2023whisper/pull/2
It would be a very low quality blog-spam level blog post.
(I'm not able to make a pull request because GitHub is giving me the pink unicorn error page when I try to fork your repo.)
What solutions are Twitter spaces and the like using for real-time live captions?
Do you happen to know what kind of EC2 instance (or equivalent config) would be required to run it?
2. Use the timestamps provided by whisper to group and align with (1).
> So, you can see my Firefox is in Italian, but you can see that it automatically detected that the page is in French and it is suggesting me to translate it to Italian. I will change it to English. Oh, fuck.
> The reason for that is because there is so many ways to extend or to define an architecture, and there is some really fucked up things that can be done in so many architectures
> We don't measure RSS because it measures anything useful we measure it because it's really fucking easy to measure.
> There were a few customers that didn't respond, but at some point you just have to decide to. Don't give a fuck.
> Computational storage, what the fuck's that?
> like on MacOs you can have a recovery key, on Windows you can have a recovery key, and this is what we should do here as well, right, like use TPMs absolutely, and then hide entropy recovery keys that things are fucked up, things are fucked up
> we get money, we invest in open source, and then we have the problem that what we do in open source could be reused by non-European companies. And thinking, oh, my God. Oh, that's so fucking wrong. Why do you do that?
> So, you know, if Monty is not screaming at me saying you are fucking moron, Peter, that is not how it is, then probably I am not doing my job properly.
> I'm going to wail and gnash my teeth and hope it doesn't happen in practice because at the moment, you know, something like EDK2, to be clear, by the way, I don't give a fuck about non-free firmwares
> Is there anybody here who doesn't know what a kubelet is? Okay. The kubelet, oh, you fuck, anyway.
> So for now we're targeting your own CA so you can just say fuck it, I will enroll my own keys to the firmware.
> So one of the idea is to start to come to a more predictable fashion, which is one major API break and API break every year around December, January, so we're in February, and we fuck this year.
> It's really not sure whether we are going to be able to do that because you're transpiling assembly. Like, what the fuck are you talking about?
> your problem was where you called, in this case, pop on an empty stack, and not the completely pointless unwrap in there, because you want to know where you're fucked up and not here, obviously.
> So if it is, then say, say like, this is fucking wrong, Peter, you know, so I can fix my slides when I talk next time, I have the wrong stuff, right?
Used a small bash script to run `whisper $f --output_format json --language en | tee $f.txt` over every file.
Ran a small python script to make the HTML. All totally trivial, it just required some GPU-hours.
curl https://jonatron.github.io/fosdem2023whisper/files/celebrating_25_years_of_open_source.webm.txt | ( echo "WEBVTT"; sed -e 's/^\[\|\] \?/\n/g' ) >celebrating_25_years_of_open_source.webm.vtt
Plays in the browser with HTML like this: <video
controls
src="https://video.fosdem.org/2023/Janson/celebrating_25_years_of_open_source.webm"
>
<track
default
srclang="en"
src="celebrating_25_years_of_open_source.webm.vtt"
>
</video>https://developer.mozilla.org/en-US/docs/Web/HTML/Element/au...
> The <audio> element doesn't directly support WebVTT. You will have to find a library or framework that provides the capability for you, or write the code to display captions yourself. One option is to play your audio using a <video> element, which does support WebVTT.
You can access the resulting WebVTT files here: https://github.com/tuukka/fosdem2023whisper/tree/feat/webvtt...
-E Interpret regular expressions as extended (modern) regular expressions rather than basic regular expressions (BRE's)
-r Same as -E for compatibility with GNU sed.
Thanks for providing the complete set of vtt files. I found some interesting transcription errors in the talk on Elliptic curves in FOSS. For example "jealous two curves" instead of "genus two curves" and "ethyl curbs" instead of "elliptical curves"I'm trying to learn some Spanish, and a fun and easy way to study is to watch films in Spanish with Spanish-language subs.
Finding Jurassic Park on piratebay was easy, finding the right subs was impossible. Enter Whisper.
My favourite little hack this year.