SeamlessM4T, a Multimodal AI Model for Speech and Text Translation
about.fb.com
about.fb.com
One limitation that seems undocumented, the current code only supports relatively short clips so isn't suitable for long transcriptions:
> ValueError: The input sequence length must be less than or equal to the maximum sequence length (4096), but is 99945 instead.
Edit: unless there is native speaker diarization. That would be a huge value add.
It's especially interesting how you could combine different model types - e.g. translation + text completion (or image generation) – it could be a pretty powerful combination...
Much smaller language matrix though.
[1]: https://github.com/facebookresearch/seamless_communication/b...
How things change dramatically in one year with such exaggeration of Meta’s collapse in 2022.
Not only they are in the lead in $0 free AI models, they are also at the finish line in the AI race to zero.
That’s so embarrassing - especially for something to show how good their stuff is (although I think it’s probably not the ai’s fault) - just shows how sloppy their people are.
I know they have plenty of Vietnamese engineers there. Did the PR dept just throw this final version of the video out without reviewing with them?
To more accurately evaluate the system without depending on text-based metrics, we extended our text-less metric into BLASER 2.0, which now enables evaluation across speech and text units with similar accuracy compared to its predecessor. When tested for robustness, our system performs better against background noises and speaker variations in speech-to-text tasks (average improvements of 37% and 48%, respectively) compared to the current state-of-the-art model.
SeamlessM4T also outperforms previous state-of-the-art competitors.
https://github.com/facebookresearch/seamless_communication/b...
I don't really know how the metrics they list compare to whisper, I'm very curious if these are fast enough for realtime speech2text? I think whisper technically could but it was difficult to do or something like that?
That said, I fully support open releases and look forward to future versions and improvements.
The CC BY-NC 4.0 license allows for the following uses of the licensed material:
* Reproduction: You can copy and distribute the licensed material in any medium or format.
* Distribution: You can distribute the licensed material to others.
* Public performance: You can perform the licensed material publicly.
* Public display: You can display the licensed material publicly.
* Modification: You can remix, transform, and build upon the licensed material.
* Derivative works: You can create derivative works based on the licensed material.
However, there are some restrictions on how you can use the licensed material under the CC BY-NC 4.0 license:
* Commercial use: You cannot use the licensed material for commercial purposes.
* Sublicensing: You cannot sublicense the licensed material.
* Moral rights: The licensor retains all moral rights in the licensed material.
Here are some examples of how the CC BY-NC 4.0 license can be used:
* A teacher can use a CC BY-NC 4.0 licensed image in a presentation for their class.
* A student can create a CC BY-NC 4.0 licensed remix of a song.
* A software developer can use a CC BY-NC 4.0 licensed library in their open source project.
* A photographer can share their photos on a CC BY-NC 4.0 licensed website.
To be fair to you, i agree. HN is probably the last place you want to use low-effort AI comments, regardless of how helpful they may be. Let’s leave the AI comments for Reddit.
The AI research environment has changed from the earlier default-open publication - unlike it's competitors, FAIR is still releasing model weights instead of serving the models behind an API.
> this is the equivalent of a kid licking a cookie so the others can't eat it.
More like the other kid baking a cookie with the words "Free Cookie" on it so others can eat it if they are hungry, but can't sell it for money. It'd be foolish for FAIR to donate preconfigured homing-missiles to OpenAI and others via one-way tech transfer.
It'd be foolish for FAIR to donate preconfigured homing-missiles to OpenAI and others via one-way tech transfer.
No, they could GPL it, and I don't think they're worried about competition taking the models anyway, there's nothing particularly special about the weights or training data, just the compute. I think part of it is pressure from AI "safety" hangers-on who pretend that AI is dangerous so only those who don't want to abide by license terms should have unfettered access. The other commercial reasons are harder to understand. With pytorch they became the standard that everyone builds off of, they could do that with their recent AI, particularly LLaMA but they chose this silly route.Also, LLaMA has a more permissive license than this translation one, and is a more powerful model, so I don't really see the "homing missiles to open AI" angle.
If that is the case, then what do you suppose is the reason most researcher outfits stopped releasing model weights, or offer more restrictions when they do?
Using the GPL won't prevent the larger AI competitors from using your model outputs from tuning their non-public models to consistently beat yours, but a non-commercial clause does.
> Also, LLaMA has a more permissive license than this translation one, and is a more powerful model, so I don't really see the "homing missiles to open AI" angle.
LLaMa lags ChatGPT 4, but SeamlessM4T is ahead of WhisperX in some ways
Their models are not copyrightable and I would not waste this opportunity to establish this fact.