The Seamless Communication models
ai.meta.com
ai.meta.com
The "universal translator" which was part of Star Trek and a lot of other Sci-Fi I was exposed to as a kid was something I was really fascinated with. My Dad worked as a simultaneous French->English translator and sadly spent long hours away from home and, as a kid, I started trying to build a translator so that it could do his work and he could be home more.
Translation is important work and one that could help a lot of people. It's my hope that we get to the point where these models work entirely on locally carried resources.
By keeping creating speaker specific tonal ranges and profiles you maintain the better cohesion on the final product.
I grew up bilingual outside the US, and speak English with a hybrid British/Indian/Middle Eastern accent (with some of my personal quirks, and mixing increasing amounts of various American accents over time). I can understand English in nearly any accent (Singaporean, Chinese, Vietnamese, Indian, Nigerian, eastern European) as long as the words involved are globally used and the grammar is passably queen's. Especially after hearing it for about an hour. And people who natively speak English with these various accents usually can understand my English better than they can an average American accent. Yet in this country, my accent is belittled, despite being perfectly understood and more versatile. Even by others who don't speak with the American accent!
This is the problem of the "default accent" anywhere being referred to as "no accent", and therefore anything deviating is considered "having an accent". This makes "accent" a negative trait, scaling from 0-bad to heavy-bad. But if the vernacular were such that we said "American accent" instead of "no accent", then noone's accent is bad, just not used to.
Most of my non-American peers who were raised on English have a better command of the language than my American ones, yet they are mocked for their accents as if they don't know the language, when in reality it's the Americans lack of familiarity with the language (as its used globally) preventing them from comprehending the language.
So yes, put in more work, the world is shrinking and English is the global language (for better or worse). What you're saying is spoken from a position of privilege because the culture allows you to mock others' accents and imply your version of it is the correct one that everyone else should put in work to provide you with, rather than the other way around.
Every time you hear English with an accent other than British, American or Australian, remember that it usually means the speaker knows at least one entire other language as well, probably one that you would sound like an idiot if you tried to speak it. Don't be rude or dismissive of their command of English.
In fact, you were so close — you called it a "no accent am-english", when you could have just called it what it is — "an american accent".
I could of been more specific, but my request for the tech to vary, I think would lead to specific options for different people.
And actually to be even more.. not sure the word.. I want 'the Chicago accent' I think it's called, or midwest / no accent. Personally as much as I enjoy some entertainment from Jersy / NY accents, I would not volunteer to watch tutorials on tech taught by the Sopranos cast - as funny as that might be (and I get if you are from the NE, you may be learning just fine being taught with such a language style).
As annoying some of the Cali style of language is, I can understand the words and meanings without squinting my ears and spending double the brain cycles trying to understand the words, while then interpreting the meaning, and then trying to put together concepts for understanding new ways of coding or using tech.
I've run into folks in Louisana that I could not understand at all and had to ask for an interpreter at a gas station. From Florida to Chicago to Seattle down to Miss and Ala - I can hear what people are saying and learn without spending lots of extra energy trying to understand.
With that being said, I understand there are parts around Miami where accents may be thicker (or not) - and with some folks even if using the rights words and grammar, I may need to slow down the speech to actually learn if they were teaching a class.
The slow down and speed up options already exist with youtube.
"So yes, put in more work"
- I do try a bit. I don't mind accents with some folks and media.For example I can listen to and enjoy Shankar sharing via the 'hidden brain' series, partially because his accent is limited but also because the media requires less thought intensity.
I have tried many youtubes, and bought a few courses taught from folks in India and other places where I just could not muster the energy. I literally squint with my ears and feel like my head gets hot trying to decipher what is being said, translate into what is meant, and how it should create new patterns of understanding in my brain.
I can only do that for so long and I am done. Now I just skip any learning video that has non-am English speakers. When I consider courses to sign up for or buy, I have to research the authors / speakers and find video of them to hear the audio, because I just can't learn well that way.
"other than British," - True story, a few years ago I had to call an ISP in Britain(?) and the person I got to to file an issue with, I could not understand them. I had ask 'what did you just say' many times. I laughed at myself for even thinking of saying 'can you slow down and speak clearer English please' - I mean, crazy... I was paying by the minute for the long distance at the time and it ended up being a 25 minute call that could of been 10 if I had a magic translate without accent device.
"a position of privilege because the culture allows you to mock others' accents"
- This is truly not about mocking accents, this is truly about my lack of ability to learn well.
Yes, I would defintely sound like an idiot trying to speak another language. Like I said, I do not learn as well as some others.
Truly not my intent to be rude. I apologize if the shortness came off that way, I was trying to be brief in the hope that there's a chance that some tech like this exists and someone here could point me to it. Before I posted, I DDG'ed it and found a couple of things attempting to be in that space with a 'speak to sales' type of 'you'll never afford this' button for info.
I will never be dismissive of anyone's command of English, or other spoken language, or computer language or anything like that. There is no way for me to know someone else's situation and circumstances led them to their current command of whatever language. If someone is trying to learn more at any age; I applaud and encourage them - being rude or dismissive does not encourage more learning.
"no accent am-english", when you could have just called it what it is — "an american accent". - Well maybe, but actually I meant to be more specific, as mentioned a bit above - I mean '"no accent" American accent' - because there are plenty 'American accent' types that I would want removed by a magic earpiece to make it easier for me to understand and learn.
There is no "no accent". An accent is a baseline feature of intelligible human speech, like a voice, or a volume, or a language. You can't say stuff without those features. When you say "the Chicago accent", or the "Midwest accent", that's an accent! Not "no accent".
I understand it's common usage to refer to the default "radio accent" as "no accent", but in a country like America, all kinds of people with all kinds of accents speak English. Reinforcing an expectation that a certain (usu. majority-white-spoken) one is the "default" by referring to it as "no accent", implicitly suggests all others are erroneous affectations, even if I trust that is not your personal intent.
All that said, I think your idea for a translation device capable of revocalizing what is said with an unfamiliar accent into one you are used to is not a bad one, and likely easier than translating between languages while retaining expressiveness.
Wow, you just keep digging in don’t you? When these Americans you deride say “no accent”, do you think they are referring to the “majority-white-spoken” Scottish accent?
No, of course not. Get that race baiting out of here.
I believe I had heard the term 'the chicago accent' as a term for radio broadcast that may have come more from 'the chicago market' or mid section of the US.. not meaning specific urban/city speak but the market segment for that part of the country as opposed to the east coast market, etc.
Looking at the top 12 radio markets; New York, Los Angeles, Chicago, San Francisco, Dallas-Ft. Worth, Houston, Atlanta, Philadelphia, Washington, Boston, Detroit, Miami
If you pick any of those and think of broadcasting the base accent of that city to the rest of the country.. I could see most annoying the other cities.. maybe Philly being a close 2nd.. and I don't think anyone can determine what a base accent for DC is.. that would depend on which neighborhood or however to median that.. (and I have no idea what the accent / tone / language is of Frisco in general, never been, and maybe it was different 20 years ago vs today)
And I mean, yes, there are people who know they don’t sound like whatever ideal accent they have in mind, and there are people who will make fun of accents - but, and I can’t stress this enough, depending on the context literally any accent can be made fun of, sadly. I’ve had people mock my “American” accent while travelling, for example. It sucks, but it’s not easy to single out any accent as “default” unless it’s literally enforced by a government and taught that way in schools. Last I checked, the US is not one of those countries and English is not as centrally controlled as e.g. French can be.
What accent? Whose accent? Brits are as diverse accent wise as Americans, London, cockney, New England, Southern...
A lot of Indians that I know have a very "proper" British accent, one that maybe a bit aristocratic, its quite an irony for a former colony. https://www.bbc.com/future/article/20220915-what-the-queens-...
The context matters, but so does history.
There are many american accents. Your suggestion makes the sentence much less clear.
And by specifying "american" they're already making it clear there is no such thing as a universal base accent for english.
Question is, what will happen to the tribal / local languages? Will they survive?
This might be required to get full buy in for a unified language, which is a bit sad but makes some sense - if you ensure it's taking up more and more of media and culture more people know it from immersion, and other languages are reduced to being spoken at home / with friends and that's going to cut into how many people really are fluent in them.
In my other languages that I am actually fluent in, it’s kinda the same — you use specific suffixes to soften or embolden your point and so on. Maybe add “exclamation making sounds in specific language” too. Eventually your nouns and verbs end up in different languages, with different suffixes where it “makes sense”, yet the person whom you’re talking to will “get it”.
Would be curious to try the new Seamless model on such speeches.
The only thing that doesn’t work is if you talk to people too young to remember Skype. Then you feel old.
Verbing a noun is pretty simple; French borrowing English nouns as loanwords then verbing them is another. I was expecting something intermixed and trilingual.
"The word “brunch” doesn’t exist in French": well not the official French language per the Académie Française, but functionally it does, once everyday French-speakers start using it [0] "Many were quick to point out, as Reuters reports, that French President Emmanuel Macron commonly uses English idioms, including “start-up nation” and “bottom-up.”". I'm guessing Québecois people are even more fluid about this, since they have to be functionally bilingual(/trilingual) in daily interactions, and France's Toubon law doesn't apply to them.
For example, it would be interesting to chart the ratio of the Académie-mandated 'courriel' vs 'email' in majority-French-language posts in various regions. [1]
[0]: ["How France Tries to Keep English Out of Public Life" (8/2019)](https://news.ycombinator.com/item?id=20730219)
[1]: [Should You Use "le Courriel" for "Email"? (2019)](https://www.thoughtco.com/le-courriel-vocabulary-1371793)
Ne smotra na fact that tomorrow bir o kadar da sunny diyildir, bilo bi horosho progulyatsa around the seawall.
(Despite the fact it's not really sunny tomorrow, it would be nice to go for a walk around the seawall)
It's very weird to type it out, as we only use it when we speak. At least, when I type, I tend to think in one language, so there's less organic mixing happening. Since it's speaking-only thing for me, I tend to add the usual filler words, suffixes for superlative forms and etc., but from all three languages. Apologies for using French/English, as that didn't convey my point properly.
It's not about loan words, it's exactly just mixing random languages together as they roll off your tongue. I know it sounds stupid, but I've been talking to my siblings that way my entire life, so it's very "natural" to me. Most of my university friends are at least bilingual as well, and from what I've been told, they do the same with their friends in their own languages.
I don’t know for cockney but verlan is very alive.
I’ve been in mixed language communities in which I wasn’t sure who spoke what, and I have found this to be quite effective when done right.
Good time to reference st:ng “darmok” episode and quotes like “darmok and jalad at tanagra”.
I have a pair and have only asked it that so far...
Kind of like having an expert next to you all the time.
So no matter what, conversations in your native speech would have to be delayed before translation.
I'm reminded of Mark Twain complaining about verbs arriving at the very end of sentencess in German (among a myriad of other complaints)
"The Awful German Language* -Mark Twain https://faculty.georgetown.edu/jod/texts/twain.german.html
Pretty sure there is nothing universal there though as you say.
I did several experiments recording from all the microphones I could on my iPhone and AirPods while out in the wild. My conclusion: it's impossible right now for that hardware given the microphones we have and what they pick up.
So much of what's spoken is at a combination of (a) high distance (b) low volume (c) background obscuration. Something that was clear as day to my ears would barely register on the mics. While context is of course an issue, the raw audio didn't have enough to even translate.
The one caveat is that there might be low-level (i.e., Apple-only) access to headphone microphones that capture the environment to do noise cancellation. I'm not sure though---I couldn't find them on any API.
For cases where you do have clear audio, existing apps (e.g., Google Translate) are so close to achieving this, but don't let you specify audio outputs with enough fine grained control. By default, it will start screaming out of your phone what you were attempting to silently translate.
That is, they are able to translate (in all directions) novel languages that were not previously heard[0]. It is an open question, with likely a negative answers, that there is a universal grammar even among humans[1] (the definition itself is vague but even the most abstract version is suspect and highly likely to not be universal across species). I think no one will be surprised if it is always impossible to interpret an entire language based on only a few words (let alone do it in real time)
This isn't a knock down, because even a trained device is insanely useful, it's just a note about limitations and triage. This is awesome stuff and I can't wait for the day we have transnational headphones. It's an incredibly complex problem that I'm sure is not short of surprises.
[0] There are a few exceptions such as Star Trek TNG's episode Darmok, S5E2, where the Tamarians' language is unable to be translated due to its reliance on cultural references (the literal words are translated but the semantic meanings are not). It's a well known episode and if you hear anyone saying "Shaka, when the walls fell" (translates to "Failure") they are referencing this episode (often not using the language accurately but who cares (nerds. The answer is nerds)).
> The Babel fish is small, yellow and leech-like, and probably the oddest thing in the Universe. It feeds on brainwave energy received not from its own carrier but from those around it. It absorbs all unconscious mental frequencies from this brainwave energy to nourish itself with. It then excretes into the mind of its carrier a telepathic matrix formed by combining the conscious thought frequencies with the nerve signals picked up from the speech centres of the brain which has supplied them. The practical upshot of all this is that if you stick a Babel fish in your ear you can instantly understand anything said to you in any form of language. The speech patterns you actually hear decode the brainwave matrix which has been fed into your mind by your Babel fish.
“The argument goes something like this: ‘I refuse to prove that I exist,’ says God, ‘for proof denies faith, and without faith I am nothing.’
“‘But,’ says Man, ‘the Babel fish is a dead giveaway, isn’t it? It could not have evolved by chance. It proves you exist, and so therefore, by your own arguments, you don’t. QED.’
“‘Oh dear,’ says God, ‘I hadn’t thought of that,’ and promptly vanishes in a puff of logic.
“‘Oh, that was easy,’ says Man, and for an encore goes on to prove that black is white and gets himself killed on the next zebra crossing.
“Most leading theologians claim that this argument is a load of dingo’s kidneys, but that didn’t stop Oolon Colluphid making a small fortune when he used it as the central theme of his best-selling book, Well That about Wraps It Up for God.
“Meanwhile, the poor Babel fish, by effectively removing all barriers to communication between different races and cultures, has caused more and bloodier wars than anything else in the history of creation.”
That said, given the Heart of Gold improbability drive, I don't think information theoretic violations are your biggest problems.
"Can you stand up?" would be translated differently into Japanese depending on whether you're implying you need them to move their butt off your cell phone versus directly inquiring as to the function of their legs after a car accident. If you speak English and hear it as a background without the rest of the context being picked up, your brain instinctively knows it can interpret it either way, no problem.
But if you're Japanese and the AI picks a specific way to translate it, then you are completely unaware of the ambiguity because the AI resolved it with a 50% chance of being wrong.
nitpicky, but is it though? not really. and it's as much 'difference depending on what you're implying' as there would be in english comparing just saying 'can you stand up' or specifying 'from the seat/at all'.
Google: "A bit sticky, things are pretty sticky down there."
I think it's totally doable but you'd need many more microphones in order to deal with real world noise. As MEMS microphone quality improves, this should eventually be possible with a combination of smartphone/headphone/some other device like something around your neck.
That's a nice way of saying unemployed.
Everyone gets a personal tutor for hours a day.
I would absolutely love a VR game where I just need to work in China or Mexico all day and pick up the language that way.
I think a lot of people are working on similar things right now. I know of one called http://yourteacher.ai
I don't have the expertise to judge the quality of Mandarin pronunciation myself, being a beginner. But it sounds OK in English and it's made by native Mandarin speakers in China so I expect that it sounds better in Mandarin than English.
> coming soon: "We've been getting a lot of these tourists here lately, they're eerily fluent, but all seem to have the same minor speech impediment"
Haha, if that were to pass, that would still be a far better outcome than our current situation of completely blind machine translation (this is especially for various Asian languages that are very sensitive to phrasing) and mispronunciation by non-native speakers.
Ah, that is called an accent.
This would be something quite novel as the speech irregularities would not have their origin in people
I don't know what you would call it but it needs at least some adjective before accent to differentiate it IMO
i’m hoping a voice as realistic as this becomes a local app soon, but i’ve not found anything that’s nearly as natural sounding yet. (also, honorable mention to chatgpt’s “sky.” she pronounces mandarin with a funnily american accent, but it sounds natural and not as robotic as the open-source alternatives i’ve tried)
You end up with the Duolingo problem where you know to say the names of 20 different fruits but not how to introduce yourself.
Not sure if this is a duolingo problem. There are of modules in duolingo specifically for saying your name. I think its the travel module.
The parents suggestion is that if we don't have to learn languages that will lead to us all laying down drinking big gulps while robot slaves take care of us. Their take is the extreme example. People have literally made this same suggestion about every technological advance and it never comes true.
Obligatory reminder that the movie itself explains that people are what they are not because of their lifestyle, but because of the time spent in low-gravity environment.
But there's a point in language learning where you can come to express yourself directly in a new language without intermediary "thinking" in your first tongue. The communicative and expressive potential of that mode is much higher than trying to squeeze one's intent through any kind of translation, machine or internal.
Plus, you know, it's fun.
However, translation has a great deal of subjectivity embedded in it, particularly when there aren’t 1:1 translations. Case-in-point: there are many English translations of the Christian bible, all similar enough, but there are enormous variations in some cases. And there are at least as many branches of Christianity as there are English translations of the Bible. Some of them strictly recommend the same translation, and they still disagree on the meaning of various passages.
Besides the problems inherent to translation, learning another language gives you another paradigm of thinking. The words we use, the way we construct sentences, etc., all impact our view of the world. Here’s a paper that discusses the impact of the over-reliance on English in cognitive sciences, and how this has downstream effects: https://www.sciencedirect.com/science/article/pii/S136466132...
Learning languages as an adult also has protective benefits. It reduces the probability of Alzheimer’s (maybe dementia, overall?).
I could integrate this instead of Polly pretty easily.
https://fluent-forever.com/product/fluent-forever-pronunciat...
I resisted using Duolingo because I knew their speaking feature sucks. But the only reason I need an app rather than books or audio tapes is that I need something to correct my prononciation.
https://github.com/adrianmfi/gpt-tutor
I look forward to testing if switching to Seamless can improve it further, Seamless supporting nearly 100 languages is a nice improvement.
Yes! Better yet, you're a spy, or a hostage negotiator, or the leader of any kind of enterprise (army, business, aid organization) ...
Programming games like that will resemble directing improv theater. You can't program every response; you'll have to instead fit each character with beliefs and motivations.
I can hardly wait.
What would be really cool is something that can autodub videos or audio into your target language. The hardest problem learning languages that aren't English is often finding content to consume in them.
Disclaimer : I am Krashenist so this take is biased
I'd love to learn French and the game would take place in locations all around modern France.
It would have to a good story. Maybe something in the style of Professor Layton series could be interesting, or something more open world.
He ended up rolling his own solution by standing up Whisper in one of our clusters and writing a basic front end and API to take his laptop’s mic input and chunk it every few seconds to send to the model and get back text in pseudo-realtime. We got him a pretty beefy Alienware so he wouldn’t be tied to the cluster GPUs. I can’t wait to see what he does with these new models!
We need more support for employees like this!
Also, what about Apple’s latest M3 series chips? Are this in the same realm as Alienware in terms of AI compute?
So gracious, to give a software developer some hardware to run the software they need to work, that costs a whopping nothing more than what other people in the industry get on the average.
>and let them roll an accessibility solution
"You're such a good employer! You let your employee build their own accessibility ramp to the back entrance in their own time, and even got them a mortar spatula to do so!" We need more support for employees like this!
>We need more support for employees like this!
And less support for employers like this.
It reminds me of the Homer Simpson quote, "I don’t mind being called a liar when I’m lying, or about to lie, or just finished lying, but NOT WHEN I’M TELLING THE TRUTH!" I would be equally critical if it was warranted, but when it isn't it's deeply unfair to the accused.
If the person wanted to build their own ramp, and the employer let them do it on the clock, that's a completely different scenario than the employee having to come in during their off-hours to build the ramp just so they can go to work.
So you're saying, not only you didn't pay him extra, but that the company got to benefit from him building the system as it was already in line with your other projects?
Unless his working hours were reduced, he did it in his own time.
Unless his pay was increased, he did it in his own time.
Unless the expectations for him were scaled down, he did it in his own time.
Merely allowing them to hack on a project that is in line of his work is exactly like having someone build their own ramp because "it's already in line with other construction they did on the project".
Nowhere did I see that GP said the employee was paid extra, had the expectations on other projects reduced in writing and deadlines shifted, or had his working hours reduced at same pay.
Saying "oh hey, you can work on this during 9-to-5 as long as you get your other shit done on time" means the project was done in his own time.
Because we are on Hackernews, where everyone likes to think themselves a scrappy startup owner, and not a person with a disability who might need accommodations from one.
As someone who’s profoundly deaf myself, another less technical approach is to install Rogue Amoeba’s Loopback, and use it to pipe audio from a given app into a tool like Google Meet or Otter.ai using the Loopback device as the audio source. This effectively provides real time captions for anything running on your existing machine.
The nice thing about Chrome feature is you can move the caption box around and keep it in the foreground while doing other things, although styling options seem limited (the text might be a little small for some).
[1] on desktop, not sure about mobile
[2] via chrome://settings/accessibility -> Live Caption
The extent of the effort being getting their employee a slightly-more-expensive-than-average tool that would enable them to do their job better regardless of the disability?
Such inclusive, much pat-yourself-on-the-back, wow.
"We gave our woodworking shop employee a quality saw so that they'd make their own accessibility ramps!"
Why instead? They didn't a bad thing, just nothing beyond what the law requires.
A simple answer for you, BTW:
Pay a student $15/hr to transcribe all meetings in real time for the employee.
Or pay money to professionals in the field to set up a solution. You know, people who get paid to do just that.
Jesus.
That is literally grounds for a lawsuit.
>so yes, making an effort to support an employee’s disability and their needs is worth recognizing.
...if the effort were actually made. Which is not the case.
Merely being compliant with the law is a very low bar, and we should recognize that they did not do anything beyond what the ADA requires them to.
But hearing loss does not impair standing up servers and software. They can pay the employee who probably is the expert at this, the guy with the hearing loss, or go task Emil to go do it to ... avoid 'appearances'?
Aaaaaaaaand who said they paid him?
Not the OP, for sure. It seems like the disabled person set the system up in their own time to help the company communicate with him in meetings.
>or go task Emil to go do it to ... avoid 'appearances'?
They could hire Emil to live-transcribe the speech into text.
They could hire the services of a professional company that would set up speech recognition for their meetings.
They could have done a lot of things.
What they did is nothing. Gee, a software company gave their employee a laptop, how totally unheard of!
Why are people here so willing to pass this off as some act of charity?
Additionally, it supports multiple languages (only one at a time sadly), so I also use it for Japanese captions and it's equally great there.
>That's very nice of you
...doesn't compute.
What exactly was nice here?
Probably this.
An "Alienware" is not a Ferrari, those things cost about the same as any other professional-level workstations.
There you go. https://github.com/dictation-toolbox/dragonfly
Simply voice to text is not what's needed for dictating commands. Unless I can load commands of on the fly and decode utterances that may be useful.
The client would need to be able to send its commands to the server on the fly.
- Whisper processes 30 second audio chunks. So if you process 5 seconds of audio you have to pad it out with 25 seconds of silence. Hence a loss of efficiency with wasted CPU / GPU cycles on 25 seconds per chunk in the case above.
- Whisper most likely can't handle hundreds of commands much less than a thousand performantly.
- Whisper doesn't handle short commands very well with a degree of accuracy post processing commands from free dictation utterances.
Command dictation should be weighted higher than general dictation when decoding.
I work with a little under 1500 of commands dragon naturally speaking. DNS is hot garbage as a program despite it has the best accuracy to date with the feature of commands and dictation in one utterance. You get to pay $750 for the privilege m
I've yet to see a free and open source speech recognition engine that can handle both dictation and commands with a high degree of accuracy.
Please please let me know if there's alternatives out there. I would definitely pay to support an open source project like this that focuses on command and dictation.
Most solutions out there that are open source nowadays focus so much on iot command recognition with intents. That's not well suited for controlling your computer with grammars containing voice commands.
> Input audio is split into 30-second chunks, converted into a log-Mel spectrogram, and then passed into an encoder.
Computer: webcaptioner.com Android: Live Transcribe (g.co/livetranscribe) iOS: Live Caption with the 'mic' icon enabled.
Web conferencing: Meet, Zoom, Teams all support realtime CC, which is pretty good.
I will note though that I feel safer getting an occasional bad word than I do having a translator straight up deceive me.
For example, "what the fuck" in English->Spanish is giving "qué diablos" output. Definitely toning down the meaning there.
If someone says something mean to me, I want to know it.
Would be interesting to see some work stress-testing the ability to convey ill-intent across multiple languages. Accurately conveying ill-intent is safety-critical for the person being threatened.
I'm the most excited for an open source one though, and it would be incredible if this could become it. I do 95% of my compute on desktop linux and it sucks being behind.
I told her then that the industry would be disrupted by AI before she retired.
Glad she pivoted. Really impressive results.
a few more years of improvements if they happen could be disruptive
Typesetting. Music engraving. Bookbinding. The quality of all these fields have been materially harmed by advancements.
Computer typesetting has, by and large, been a significant regression, though the gap has largely been made up now if you make the right choices.
Published music scores used to be set by experts. Now they’re set by novices using software that is mechanical in method and generally quite insipid. Most are atrocious compared to the old masters, and mediocre at best compared to the typical published scores from a hundred years ago; and very few popular scores are really good (… and if they are, there’s a reasonably high chance they’ve used GNU LilyPond, which has focused on this problem). But the barrier for entry is so much lower, and people have got used to the inferior results, so I don’t know if anyone engraves music the old way, and even people that know better largely just shrug and make do with the new. Like with computer typesetting, there is hope because things have slowly improved. But most will continue to be mediocre.
Books used to be bound with cold glue. It takes time to set, but the results are very good, supple and long-lasting. Then along came hot-melt glue, and it’s just so much friendlier for cheap manufacturing because books are finished within a few minutes instead of a day or two, that I don’t think anyone produces books the old way any more, even though the results are abysmal in comparison (compare the binding and reading experience of a paperback from the ’40s or ’50s with one from the turn of the century; no one after tasting the old will desire the new; for he says, the old is good). But they’re just (barely) good enough. Unlike the other two, I don’t think there’s any hope here—the regressive advancement crowded out the superior but dearer option so that no place was found for it.
What seems to universally happen is, the market bifurcates - one part is in a race to the bottom, the other (much smaller) aims for super premium tier (overpriced quality), because only those two positions are sustainable, once the race-to-the-bottom side drags all the economies of scale with it. So as a consumer, you get to chose between cheap low-quality garbage that's barely fit for purpose, and rare, super-expensive, professional/elite high-end products. There is no option for "good value for reasonable price".
This has been happening to everything - software, furniture, construction, electronics, vehicles, food, you name it.
Sure the final result sounds slighty robotic. 99% of people wouldn't care, and you can get more training videos done, faster for a fraction of the cost.
[Edit] And I'll add the difference from 6 months ago is noticeable to today. I imagine every 6 months we can just re-download updated voiceovers and every 6 months will sound just slightly more polished..
Yes. I just discovered there is a text-to-speech addon [1] (now a few months old) for World of Warcraft that adds voices for every NPC in the game... It is so impressive and game changer (pun intended) that I naively asked in the chat of the Twitch stream I was watching "when did Blizzard add voices to the NPCs??". For an instant I really thought Blizzard contracted actors, but no, someone like you and me just used AI to generate realistic voices for every character in the game. I don't think it's ready yet to completely replace actors in video games (surely it will in the near future tho) but voice acting is something so expensive to do that I can see studios and developers in 2024 already use this tech for all the optional dialogues and secondary characters' voices.
(quote from the video presentation in the link)
I thought it was meant to be "toxic words, hallucinations, etc" in the script
Just text to speech has gone too far. Audio books would be mainly generated on the fly like this?
I think some RPGs in some 5 years time might have something like this:
- A text file that outlines characters and a lose plot/Story line. Human written.
- 3D Mesh Generation based on character description via Transformers based models. Auto generated.
- Dialogues for each NPC via LLM.
- This TTS engine again based on such models.
Result - almost unlimited replayability. Or even edit text file, have a new world based on a new story line with characters having different personas.
It's not just where the verb is; sometimes I say something ambiguous, and my next utterance is supposed to acknowledge and remedy that. But if that ambiguity doesn't exist in the target language, I don't see how a simultaneous translator can convey the ambiguity, without knowing how the next utterance is going to refer to it.
Maybe that's why human simultaneous translators often seem to stumble or backtrack. I've never met someone whose job was simultaneous translation. It must be very difficult.
I'm impressed by this effort to convey non-linguistic elements of speech in translation. It's quite an achievement, and a very ambitious goal.
Aside: I wish I knew how speakers of tonal Chinese dialects express feeling, when tonality is supposed to convey semantics. When I hear chinese speakers, I can "hear" the feeling, but I don't know how they do it - it can't just be down to emphasis. (I learned some mandarin 50 years ago, at school. I learned the tones, but they didn't teach expression; and I was never taught by a native speaker, although there were language-lab tapes.)
It would be nice if they could layout what exactly is missing in terms of data to make a language work better, while the actual AI bit is out of reach for most of us maybe we could provide more data.
There is also a 60 sec limit and wonder if this is HuggingFace limitation or Seamless?
If you want to contribute by recording yourself speaking Swahili, https://commonvoice.mozilla.org/sw is the place to go. Although Meta has access to much larger data sets, they nonetheless use Common Voice as a "known good" source. E.g. the paper on their SONAR speech encoder reports experiments on Common Voice data, coincidentally involving Swahili https://ai.meta.com/research/publications/sonar-sentence-lev...
As a student in the mid-90s I worked on a system called Verbmobil at the German Research Center for AI and it did speech-to-speech for English, German and Japanese in very limited domain.
This was done via "classical" NLP: You had to model the domain with concepts, you needed sentence parsers, semantic engines, speech-to-text hand-crafted for 3 languages etc.
As it turns out, this approach is/was a dead-end.
I don't think this would fool anyone that I was a real native speaker of the target language, but for casual conversation this would work pretty much perfectly. It basically avoids all of the traditional pitfalls of machine translation, like the unnatural robotic voice that it outputs, the slow translation speed and huge latency for realtime conversation, and the loss of emotion.
Yeah, I kinda agree with the spirit of your comment; it sure would be nice to see a major Indian language like Telugu on their landing page for sure. But that's just my Indian-person bias speaking.
is there an example notebook that shows how we can use this model with own sample audio and text? thanks!
This may be apocryphal but I’ve heard that in formal settings (e.g. UN) they won’t translate it and will instead give instruction on when to laugh.
That said, I suspect that real-time language translation is always going to be somewhat imperfect due to its nature. Non-real-time translation of literature is still a subjective art form even at the very high-end of human expertise.
It's all a trade off.
Either way it's extremely exciting that we get to even discuss this stuff as real possiblities.
I’ve been learning german since 8 years, and the amount of expressions and different ways to say things around the country is impressive. There’ll be a “interpretative” real-time translation, but it won’t guarantee fully understanding in so many cases, maybe ever.
Other thing, and we have this in common with all languages, is the context and this is difficult to address i believe.
Nevertheless, it’s impressive how far we’ve reached and i acknowledge the usability of these tools. However, human knowledge will be always crucial and primordial if we want to guarantee full understanding.
"Since", as used here, would lead me to guess you are not a native English speaker?
So it can't be trusted at all then
Also: "This research demo is not open to residents of, or those accessing the demo from, the States of Illinois or Texas"
It's not dissimilar to some kind of a "ch'ti" / "chtimi" accent or a belgian-french accent (which is not dissimilar to the french ch'ti accent, heard in some part of the north of France. "Ne partez pooooo" (with a longer "a" which sounds nearly like an 'o': that's not proper french at all) instead of "Ne partez pas".
That's said I'll take the non-expressive accent any day over subtitles for when watching video in a language I don't understand: it's clearly good enough.
Besides the ACCEPTABLE_USE_POLICY, there's a CC BY-NC 4.0 (NonCommercial) license, a 'SEAMLESS_LICENSE' (NonCommercial), but also an MIT license? It would seem these other licenses contradict the MIT license, could somebody help clarify how these all interact in practice?
https://github.com/facebookresearch/seamless_communication#l...
https://seamless.metademolab.com/expressive/?utm_source=meta...
Interesting mix.
In Texas it seems to be part of AG Paxton's culture war stuff. https://www.texastribune.org/2022/05/12/texas-face-filters-i...
> According to the Settlement Administrator, payments to class members between $200 to $400 started going in the mail May 9.
I got a $0.19 check from an iTunes settlement once, but this wasn't one of those cases.
Facebook has had to pay out hundreds of millions of dollars in settlements for related class-action lawsuits, and rather than trying to get informed consent, they’re deciding not to collect biometrics from residents of those states.
It runs but any audio input (you will need to provide wav not mp3's) I tried (tried 20s/40s/300s) I get just one short sentence returned in target language that seems not related at all to my audio input (i.e. Tous les humains sont créés égaux).
Seems like some default text but it runs on full GPU for 10 minutes. Tons of bug reports in GitHub as well.
Text Translate works but not sure what is the context length of the model. Seems short at first glance (haven't looked into it).
Oh and why is Whisper a dependency? Seems not need if FB has their own model?
> It runs but any audio input (you will need to provide wav not mp3's) I tried (tried 20s/40s/300s) I get just one short sentence returned in target language that seems not related at all to my audio input (i.e. Tous les humains sont créés égaux).
You might want to open an issue on github for that one. The model is made to work on short utterances, if you have a long speech, you'll want to segment it first. I've tried "tous les humains sont créés égaux" on the demo: https://seamless.metademolab.com/expressive (which runs the same code as in the repo) and the output was correct. Maybe there is something wrong going on in the conversion of the input audio?
> Oh and why is Whisper a dependency? Seems not need if FB has their own model?
Whisper is a dependency as it's used as a baseline for evaluation. You can check out the paper for explanations.
Edit: I can’t find the weights but if I’m reading the paper right anyone could train their own detector.
Yes, we chose not to release the watermark detector to safeguard against adversarial attacks. This decision helps prevent any attempts to erase the watermark by malicious users.
The watermark generator and detector are trained together, one can use the information in our paper to train your own generator and detector model, however in this case the watermark signature created will be distinct from the one we use to protect our seamless translation models. This approach ensures each model maintains its unique security features.
My ultimate goal is to have realtime translations of video conferences. I've moved to a new country, and while I'm super privileged that most of my colleagues speak English, we still have a number of "all hands" meetings that I get lost in pretty easily.
Of course this is all before the new seamless v2 drop, so I'm hoping it will become easier! There's a lot of interesting things in this new seamless release, among them ggml inference and the streaming model. Poking around the code, there is even some whisper binding in there too, so it seems like a possible integration is already in use (I haven't had time to really dive in).
That would have been the knockout punch.
Especially because the head of AI at Meta is a French guy AFAIK (Yann Lecun).
They also mention it in one of the videos about the streaming variant of their translator. But I guess 2s delay or what they mention is close enough for practical purposes.
I feel like for personal relationships where true real-time is required, having a computer intermediary would be weird anyway and you have to learn the language, at least for the time being and as long as personal relationships are still relevant (in the post-AI world they might not be).
I guess it's possible that the AI learns about a specific person over time? That way it can be confident about what's being said as soon the person starts saying it
What I would love to see is an ability to add my own voice (yes, at the risk of deepfakes) so that the model could "speak" in any language and sound more like me, not some random voice actor it was trained on.
Would allow more interactions of people that don’t speak the same language
Edit: by the upvote I guess it wasn't just me?
Attribution-NonCommercial 4.0 International
https://github.com/facebookresearch/seamless_communication/b...
Thank you Meta!
"The Babel fish is small, yellow, leech-like, and probably the oddest thing in the Universe. It feeds on brainwave energy received not from its own carrier, but from those around it. It absorbs all unconscious mental frequencies from this brainwave energy to nourish itself with. It then excretes into the mind of its carrier a telepathic matrix formed by combining the conscious thought frequencies with nerve signals picked up from the speech centres of the brain which has supplied them. The practical upshot of all this is that if you stick a Babel fish in your ear you can instantly understand anything said to you in any form of language. The speech patterns you actually hear decode the brainwave matrix which has been fed into your mind by your Babel fish.
"Now it is such a bizarrely improbable coincidence that something so mind-bogglingly useful could have evolved purely by chance that some thinkers have chosen to see it as a final and clinching proof of the non-existence of God.
"The argument goes something like this: 'I refuse to prove that I exist,' says God, 'for proof denies faith, and without faith, I am nothing.' 'But, says Man, the Babel fish is a dead giveaway, isn't it? It could not have evolved by chance. It proves you exist, and, by your own arguments, you don't. QED.' 'Oh dear,' says God, 'I hadn't thought of that,' and vanishes in a puff of logic."If anything, the evidence is that it isn't true, see https://journals.plos.org/plosone/article?id=10.1371/journal...
Any apparent causality of age of acquisition seems to be a proxy of hours of exposure. It may well be that it is easier for young people to rack up a lot of exposure to a second language, but not much evidence that age plays much of a factor for people of different ages who had the same degree of exposure.
> we show that the ERP signal in response to grammatical violations depends on the AoA of an L2 learner, as well as on the regularity of the structure under investigation. In (lexically determined) syntactic constructions different from the L1, we found a gradual change in processing strategies that varies by AoA, with a native-like effect for early learners and a less efficient neural processing strategy for later starters.
Although they do clarify that these effects could be confounded with age of acquisition instead of it being the cause.
I'm not sure I want the latest twitter trend to be involved in the design of my translator...
There are more details in the paper if you want and the mitigation code is all open source if you want to check what it actually does.
Oh, well that clears it up! </snark>
I don't see any definition of 'toxicity' on the landing page - it seems to be one of those 'I know it when I (hear) it' kind of words... unless there's some widely-accepted definition in this area of study?
The tldr is that if you say: "Thank you for this job offer." you wouldn't want it to be (mis)translated as "Go F*k yourself.". But if you do say "Go F yourself", you still want it to be translated as that.
The Hitchhiker's Guide To The Galaxy claims the opposite:
"Meanwhile, the poor Babel fish, by effectively removing all barriers to communication between different races and cultures, has caused more and bloodier wars than anything else in the history of creation."
e.g. "geil" (either cool or horny depending on usage) in German
It's not fundamentally different than e.g. "wicked" in English, but the biggest bias that potentially all these ML models exhibit is predisposition towards Anglophoneism
You can find details on how the multi-language creation of the toxicity lists was done in section 7.3 of the NLLB paper: https://arxiv.org/pdf/2207.04672.pdf. TLDR: it's not just a translation of a base English list, even if we started from that, each language has a curated list that was built by professional translators.
Can it make sure that the output toxicity level is not lower than the input?
If not (which I strongly suspect is the case), then that is unacceptable. We cannot fight toxic narratives with ignorance.
Perhaps it's taking a Big List of Naughty Words and weighting them so that the system must be "extra sure" that's what the speaker said, or else fall back to a G-rated word?
1: https://www.google.com/search?q=engrish+fucking+sign&tbm=isc...
There is no moral superiority to deny or force label other people's identities. You're an attack helicopter? Great, roger dodger, let's go get coffee Seahawk.
No one is seriously asking for litter boxes in school bathrooms or helicopter refueling stations.
This feels a bit out-of-nowhere.
My read on parent comment was that "Twitter trends" are fast-changing norms about what language is (un)acceptable. They were not saying that LGBTQIA+ identity itself is a trend.
The original comment you replied to made the point that they don't want their own personal expression curtailed or modified according to someone else's opinion of acceptable speech.
As someone who repudiates Russia's policies, I support and agree with their point.
Taking what they wrote as harshly as possible, a translation model's output might include narrative elements from the transphobic judgement you are concerned about. That would be a problem, because it would amplify transphobic narratives.
Taking what they wrote as favorably as possible, a translation model's output might rephrase what was written, such that a pro-LGBTQIA+ inclusion narrative is more eloquently expressed than the author actually intended. That would be a problem, because hiding the reality of transphobic narratives would remove our ability to recognize and talk about them.
To make this even more complicated, what if we are using this model for real-time dialogue? What happens when someone says something vaguely transphobic, their words get translated to an inclusive narrative, and you continue that inclusive narrative in your reply? Should the translator alter your words to be transphobic? If it doesn't, then will the entire conversation go off the rails, or will both parties continue, oblivious of each others' ideological subtleties?
---
I don't believe for a second that a model could be trained to avoid toxic narrative and translate accurately.
Hallucination is a feature, not a limitation. The sooner "AI" narratives can accept this reality, the better.
From the hackernews guidelines
None of the videos shows any modified/lip-synced footage. There doesn't seem to be a reason for this thing to need access to my camera.
Also, using it with tape over the camera doesn't seem to work either. (Perhaps it needs to see facial expressions in order to work?)