15.ai
15.ai
15.ai
To clarify: I have not been with MIT in years. I was paid the minimum hourly rate (roughly $14 an hour) to work on a related project during my undergraduate years, which eventually evolved into this project years down the road. (In fact, I had to pay for my own compute to get my work started - MIT never offered me any credits.) Everything else has been paid out of pocket since then. Yes, it does indeed cost several thousands of dollars a month - that is not an exaggeration. This has been optimized so many times that the technology needs to improve first before I can cut down on costs.
The timer in the "TOS" was put there in hopes that people could understand where I was coming from in regards to misappropriating this kind of research. I did not expect people to get this riled up over a 10 second timer. (Especially not on Hacker News, of all places...)
Edit: I suppose it makes sense to include other information in this post so that others don't have to go hunting for my comments in this thread:
- >This is unrelated but what's with the fascination with HN users and My Little Pony? I've noticed this on a lot of posts in the past few months.
- Twilight Sparkle's voice is indispensable in getting emotional contextualizers to work properly. The logo and profile picture is an homage to that fact.
- >But seriously: how did you get the domain `15.ai`?
- I purchased it. It was definitely not cheap.
- >They named their tts "Deep Throat"? Why would you?
- It was a suggestion from a Twitter user, and I found it clever.
- >I heavily doubt it's "several thousands of dollars"...
- It is indeed several thousands of dollars a month. I can show you AWS invoices, if you're skeptical. Just send me an email and I'd be happy to show proof.
- >The disclaimer is a little ironic considering the site owner doesn’t own the model (MIT does) and doesn’t own the training data (the various shows and games do)
- I'm sorry to tell you that I do, in fact, own my own model. I have not been with MIT in years.
As an undergrad, I was completely broke. I figured that keeping the project free to use was the best thing I could possibly do with my research as I continued to work on it.
They used to sell them openly, but coin mining ruined that :(
I can afford it.
>Can't you continue/do your research without a public website?
Yes, but the website has multiple purposes. It serves as a proof of concept of a platform that allows anyone to create content, even if they can't hire someone to voice their projects.
It also demonstrates the progress of my research in a far more engaging manner - by being able to use the actual model, you can discover things about it that even I wasn't aware of (such as getting characters to make gasping noises or moans by placing commas in between certain phonemes).
It also doesn't let me get away with picking and choosing the best results and showing off only the ones that work (which I believe is a big problem endemic in ML today - it's disingenuous and misleading). Being able to interact with the model with no filter allows the user to judge exactly how good the current work is at face value.
I'm no stranger to passion projects, I have a lot of respect for people like you. This is great stuff.
So, excellent job.
Edit: also wanted to thank you for the Chell voice, it sounds completely true to life to me! (minus some jumping noises)
I don't feel comfortable publishing or releasing anything until I know for a fact that I can make no further improvements. It's not out of corporate greed or anything like that - I'm just really paranoid about getting out the best work possible.
The results are clearly synthetic and need work. However what's cool is that there are a ton of characters (from popular shows and video games) and there are useful statistics like inferred emotion (which is also in the output).
Honestly it's a big problem how a lot of AIs are like "black boxes" where you really can't customize or see anything. Yeah we have DALL-E and GPT which can generate text images but the lack of customization or fine-tuning the image afterwards severely hinders what's possible with them. Ultimately what you want is something interactive, where you can control how much or little the AI generates, and give it really specific criterion.
But seriously: how did you get the domain `15.ai`?
I purchased it. It was definitely not cheap.
it's an MIT project so I'm sure that was a factor
You can remove or add things etc.
And for GPT you can also specify more details.
Only a question of time until you can work with the ai on your art/thing.
There are ai models which keep track of context and others which generate a plan of actions.
AI is not a blackbox
Unless this is a recent change, their mission isn't that:
> OpenAI’s mission is to ensure that artificial general intelligence (AGI)—by which we mean highly autonomous systems that outperform humans at most economically valuable work—benefits all of humanity.
Blackbox would mean in a normal setting that no one knows how it works
If the weights were public, the community would figure out a way to fine tune it.
A friend of mine was going to build a writing assistant on top of GPT-3. She got a lot of encouragement from them. Then one day the social media storm hit OpenAI, and suddenly safety became a non-negotiable feature of their api. And along with that came the restriction “absolutely no products that can generate unlimited amounts of text, even if you can pay for the credits.”
Poof, no more business.
Imagine buying a keyboard that restricted what you could type.
All of these problems go away when you have access to the weights.
Have you used GPT-3 with any of the methods mentioned in the docs?
I’ve seen that GPT-3 can produce quite starkly different results when prompted differently and when samples have been uploaded.
The whole history behind the project is fascinating: 4chan had a huge role in its development, and the project's work was stolen by an NFT company that a famous voice actor endorsed not too long ago.
The truth is that, today, if I was going to use a tool to generate voices (say for YouTube), I wouldn't necessarily pick a small SaaS tool. I'd use Amazon Polly or some other GCP-style platform voice creation tool. There are already a few products in the space, and their costs are so low as to be almost negligible (example: Polly, 5 million characters free). For a commercial project, I could probably stay on a free tier for a whole year.
With Dall*E, it seems like the only option, and it's such a superior option that a website could abuse it for commercial profits. But for voice synthesis, it's already dirt cheap and commercially available without limitations.
15.ai seems to beat them on some sentences, but not all. Looking forward to the day when we can have real human-level quality of voices on-demand.
> OOPSIE WOOPSIE! Uwu. We made a fucky wucky! A wittle fucko boingo!
I was a bit taken aback because I'd never seen this meme before. So I looked it up and I found the origins of it pretty funny: https://knowyourmeme.com/memes/oopsie-woopsie
Also, hearing the voice of Spongebob saying that really brightened my day. Bravo, MIT.
[1] https://www.youtube.com/watch?v=Bb5eHXavk34
Edit: more links because this made me laugh way more than it should have https://youtu.be/zIUiOUGiB7s?t=6 https://youtu.be/0VovEGbYsI4?t=41
[1]https://www.urbandictionary.com/define.php?term=Rick%20Roll
I dunno. Feels a little gross to me. Eventually there is going to be a big copyright case about a model trained with copyrighted material. I have no idea how that will be resolved. Or maybe there will simply be new laws passed to make it either explicitly ok or explicitly not ok.
I highly suggest reading into the project first. The Wiki article I linked before (https://en.wikipedia.org/wiki/15.ai) answers all of your questions about copyright infringement.
In terms of copyright infringement, your wiki link answers nothing. A court ruled that Google could use copyrighted book text to train an algorithm to improve search results because the copying was highly transformative and did not serve as a market substitute to the original work.
Meanwhile 15.ai is using copyrighted voice recordings to train an algorithm to synthesize new voice recordings that sound like they came from the original speaker. This is radically different from the Google case. Just because one instance of using copyrighted material to train an algorithm qualifies as fair use does not mean that all use of copyright material to train any algorithm also qualifies as fair use.
There is absolutely nothing about this that is settled law. In the next 20 years there are going to be lots of lawsuits, lots of settlements, possibly a few rulings, and maybe even a few new laws. I find the whole topic very interesting. YMMV.
In many instances a policy of "ask for forgiveness rather than permission" can get you further, faster. While Nickelodeon are unlikely to grant you a license to the Spongebob voice because that has broader licensing and IP repercussions, they are likely to tolerate a research project using their characters (e.g. just as they have to-date tolerated The SpongeBob SquarePants Movie Rehydrated, which was a fan re-creation of one of their actual movies).
So, creepy thought: should we be recording audio of our parents, so we can still "hear from them" once in a while after they die? People are going to want to reconstruct their lost loved ones with AI. This project seems to imply you only need an hour or so of audio.
https://www.wired.com/story/a-sons-race-to-give-his-dying-fa...
Wouldn't do it for me. My mother is an artist and makes the most outlandish connections between seemingly unrelated topics on a regular basis.
I think this is kinda edgy humor that can work in small, select groups.
But maybe in most contexts in which people will be looking at AI method tech demos (e.g., within a company, or researchers at a university), we're still feeling ongoing effects of multi-generational injustices. In such a context, no one wants to be associating to ideas of women around some infamous '70s porn film. Doing that, and making light of it, seems like it'd rightly bother a lot of people.
When you're focused on a project, and maybe first discussing/showing in one small group, it can be easy to forget there are many additional things going on outside that group, some of which we also want to consider. I've made that mistake multiple times, including with the wrong humor for a context, still cringe when I think of instances of it, and this looks like that to me.
Maybe the developer will see this discussion, and decide to change some things, ASAP. It might still be relatively easy to change. (Maybe doable before Monday business hours, in the time zone of the university mentioned prominently.)
- Deep, because it's deep-learning based
- Cream, referring to its pornographic application
- Py, as it's written in Python
- DeepCream is close to DeepDream
- CreamPy is homophonic to creampie
- As a whole, homophonous to deep creampie, a meaningful pairing of noun and adjective
Shut up and take my money.
It’s really strange reading these ignorant comments from HN…
> The DeepThroat model is able to generate voices of varying degrees of emotion despite never having been exposed to emotive data of the character during training. Furthermore, multiple characters can be trained simultaneously, significantly reducing the amount of time required compared to if one were to train the character models individually.
Can't wait to read the paper.
Brings back fond memories of /prog/ and all its nonsense. Praise The Sussman!
One thing I noticed experimenting with it, questions don't have the typical rise in pitch at the end of a question in English.
On desktop, maybe I’d open dev tools and remove it. On mobile, I won’t be bothered. I hate that this is what the web has become and I choose to simply miss out on websites that behave this way.
document.querySelector('.vm--modal').remove();document.querySelector('.vm--container').remove();
Or add it to some default script that is applied to all pages, that way when you visit 15.ai you never even see the TOS box in the first place.But also, 10 seconds is extremely annoying.
Aren't we all appropriating the work of Newton, Maxwell, Einstein, and others? It's not like Maxwell's equations are copyrighted.
You're entitled to your opinions, but as a PhD myself I'd rather my research get used by people than end up in a copyright junkyard of things people can't use.
Does anybody else find the irony in this statement absolutely amazing lol.
I don’t know why, but I honestly expected more from HN.
But it's not just about the popup - it 's more that when your work is fundamentally about using reusing someone else's character, it feels pretty hypocritical to be so focused on making sure you get credit.
The problem is that after having to wait for 10 seconds to reject their terms of service (which you should be able to reject right away) before even being able to see what the site is about, they are rickrolling you, effectively giving you the finger for not wanting to agree to their terms without context. That‘s quite unprofessional, counterproductive and antagonistic.
Shame to see the toxicity over a passion project, whos creator generously went out of his way to answer the questions and ridiculous comments.
"...MIT owns inventions made or created by MIT faculty, students, staff, and others participating in sponsored research projects or in MIT programs using significant MIT funds or facilities or those inventions developed pursuant to a written agreement with MIT..."
I got RickRolled as soon as arriving to the page. :-)
In other words, if I deep fake someone's photo on someone else's body, I own the rights of that 'model'?
> Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the "Software"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so
His TOS is literally the antithesis to the free nature of the MIT license :)
To clarify: I have not been with MIT in years. I was paid the minimum hourly rate (roughly $14 an hour) to work on a related project during my undergraduate years, which eventually evolved into this project years down the road. (In fact, I had to pay for my own compute to get my work started - MIT never offered me any credits.)
And to address the philosophy behind the MIT license (also copy pasted from another comment):
For the past three years, I have done nothing but work on this project nonstop. I've been working on massive improvements (that some have pointed out in this thread) that I've been stuck on for the past several months, but I'm getting close to finishing that up.
I don't feel comfortable publishing or releasing anything until I know for a fact that I can make no further improvements. It's not out of corporate greed or anything like that - I'm just really paranoid about getting out the best work possible.
https://pjreddie.com/static/Redmon%20Resume.pdf
And in case you were wondering what this little pony did next...
https://scholar.google.com/citations?user=TDk_NfkAAAAJ&hl=en
AHHHHHHHHH
In the greater sense, though? Ponies have always been this weird relic of internet absurdity and bear-baiting. Some people rep it ironically, other people are dead-serious, but the community has significant overlap with the STEM field. As a result, a lot of pony-related stuff would end up propagating into the tech world, much like this very project.
The extremely dedicated brony subculture voluntarily put in a lot of work to get a corpus for the AI to learn from.
There's also another factor at play: this AI works best with highly pitched voices, which my little pony is just full of. Not only did MLP provide such a generous source of training data, its results were also much more impressive than the dry dictation many other corpi would've resulted in, adding to its fame.
I personally haven't seen any significant rise in MLP references, though that could be because I don't know the show so I don't catch references to it. It's also very possible that you've caught the Baader-Meinhof phenomenon.
Weeaboo/furry data scientists are always ahead of the industry - I seem to recall an effective decensoring model that was called "DeepCreamPy" and had almost 10K github stars before it was nuked and rehosted.
I'm convinced that learning Statistics is in a zero-sum game with social skills.
And it's not really specific to HN. For instance you have well-known people in the community who do vaccine R&D, or cryptography, or contribute to the C/C++ standards at ISO, or several other STEM things that are pretty outspoken about their interests.
This is made more obvious on Twitter, where people tend to blur their personal and work identities a lot.