DeepMind introduces ‘EATS’: adversarial, end-to-end approach to text-to-speech
syncedreview.com
syncedreview.com
EATS 4.083
WaveNet 4.41
Tacotron 2 4.53
Human Speech 4.55
It seems like a stretch to me to say it's "a comparable to the state-of-the-art models".[1] https://google.github.io/tacotron/publications/tacotron2/ind...
Clarification: DeepMind never fully 'beat' SC2 pros the way they did in Go.
Their AI was super impressive, certainly far beyond any other that has been developed. But they only went so far as consistently doing 'okay' against top tier players on the ladder, and that with some significant caveats to their performance. In particular, it's debatable to the extent that they're overpowering some players through sheer APM, rather than superior tactics or strategy. Yes, this is true even after they put in certain restrictions.
And given their recent silence on the issue, it looks like they may have given up, which would be unfortunate.
The implication here is that the AI can select units and buildings arbitrarily much faster than humans can, and doesn't need to spend APM on creating and maintaining control groups themselves.
That was a version that was trained for shorter time, though.
I'd say AI won mostly because of better micro and less mistakes, but I haven't seen the most recent games, maybe that changed?
Blink stalkers are very cost-effective units that are balanced by being very hard to micromanage perfectly. It's basically impossible over large area (during battles that take more than 1 screen) because it's pseudorandom which of your stalkers is targetted and needs to blink in any given moment, and you need to move the screen to that place to blink the stalker in time. AI knows health of all units without the need to move the screen to see it so it knows where to move the camera in advance.
AI perfectly micromanaged full limit army of blink stalkers during a battle that was 3-screens big. No human could physically do it with the current starcraft UI.
It's like chess pieces were 100 kg each and AI was steering an industrial robot :) Interesting, but it's mostly measuring stuff that's not about AI :)
In the next game MaNa went into immortals (units that specifically exist to counter stalkers in pvp), and still lost, because AI blink stalker micro was just so impossibly good.
To be fair to AlphaStar, figuring out that it's a good idea to do this is not trivial. You could say it learned a kind of exploit (of the artificial restrictions placed on it).
Yes, the APM, in terms of raw numbers, was constrained. But the computer used its APM in a completely different way, one which players could not possibly replicate.
One obvious example: the AI doesn't use control groups. To anyone who's played Starcraft in a somewhat competitive fashion, that's a HUGE tell. Because the implication there, is that the AI doesn't need to make control groups, because it can just arbitrarily select whatever subset of units it wants to control, whenever it wants. Humans can't do that.
IIRC, the AI also tended to 'burst' its APM in the middle of fights. That's the opposite as what happens with human players. Human players often have higher APM outside of fights, when they're doing repetitive, simplistic actions (e.g. Zerg players slamming the "make lings" or "make drones" hotkeys on all hatches simultaneously) and/or spamming to keep up a rhythm. When they go into a fight, where they don't want to waste time on spamming and where commands to improve their position are less obvious, they actually tend to slow down.
"MaNa: I would say that clearly the best aspect of its game is the unit control. In all of the games when we had a similar unit count, AlphaStar came victorious. The worst aspect from the few games that we were able to play was its stubbornness to tech up. It was so convinced to win with basic units that it barely made anything else and eventually in the exhibition match that did not work out. There weren’t many crucial decision making moments so I would say its mechanics were the reason for victory."
I wonder if DeepMind will tackle emoting; they certainly have something good enough for voice overs in a business context.
https://paperswithcode.com/paper/end-to-end-adversarial-text...
The whole industry benefits from this sharing of knowledge.
Also, the best aren't going to work where they can't publish.
1) Everyone benefits greatly from collaboration across companies and researchers at univiersities 2) There is a lot of value for google (and everyone else) as a company to show their achievements for many reasons, such as attacting talent, partners and so on. 3) And lastly, they don't expect the secrecy of their research to accomplish much, i.e. someone might re-discover similar results not that far into the future
I'm wondering why MacOS which used to have a superior set of TTS voices didn't implement any of the new neural voice engines. It's been years since Tacotron and WaveNet, and still the same crappy system voices.