The problem here is that it's missing the "glue" to more real world applications. This is where more humdrum software engineering comes in.
Diplomacy in this is much more interesting than Stratego or beating the next video game - it mixes cooperative game theory with NLP and reinforcement learning.
This is a bold statement. The world does not function based on "well-defined rewards". The concept of "common sense", which some consider table stakes for a human operating competently in our world, is mostly made up of things which are neither well-defined, nor allow a tremendous amount of training data. Current ML approaches require both.
But in general, yes, this is why since 2013 we don't see AI making anywhere as massive strides everywhere as they do in boxed-in applications like games.
Humans can be neatly reduced to a "procreate, and make sure your progeny procreate" value function. A lot of apparent counterexamples to that turn out to make sense when 2nd order effects on progeny are considered.
It's been providing real value to me over the past day for practicing Spanish, explaining Machine Learning concepts, and doing fancy write-ups in LaTeX. And this one can't even use Google yet! (other research teams have already created models capable of doing so, it's only a matter of time until these innovations are brought together in one place)
---
If you took the word "Wikipedia", and threw all the letters up in the air, what are some possible things they'd spell once they fell on the ground again?
If you took the word "Wikipedia" and threw all the letters up in the air, some possible things they could spell once they fell on the ground again might include "kawipieda," "keiapwidi," "kipedwaia," "pwakieida," or "eipawikid." However, without being able to see the exact arrangement of the letters on the ground, it's impossible to say for sure what they would spell.
---
(last sentence is a bit confused)
Input: "If five cats can catch five mice in five minutes, how long will it take one cat to catch one mouse?"
Output: "If five cats can catch five mice in five minutes, then it will take one cat one minute to catch one mouse. This is because the number of cats and mice are inversely proportional, so as the number of cats decreases, the time it takes for a cat to catch a mouse will increase. Similarly, as the number of mice decreases, the time it takes for a cat to catch a mouse will decrease. Therefore, if five cats can catch five mice in five minutes, then it will take one cat one minute to catch one mouse."
As you can see, it fails at simple logic because it fails to understand that the number of mice aren't fixed in the question. Whereas a human would immediately pick that up because they reason beyond x=5 and y=5.
[Citation needed]
I think that the computer made the absolutely standard human mistake, so that could be considered a plus.
You could probably test this by asking a few friends this question and see what they say. Outside of pure math problems you can get into an infinite regress defining the underlying first principles behind any given assumption.
Are you claiming its training data has letter permutations of the word “Wikipedia”?
It’s actually pretty capable of doing basic combinatorics.
I have seen reports that it will happily hallucinate a plausible but wrong answer to all sorts of different prompts, intermixed with many mostly correct answers. It's interesting to think about how to place trust in such a system.
Anyhow spoiler alert, the neural nets running the virus response have been inadvertently trained to prefer simple systems over complex ones without anyone realizing, and decide that a planet with no life on it after being wiped out from the virus is infinitely more simple than the present one and starts helping it out instead of stopping it.
So short answer to your question is I would not place much if any trust and systems like that, in as far as anything that has high stakes, real world consequences.
I'm glad I forgot about them and opted out of Copilot. Fwiw, I'm currently in Cambodia.
As others mentioned, AI is making headspace in enterprise and accounting, and achieving the “last mile” of human work. Better image recognition for handwritten forms and mail, better content and sentiment analysis for reducing spam, robot arm tasks which are more and more complex (yet still tame compared to humans)
“AI” hype is indeed overrated. If you think we’re close to reaching the singularity or anything resembling skilled human work you will almost certainly be disappointed. We probably have decades of slow improvement, more and more of these “breakthroughs” which aren’t really amazing compared to a human 5-year old, and aren’t really going to revolutionize industry, but will nonetheless have practical benefits
The actual models work fantastically well.
The board games are merely a cover to advertise to AI Researchers and portray AI as "innocent" in the public eye.
Stratego is Google goofing off.
The Ferrari AI models are being used by Google to absolutely swindle money in some ad tech niche.
I agree that these specific models are not going to be useful outside of board games. But in the future when there is the opportunity for AIs to interact with the world for real, the this kind of research will allow AIs to dramatically outperform humans on these tasks.
That's what one lab is doing.
You cannot be blind to many many applications that are finding their ways to consumers and earning people money.
At least this new model is very different to the approach taken in Alpha Go.
Don’t get me wrong, using AI for that purpose is pretty amazing (but can also lead to some sketchy results if you don’t know what you are doing[1]) but pretending it will lead to some “general AI” is nothing but hype IMO. And teaching AI to play these board games better then a grandmaster only serves to increase that hype.
1: https://www.vox.com/recode/2019/8/15/20806384/social-media-h...
There are for sure use cases for inference models in generalized (or rather ill-understood; or even highly dynamic) non-linear systems, and deep learning models kind of ace at that—given enough training data and a lot of computational power. However I’m not really sure what we will use AGI for.