Custom bots for Unreal Tournament 2004 pass Turing test
eurekalert.org
eurekalert.org
For example, a chat bot can certainly behave as the worst chat user, but that doesn't mean they can hold an intelligent conversions, which is really what we should care about. AI should simulate human intelligence, but very often all we're seeing is simulation of human stupidity. That's what happened in this case as well.
To me, good AI would be characterized by ability to handle highly unusual situations, not by mimicking irrational behavior.
This is a much simpler set of behaviors for a computer to emulate. It's an interesting test of AI for sure, but it shouldn't be called a Turing test.
What was interesting though is that the bots and humans did not converge on optimization strategies, and part of appearing human was programming irrational behaviour (such as grudges) into the bots.
In this case, the fact that on average bots seemed scored as more "human" than actual humans is more of a sign of a critical flaw in the judging system than any great progress. It looks like they reduced typical human behavior to some very simple things, such as holding a grudge or other irrational behavior. If that was a major part of the judges' criteria (consciously or subconsciously) then all this contest proved is that bots can be programmed to be more irrational than human players.
Bonus points if the bot can sing annoying, off-key pop songs in the pre-game loading screen, to mimic the true CoD Xbox Live experience.
If the judges were at a competitive level, colour me impressed - but if it was their first time, or even their first week, I'm a little more skeptical. I don't think a novice player would understand the game well enough to judge well. It would be like attempting a traditional Turing test with humans who can't speak English fluently and were raised in a non-English culture: impressive, but no indicator of bots reaching human-like levels.
I use to play UT/CS competitively and worked for a startup that licensed our technology to id Software and Riot Games / League of Legends a long time ago.
UT2k4 was an amazing FPS game with a really steep learning curve. It's one of the only FPS games that I refer to as the "basketball" of online gaming. The diversity of movement, weapon tactics, and map control meant a seasoned gamer could really define their own style. But it also meant few people ever transitioned from public servers into competitive play because 1 pro could easily go Godlike and demolish an entire server, making it extremely frustrating and unexciting for casual gamers.
That being said, watching the videos included in this article signaled that these judges had no experience with UT2k4.
In a match with professional gamers, it wouldn't surprise me if those judges thought WE were the bots. 50%+ accuracy was not uncommon with prim shock or lightening gun.
There's a pretty broad area between those two extremes - I'm nowhere near competitive level, but I played (very casually, mainly with friends rather than online) a few hours a week for years, and I definitely understand the game well and can identify good and bad plays, and humans and bots.
I do thoroughly agree, though, that this is nowhere near the level of a true Turing test - but I thought it was still quite interesting!
--
Aside: I think these days with the rise of game spectating and commentary and analysis of matches, more and more non-competitive players are gaining deeper understanding of games they like even if they couldn't necessarily pull it all off themselves.
For a more modern game like Starcraft II for example I think you could find a very large number of non-competitive players that could reliably identify bots from humans.
What was the criteria? "That one must be human because it can move and shoot at the same time!"
This reminds me of the awesome days of the ReaperBot, that cheating bastard.
The judging system needs to be seriously reevaluated.