I think companies like OpenAI are aiming for something far more ambitious, like solving a millennium prize problem (even with human assistance). That's the kind of news release that'll add another $100 billion to your market cap.
I think companies like OpenAI are aiming for something far more ambitious, like solving a millennium prize problem (even with human assistance). That's the kind of news release that'll add another $100 billion to your market cap.
An analogy: If it takes 20 years to create AGI equivalent to the village idiot. It take another couple of hours to go from that to Einstein.
Imo we are not currently in the beginning of an exponential curve re solving math with AI, and def not on the path of AGI. I understand that if one believes that we are on the path to AGI soon then we shall have these math-AI advancements quite soon, but I disagree with the premise.
Yet it only took the AI world 1.5 years to go from drawing child scribbles to replicating top artists with like 90% similarity (I can barely tell the difference between AI and human drawn art anymore with the new NovelAI model). It won't be long before AI starts to go superhuman in art skills.
It won't take long from a school-math model to math olympiad model (I'd say 1 year is enough), and going to unsolved conjectures won't be that long either (2-3 years?). We know from AlphaGo that its possible to make AI systems far superhuman at solving some abstract math problem.
Not sure. Art is about approximate pattern recognition and if you have a large enough dataset it seems that you can definitely reproduce some of that.
For math... it involves consistent reasoning from A to Z - which does not allow for any kind of mistake in the way. In Art you won't feel too bad if the shadows or the lights are a little weird or if a character has 7 fingers instead of 5 on one hand, but this kind of mishaps break everything in Math.
I think artists would disagree with your assessment of 'approximate pattern recognition'. Its more like:
1. Given a set of words describing what the user wants. 2. Arrange pixels in a grid 3. That maximizes the user rating
On one hand it is tolerant of small errors. On the other hand its an extremely broad problem. Also to get a good user rating, it has to do 99 things right for every 1 thing it draws wrong.
Now, if we end up seeing mastery, that would be extremely interesting.
No no, you see I prompted the AI with "masterpiece, photorealistic, 35 mm photography, cinematic, dslr, volumetric lighting, trending on artstation, 4k, 8k, hyper-detailed, epic digital painting by Greg Rutkowski" /s
Someone can tell you when a math problem is solved. Someone else can't tell you when you've successfully art'd with remotely the same degree of confidence. In so far as they can, however, many experts claim that AI cannot, indeed "do an art".
Also, it's easier to formulate a hard math question (there are plenty of unsolved problems), but that's (IMHO) harder to do for art. Sure, you may think this is the first time the phrase "Astronaut riding a Llama and holding an avocado" was writ, but those are all well represented concepts in the dataset. For more abstract prompts, there really isn't a way to verify "correctness".
If I was asked to draw
a. the emptiness in your heart
b. the lack of furniture in your room
c. your empty bank balance
d. starvation
e. object returned by a python function with no return
...
I could just submit an empty sheet of paper, & an artist would argue that my empty sheet of paper represents any/all of the above.
Now, if I turn in the same empty paper at a math qualifier and argue that it represents the infinite set of real and complex numbers, ergo the answer to the posed qual problem must be in there, I'll get kicked out of that phd program in a jiffy.
And they'd be taken about as seriously as the ads taped to a urinal in a Museum of Modern Art washroom.
In a way it is, in a way it isn't. You have to remember what is easy for machine isn't going to correlate to what is easy for us humans. Look at AI art. Closely. No, closer than that. All the detail is fucked up. Not just the hands, but the tiniest of things. Strokes, lighting, reflections, and consistency, and all that. But can I turn my friend into a convincing werewolf? Yes. Can I turn my cat into a human or Wonder Woman? No. The system isn't a "fancy copier" but it is a compression algorithm and the aforementioned tasks were only possible because lots of work training LoRAs, textual inversions, control nets, and so on (you could seriously improve GANs, VAEs, hell, even Boltzman Machines could probably do pretty well were any of these given the same research investment that diffusion has received. GANs come close but nuances like GANs having a magnitude fewer parameters).
But let's look at math, can I consistently add numbers? No. The problem is that in math, all those tiny intricate details matter. Not only that, they matter at every single step. The thing here is that these are still pattern recognition machines. But they aren't generalized machines. You can't really derive out all of math from probability distributions (or at least cleanly, but still not convinced you can). The thing is that for math to work in AI we have to address the elephants in the room: math. Yeah, math. ML people don't like it. But we gotta address the axioms in the room that we're operating under. How do we move on from machines operating on manifolds? How do we make it so data are not distributional? How do we move away from a number of unmentioned axioms remains a large open problem in AI research. One that does not get anywhere serious enough of a conversation, especially within the community. Sure, maybe transformer circuits can learn some addition by learning how to do FFTs and add in the FFT space, but you're not going to get to Abstract Algebra that way. Ideally the AI can solve problems that have no algorithms, pun intended.
AlphaGo didn't solve go (Ie, can the first mover guarantee a win?). However, it understood go at a far, far superior level to any human.
A mathbot doens't have to solve math in general. It merely has to be better at solving math than any human mathematician to be considered ASI. And it only has to be better than the 'average' human mathematician to be extremely useful in accelerating math research.
Solving Go means determining by what margin the first player can win, or equivalently, at what komi for white the game is a theoretical draw.
That also makes it dependent on the exact rule set used.
E.g. 2x2 go is a +1 first player win with the Tromp-Taylor rules of Go, while the Japanese rules are not even sufficiently formalized to allow scoring a 2x2 game.
Suppose that P1 passes on their first move (which is a valid move). Then P2 has a winning path of play in which they put down the first stone. But P1 could have made that move and then they would be on the winning path.
First player does not always have the advantage. Nim is the best example where the setup can either be the first player winning game (nim-sum of the sizes of the heaps is not zero) or the second. My understanding is also that Chess (another perfect information turn-based game) is not shown solved or even has proven first player advantage (though in practice it looks so).
So I get the argument, I just don't buy it. I would be inclined to lean towards that direction, but it's a tough claim theoretically and probably not meaningful in practice (unless a generalized strategy such as strategy stealing can be employed otherwise a lookup table is impractical as it'd contain more bits than atoms in the universe even for 100 move games).
I think we have to consider far more than strategy-stealing which is not even a generalizable strategy to two-person perfect information turn-based games.
https://en.wikipedia.org/wiki/Nim#Proof_of_the_winning_formu...
First, I thought going first in chess is generally considered an advantage. Even the wiki article states that. Or at least says there's a 10% increased win rate.
Second, I still don't get why passing is the important aspect. I thought the important aspect is symmetry. I mean I can understand this in nim since that symmetry is that killer aspect that makes for the easy analysis of a solution.
When I said I'm not a game theory person I didn't mean I have no game theory experience but that's not what I study. I'm on the mathy side of ML but not so much in RL. You can use math with me if that makes things easier (in fact, I love math. Please do. RL notation doesn't scare me but rather weirds me out that it scares others) because I think we're getting lost in the conditions.
I'm not a mathematician myself, just got into this stuff when I was working on a boardgame solver. I find it difficult to map the 'symmetric game' definition "the payoffs for playing a particular strategy depend only on the other strategies employed, not on who is playing them" onto a turn based game, but if it can work for Hex it must be compatible.
If you consider a strategy to be "a decision tree for how to place stones, when I'm playing as 2nd", then there's perfect symmetry between being 1st and choosing to immediately pass, and being 2nd. The possible strategies and resulting payoffs are the same. You add on top the extra move, which the possibility of passing means is at worst neutral for 1st player, and they cannot be at a disadvantage.
More intuitively for me: being allowed to pass your first move is the same as getting to pick which side you want to play. There's no way the side who can pick to play 1st or 2nd at their option can be at a forced loss to the side who just has to accept their decision. The picker would just pick the other side and now they have a forced win.
(I'm always assuming above any kind of infinite-pass-standoff is a draw, and not some kind of weird other thing).
If I haven't expressed my thoughts clearly enough that's probably about as well as I can manage I'm afraid.
Yeah so probably the better way to think about it might be with the payoff matrix. Because symmetry is actually about the strategy. That's why there are the notes about the laddering in Go. But the payoff of a symmetric game is actually when A = -A^T. So if we have a 2x2 game a symmetric zero-sum one is where the payoff matrix might look like [[0, 1], [-1, 0]] Where we're like an inverse-identity matrix (actually anti-symmetric) but the diagonals are opposite. Maybe it is best to think about this from a geometric perspective, this symmetry here (in this specific example) is a rotation matrix. That's what it does when applied to another matrix. Recall our standard form is [[cos(theta), -sin(theta)],[sin(theta), cos(theta)]]. Pretty easy to get our matrix from there if you remember that cos(90)=cos(180)=0 and sin(90)=1 but sin(180)=-1. So our angle of rotation is 180 degrees (or pi radians). You could also see that if we made the two columns vectors we'd see they pointed in opposite directions. That's the symmetry! Okay, yeah, maybe that's confusing lol. But I find it helpful to see matrices as transforms and I wish this was stated a bit more clearly and often.
So now that we maybe understand that, symmetry is about a __strategy__, not a player. Because our payoff matrix is strategy based. For example, our strategy for rock-paper-scissors is to pick each outcome 1/3 of the time, which gives us this symmetric payoff. But if we pick rock every time we don't get that payoff, right? So it's actually not about who goes first or second but also includes the strategy aspect.
At least that's my understanding which a lot is prompted by this conversation (thanks!)
The reason I'm finding the go argument hard is thinking of a basic "entropy" based strategy (it'll serve you well in boardgames, especially when sight reading). The idea is if you don't know the best move, play the move that gives you the most future moves. It'll trick you into thinking that this strategy is actually simple, it isn't. So in the game of Go, this isn't reasonably different from making a random move! Because there are just so many. And realistically your strategy is going to be the composition of many different strategies. Like you said, pull out the decision tree but we can actually abstract this a bit more and have a decision strategy tree that's a superset to our strategy that's a response (e.g. a ladder is set up so we play the laddering strategy). The reason I'm not buying the argument isn't about the logic, it is about the possible move sets. Even with super-ko (the board cannot return to a state it has previously been at any time in the game (must be fun to keep track of...)). So forgetting about all the extras that are played in go, passing shouldn't result in a meaningful change in the number of possible strategies. But this argument might actually be an argument in favor of symmetry, not against it. Coming back to Chess, we know that game __is not__ symmetric. Why? Because white has different strategies than black. If instead the first "move" is to flip a coin and that decides who is white and who is black, then the game actually becomes symmetric. Kinda wild...
I didn't read this, but a glance suggests that black dominates in smaller games
http://erikvanderwerf.tengen.nl/pubdown/thesis_erikvanderwer...
Does second mover in Go have some sort of artificial benefit in scoring or playing? As in -- is there something to compensate for moving second?
On the face it seems like first mover would have an advantage in any turn-based game. But maybe, in some games, seeing an opponent's strategy is more helpful than executing the strategy.
Also, are there examples of real games where second mover can always win? (Real as in, not made up with weird rules just to demonstrate it's possible.)
Yes. The second player typically gets an extra 6.5 or 7.5 points.
That's a tough thing for AI to do.
On the other hand, Terrence Tao had an interesting article on his blog a while back where he was trying to solve a problem and asked chatGPT about it in a high-level strategy sense. ChatGPT suggested several reasonable approaches, one of which turned out to work.
That's nowhere near solving a millennium problem, but it is very interesting and suggests fairly sophisticated conceptual understanding of mathematics nevertheless.
Current architecture and training methods I don't think are enough to get there. However, with enough compute, I can plausibly envision some sort of meta training of LLMs using an analogy to GANs where one network tries to synthesize new correct ideas and the other shoots them down as not novel, not correct, or not sufficiently interesting.
Such an approach I think could perhaps work, but the compute needed would probably be pretty high.
They decompose problems, solve specialized subsets, examine more general cases, use existing proofs, do some numerical analysis, etc.
Solving Millennium problems is a whole different ballgame. It's not known if these problems are solvable within ZFC axioms. (In one case, the Yang-Mills prize, stating the problem mathematically is part of the challenge.) All of the obvious applications of known tricks have been tried and failed. To solve such problems, one probably has to invent new and surprising mathematical definitions, building a framework in which the problem becomes solvable. This is something that LLMs will be crap at; the process of invention is not represented in any training data we have access to.
There are ~1000 MO winners and 1 (one) Millenial problem solver ...
Here's what Andrew Wiles, the only other person to have solved a Millennium-class problem has to say of math competition: "Let me stress that creating new mathematics is a quite different occupation from solving problems in a contest. Why is this? Because you don't know for sure what you are trying to prove or indeed whether it is true."
Nobody argues it's the same, after all MO problems are designed to be solved in ~an hour, but we are talking about mental capabilities.