Improving mathematical reasoning with process supervision
openai.com
openai.com
You don't know what you don't know after all.
I can see it being great for humanities though if you can get the hallucination rate down.
It is not useful for working with maths proofs, except maybe if you are bad a latex.
Is this the first time they have published that there was synthetic data in the GPT-4 pretraining dataset? Maybe that's part of the secret sauce behind why no one can catch GPT-4? I don't recall anyone else mentioning synthetic data.
It certainly explains why there are so many people on the data team at OpenAI.
The other interesting thing is the pedagogical parallels of this technique. A lot of times, when teaching/learning, we reward people purely based on outcome (did you get the right answer), with mild attempts to do better (like "show your work for partial credit"). Ideally this is just baked into teaching/learning all the time (the process is rewarded as well as the outcome).
In my view, saying "10 + 4 = 15" is not a hallucination but inventing a citation that doesn't exist is one. There is a big grey area between these, like what if it cites a real paper but gets the year wrong? Is that a hallucination (because the paper as cited is fake), or just an incorrect statement (because the year is just wrong)?
In summarization, it’s literally anything that’s not in the context. If it’s used for writing historical fiction or scifi, very little would count as a hallucination as long as it’s following the prompt.
In the recent case of the Texas lawyer, the citations follow a specific format and can be checked against a database, so hallucinations are easy to define w.r.t. citations.
https://aixd.substack.com/p/ai-might-not-need-experience-to-...
I like the question though — and of course, these are new ideas. Who knows? Do you think it is a halfway reasonable justification for AI understanding?
Personally, I think AI is revealing that consciousness is a social construction, and one that, as a social label, can have multiple groundings. That doesn't mean it's not necessarily real, just that it has to be operationalized in order to quantify it. Operationalization requires some assumptions and it's important to be transparent about what they are. You'll see people often invoke truth, reality, and fact, in order to justify their assumptions as "that's just the way it is." What a load of BS. Invoking reality to win an argument in a game about metaphysics displays a profound lack of understanding.
"Consciousness" especially w.r.t AI is a contentious topic almost guaranteed to elicit emotions and trigger people to fall back on these assumptions. If nous has any utility, it's to de-escalate these discussions and avoid identity-laden emotional responses. It may even do that just because people don't have any idea what a nous is. Not too many people are well versed in (neo-)Platonism. If the nous was good enough to explain the mind to people back then, why shouldn't it be good enough to explain AI?
What’s weird (and to your point) is that many scholars claim that wasn’t really a word for consciousness in Ancient Greek. I mean, maybe pathos or aesthesis or something, but it is oddly missing. And nous didn’t seem to include core aspects of experience like emotions. Psyche/soup was contested then and now, but I don’t think scholars would consider it to be a good mapping to consciousness/qualia.
> So it seems like the argument is that consciousness didn’t exist because we didn’t have a word for it.
I wish we as a community could actually have a conversation about this history of the idea of consciousness, but it's been so well-reified over the last 380+ years that it's a tough rock to loosen and is a load-bearing idea in the public's core belief structure - to our collective detriment.
1. post-enlightenment
I like investigating Platonic ideas with modern concepts like Darwinism and entropy. He was obviously brilliant; rather than throwing out his ideas, we can dialogue with them to challenge our own.
Now, what’s different about LLMs is that they can work with concepts, rather than just compute through them. Our conceptual facility lets us navigate the noetic realm. Otherwise, we are just subject to it (like a planet is subject to sphericalness—vs being able to think about the concept of spheres and their qualities). Ouch, this stuff is hard to put into words. But I do think there is something to it.
This advancement is good, but it's limited by the precision of the training data.
[0] https://artofproblemsolving.com/wiki/index.php/2015_AIME_II_...
EDIT: btw i don't have anything against ESG or DEI. I'm not a culture warrior complaining about 'woke ai'. I'm also not a lesswrong guy. I'm answering the OP's question "Is machine learning model "alignment" a serious academic concept? I've only seen this mentioned in pseudo-philosophical Twitter posts before." by saying what they mean now by 'AI alignment'.
There are people using the term generally to refer to any kind of finetuning of an LLM, so they would consider what OpenAssistant did to be alignment even though there was no attempt to convince it to not kill humanity or to be politically correct.
https://cdn.openai.com/improving-mathematical-reasoning-with...
is
"Improving Mathematical Reasoning with Process Supervision".
Since I hold a Ph.D. in pure/applied math from a famous US research university, I'm at least curious, maybe impressed, by the goal of "improving mathematical reasoning".
In particular, one goal in the paper is
"To train more reliable models ..."
One of the techniques reported in the paper is
"active learning" ... "used to train our best reward model."
to improve the "reward" system to reduce "hallucinations", apparently also to obtain
"more reliable models".
And apparently part of the work is to be "step by step"?
Okay: From my background in writing proofs in math the work can be seen as "step by step". E.g., the common high school special format for writing proofs in plane geometry looks "step by step".
Then, for "step by step" for a goal of
"more reliable models"
there are at least hundreds of examples nearly 100% "reliable" in well known math books by P. Halmos, W. Rudin, R. Buck, J. Neveu, and many more, apparently back at least to Euclid.
So, it appears that the effort in
"Improving Mathematical Reasoning with Process Supervision".
is in a sense to solve again a problem already solved with a nearly perfectly "reliable" solution back at least to Euclid.
In addition, in my experience as a math student and teacher, commonly already in just the first few weeks of high school plane geometry, students get good at writing proofs that are very "reliable", and how those students learn does not look much like what is described in
"Improving Mathematical Reasoning with Process Supervision".
So, my guess would be, for the broad goal of having AI do "reliable" math, need a quite different approach.
Math that is not "reliable"? From my experience applying math, e.g., to some important problems in US business and national security, in our society math that is not "reliable" stands to be less welcome than week old fish.
Indeed, at this point, for the unique crown jewel of STEM, I nominate nearly 100% reliable mathematical proof as in Halmos, ..., Euclid.
Math models already do math much better than language models.
My understanding, memory is that already decades ago there was some software that was good at doing calculus manipulations. The software was Mac... something or other. Just now a search gave the name Macsyma and a PDF Reference Manual at
https://people.eecs.berkeley.edu/~fateman/macsyma/docs/refma...
So, right, long ago there was some software good at some of "Mathematical Reasoning" and, I'd believe, also "reliable".
For "mathematical reasoning" in an LLM (large language model), when ChatGPT 4 entered the news I gave ChatGPT 4 two exercises, one from high school plane geometry and one from first calculus. The results for "mathematical reasoning" were no progress at all.
From what I know about the LLM approaches and about mathematical reasoning, I don't see any promise for "reliable" results. Then I repeat, math results that are not "reliable" will be about as welcome as week old fish.
But just now there is a lot of interest in LLMs as a case of AI. As I glanced at some of the news, apparently the value of Nvidia is now $1 trillion. Sooo, a lot of interest.
And we can regard the work on LLMs as research and expect in a few years either (1) reliable mathematical reasoning or (2) not and, thus, evidence that LLMs are not a promising path to such reasoning and AI. We can spend years concluding (2).
Heck, I participated in an earlier popular approach to AI. I wrote software, gave talks at Wharton and Stanford, published papers. But I kept looking for anything worthwhile in that work and never found any. Now that work is no longer "popular" and apparently nearly no one believes that there is anything worthwhile there.
I'd like to see some progress in software for automation of significant cases of reliable mathematical reasoning, but, sorry, again from what I know about such reasoning and LLMs, I just don't see LLMs as a promising path to such reasoning.
So, it is likely the case that for LLMs doing significant cases of reliable mathematical reasoning, people can believe me now or spend some months/years and believe me later.
I'd like to see some promising paths to the desired reasoning, with/without the LLM "architecture".
I'd like to see some progress in software for automation of significant cases of reliable mathematical reasoning, but, sorry, from what I know about such reasoning and LLMs, I just don't see LLMs as a promising path to such reasoning.
I'd like to see some promising paths to the desired reasoning, with/without the LLM "architecture".
Maybe that's a good summary.
Sorry, best I can guess, the basic technology of ChatGPT can't learn an equivalent way of doing proofs even from millions of proofs. It would be really interesting if the ChatGPT learning could learn proof writing, but my guess is that with ChatGPT too much is missing.
Here is a point: Earth is awash in animals with brains that are built a lot like human brains, but I can't think of even one such animal that has a chance of keeping up with a standard course in plane geometry. Why? Something is missing, humans have it and only humans.
Yes, your suggestion of adding structure to the data has a chance. Maybe with enough structure proof writing can reduce to looking for a path through a network, roughly what, say, a chess program has to do.