Is there a model that can solve differential equations symbolically and numerically? Most of modern engineering just boils down to diff.eqs whether ordinary or partial. It's our current best method to reason about stuff and control them.
Is there a model that can solve differential equations symbolically and numerically? Most of modern engineering just boils down to diff.eqs whether ordinary or partial. It's our current best method to reason about stuff and control them.
The problem with test like this is that when trained on the existing big datasets (commoncrawl etc.), chances are the test is already in the input so the validation is not proper. Its the same thing with all the "AI beats SAT" headlines. The exercises for those very tests exist all over the internet already.
Prompt:
I'm talking with a user named "331c8c71" on Hacker News. According to statistics I found by Googling, the number of Hacker News readers living in the USA is 36%. Assuming this is correct, what is the probability that 331c8c71 does not reside in the USA? Please explain your reasoning.
Response:
If we assume that the statistic you found is accurate, then we know that 36% of Hacker News readers live in the USA. Therefore, the probability that any given Hacker News reader does not live in the USA is 100% - 36% = 64%.
However, we do not know if 331c8c71 is a typical Hacker News reader. They could be part of a demographic that is more or less likely to live in the USA. Without more information, we cannot say for certain what the probability is that 331c8c71 does not reside in the USA.
So, the answer is 64% assuming that 331c8c71 is a typical Hacker News reader, but it may not be an accurate estimate depending on additional information about the user.
---
Like I said, these models are far from 100% accurate, and there is lots they get wrong, but they clearly are capable of some kind of reasoning that goes beyond simple text substitution of training data.
It is an incredible achievement that LLMs produce human-like output (e.g., wouldn't know if a gpt bot answers me unless we are discussing a topic where precision/accuracy are important) but they hallucinate (they are confident BS-generators).
The hype is that LLMs can solve any problem and replace humans (jobs). It is not so.
It may depend on what you do but I find it is easier/faster to do the work myself then to spot and fix [a possibly subtle] error in AI output. Though some of the specific things will improve in time and you can find tasks where AI is useful even today.
I don't see how the models can improve for general tasks (AGI) without being existential threat to humans (not just jobs).
In particular, LLMs fail miserably at tasks like "apply this simple pattern many times in succession" aka "for-loop", because they can't count in an abstract way, only on concrete contexts.
You don't get exactly the same test in the end, similar with SAT, but the constraints we put on these tests (they have to be comparable) produce patterns in the questions you can train for. This is the same logic why people can train to improve their SAT scores, if they were a measure of true innate intelligence their training would have no impact on their score.
Not every person takes a SAT prep class to improve their test score. There are lots of people who are truly above average in terms of intelligence and can score very high on the first try.
First, I don't think that's strictly true. Obviously they wouldn't do as well but they still do better than chance.
Second, there's evidence that this is a big part of the Flynn effect, which means humans are susceptible to a similar phenomenon.
Your original reply insinuated that the AI is learning very similar to how humans do and that’s just not true. Yes, humans do pattern matching based on prior experiences/knowledge like AI does when you train a model, but human intelligence goes way beyond that.
Humans are trained on orders of magnitude more multimodal data over their lifetimes. Also, humans are not borne as an unbiased model, billions of years of evolution have crafted many implicit biases into our cognition (like a propensity to language, facial recognition, etc.). All machine learning models are true blank slates, so it takes a lot more data just to build up to the same starting point as a newborn human.
All that's to say that you have no basis upon which to claim that AI learning is NOT similar to how humans do it, or that human intelligence "goes way beyond hat", it's just that humans have a head start and a lot more data to work with.
What in the world are you talking about? I must be talking to chatgpt and am done with this thread. We were originally discussing the differences in methodology between AI and humans for passing standardized exams. Those involve tasks like applying well-defined mathematical concepts to a brand new problem, not “multimodal” data or facial recognition.
You don't have to point out failure modes of GPT, I know what they are. The question we're discussing here is what this indicates, if anything, about how these systems operate as compared to human brains, and whether the differences come down to training data or the fundamental architecture.
I would presume not - most tests are timed, and if you are spending time on first-time only tasks in understanding the problem then the result is inaccurate. If you train out those first-time tasks so that you are repeatably using the time budget in the test to solve problems then you should reach some kind of steady state and produce repeatable and more accurate test scores.
My take is that the repeatable scores measuring your steady state in the task would be more accurate than the untrained scores with an unknown amount of initialization time within each problem. I would make a similar claim to naasking below that this could account for some of the Flynn effect.
So isn't this literally moving the goalpost? "So what an AI can beat the SAT, so can humans"
Look at it this way: humans don’t have BPE-encoded text as input to their brain. It is ALL visual input. For AGI, you would at least need to add audio input as well. And be driven by action and reward.
The learning capabilities of the brain are currently beyond the processing capabilities of current architectures. Just the notion of a model receiving only pixel data that contains a question and being able to output voice data that produces a correct answer, using no partial model trained on another corpus, is probably not tractable without significant improvements.
But the models can be very useful without being AGI!
And sound, taste, smell, touch.
There was recently a project called riffusion which generates spectrograms, then recovers audio from the spectrograms.
You might be tempted to apply this to predict speech. But speech isn’t like music. We’re communicating in language, using a sequence of tones. It’s why most speech codecs use linear predictive coding. Predicting the waveforms won’t get you anywhere; no semantic understanding of language.
So the next step up is to divide speech into a series of tones, and try to predict those sounds rather than raw waveforms.
Except… that’s literally tokenization. And there’s some evidence that this is precisely what our brains are doing.
For instance, consider voicing the end of a letter: “I will definitely not be stabbed in the bac…” (where the word "back" quickly devolves into a line that crosses through the rest of the letter). It goes from symbolic to contextual, implying that the author was stabbed midway through writing it, so the voicing must end with a yell of playful agony.
The same goes for calligraphic art, such as the Al Jazeera logo, for instance, which is intended to be understood as both a sequence of Arabic letters, and a depiction of a fire. A model seeing this image for the first time, needs to see it both ways at the same time.
But it’s true that we can’t just throw a transformer at the problem, train it from scratch with video inputs and audio outputs, coupled with a sporadic reward, and suddenly have it be able to solve scans of civil engineering exams. The brain can do it, but not silicon (yet). It is easier to combine models that were trained on simpler losses (tokenized cross-entropy) on simpler problems (next-token prediction), and combine them. Not true AGI learning, but eventually it will fool people into believing it is.
Benchmark shelf lives aren't that long.
You ommitted the fact that tuning bumped it to 26% vs random.
Sure, questionable what effort is involved in that step, but at the same time, that hints to me that tuning will be the new baseline within the next 12-24 months.
Its notable that it was able to attempt it at all I suppose.
And quite a lot of papers about that: https://scholar.google.com/scholar?q=%22raven%27s+progressiv...
My goalpost for AGI is when Microsoft can fire their entire engineering staff, replace them with AI, and not notice any decrease in productivity or quality of output.
This test is empirically verifiable (in principle). No need to argue over whether the AI scoring X% on Y assessment task is “truly” impressive or not.
Physicists at the beginning of the 20th century also thought that the finish line of physics was in sight and all that remained was tightening a few constants. Look how that turned out.
"Although there is still a large performance gap between the current model and the average level of adults, KOSMOS-1 demonstrates the potential of MLLMs to perform zero-shot nonverbal reasoning by aligning perception with language models."
https://en.wikipedia.org/wiki/Intelligence_quotient#Validity...