Like sure it didn’t have the inclination to make the sim and hardware designs, but it did make them though yes?
Like sure it didn’t have the inclination to make the sim and hardware designs, but it did make them though yes?
Thats why the math breakthrough a few weeks ago was so hotly debated. Because OpenAI is desperate to demonstrate that AI isnt just a fancy regurgitation machine, but it can actually develop novel thought. Because that would be the stock price jumps to end all stock jumps.
But then it turned out it was really just listening in on a math professors supposed-to-be-private conversations with another instance of openai, and it used his novel work as the trigger to prove the breakthrough first.
The reason people conclude AI 'thinks' is because tt can reference obscure or poorly documented things quickly (which is its primary advantage along with processing natural language prompts into tasks), which is why a lot of people with emotions confuse that action with inventing things.
It seems unlikely that recursively predicting the next word would lead to creativity or invention, but it doesn't seem impossible. Similarly, it seems unlikely that human thought works in a similar prediction loop, but it doesn't seem impossible.
Sure, but until we can turn that tautology into something more rigorous, we can't tell if the thing humans do is more or less than what some arbitrary non-human (machine, animal, or eventually perhaps alien) does.
Ignore the fact that we are human and thing our ingenuity is unique. Is definitionally linking human ingenuity and creativity useful when trying to compare to others?
I could say earth is habitable so any other planet must prove it is similar to earth to be habitable. That simply doesn't matter, though, if different environments could lead to similarly complex living systems regardless of atmospheric makeup, precense of liquid water, etc.
I'd argue the burden is on you to prove that the definition you propose is meaningful beyond our own hubris, rather than expecting an alternative to now have to prove that the hubris was unjustified.
Plausible? Absolutely. Did OpenAI behave badly in other ways regarding this issue? Yes. Does it help to assume unproven facts and then accuse people of reaching emotional decisions? Nope.
Perhaps it is going to be more like, we will see the singularity predict the future.
While “the model was trained on sessions including the ones in question” is one aspect; ‘the model produced an accurate prediction of future human thought’, I think, is another very interesting facet.
Models predicting future things might be how we see, actually, how we ourselves formulate thought - by saying, in big and small words, ‘something is about to happen’.
>unproven facts
This isn’t a court room. We’re discussing ways by which we humans both succeed and fail at reigning in our creations. The OpenAI kerfuffle is pretty much irrelevant already. Of course AI will fill in the gaps of human thought - it is literally constructed from the stuff, in every squeeze of the curd and whey.
Call that brute force perhaps, but I would consider it technically inventing something on the merit that it would at least be an abstraction above naively throwing everything against a wall to only throwing things that would most likely be sticky.
How are you measuring complexity here? How are you measuring inputs? Your average llm is trained on a corpus that vastly exceeds the amount of data I could read in my lifetime.
In terms of how it functions, because no matter how much data you feed an LLM it's still predicting tokens. That makes it incapable of any thought.
>Your average llm is trained on a corpus that vastly exceeds the amount of data I could read in my lifetime.
But that doesn't mean they are useless, they are good at consuming large amounts of data and collating it.
I don't know how correct I am but that's my understanding and it won't change, I feel pretty confident in my simplified view of things because the basics are still there.
That's a rather dismissive way of putting it. Moreover it's confusing the output format with the complexity of the output. I could just as easily say that no matter what the human brain does it's just generating stimulus to motor neurons. Such a simplification is just as wrong as saying that it's just as misleading is saying that an llm is just a token generator. What makes an llm able or unable to be complex is the process that creates those tokens
Okay but how is that relevant for measuring whether an llm is capable of creating original thoughts? Why do you think that complexity, as you say, is a more relevant Factor than simply the number of weights?
2024 called, it wants its talking points back. I don't think claims like these are defensible after all the progress we have witnessed in the last year alone.
That has nothing to do with the question of whether the mechanism that generates those signals is complex enough to say that it's "creative."
An "AI" is a box which lies dormant until a human, with motivation and agency, enters a prompt into it.
A person or company can use it in a way where it might invent something. but at the end of the day, its a tool, and its actually not doing anything on its own.
Its not solving math problems, a mathetmatician is using it to solve math problems. Just cus OpenAI is acting like its AI is solving stuff, its really not. They're just paying people to use AI to hammer problems.
I'd say that people's motivation and agency, like their taste, creativity, and opinions, are shaped by external inputs, so saying that their agency is "owned" by them is a stretch.
The alternative would be the proposition that a slave, lacking agency, has no ability to solve problems. I find that idea unpalatable.
on a surface level a human solving an (unsolved) math problem can look like this, and of course the tree of all possible symbols you can send to a proving assistant is much wider than what the human samples, and the same goes true for an LLM in a proving loop. It isn't "truly random", it can't possibly be (and solve the problem). Both humans and LLMs solving unsolved math problems are aggressively pruning mathematical syntax and logical strategy trees.
If we do a bad job building context, they do a horrible job contexting. People who have trouble working with AI have the same problem people have in general: if they can't figure out the context of the direction, then they make random decisions of doing anything. On the flip side, if you can build the proper context around a sufficiently powerful LLM, they can derive the context via the contexting they're good at.
This is why building documents, tests, and code all in some intent pattern via prompting allows them to do a significant amount of work a normal person would have a great effort t
I've heard that before but it just doesn't make sense in context of what I've seen llms do. If I ask an llm to write a poem about magnetic resonance and vampire rabbits it can do that it created a new thing. I can ask it to build a website for managing rabbit breeding that's also a new thing.
Another way of looking at it is that human beings, just like llms, can produce output based on their inputs. Most literature is inspired by other literature. Most music is inspired by other music. Most software is inspired by other software.
So I think we need to work on defining " new things" before we can definitely exclude them from llm's capabilities
No, it didn't turn out to be that. Someone made a claim, which is silly for many reasons. There's no serious support for this happening.
This is outdated. With RLVF, LLMs can create their own training data instead of relying on what it's been fed.
Regardless of the truth of the underlying assertion, this is about as textbook an example of circular logic as it gets. (at least combined with the implicit beginning assertion that "AI can't invent new things.")
This didn't actually happen.
> Following an investigation, we have confirmed that Buckmaster’s Codex prompts over the two months preceding this announcement and paper on September 8, 2026, could not have influenced the system in any way, including through training.
Getting an LLM to design something in its own simulator that is not accurate w.r.t reality is not useful nor terribly impressive.
More broadly than the existing answers (which are correct), For a layman, I'd also add that LLMs are essentially 'brains in a vat'. They can't confirm ground truth about physical reality. They only know what's in their training data and prompt, which is incomplete and can be incorrect. Even with real-time external sensors they are limited to the sensor's margin of error, range and trusting it's working correctly.
When properly trained, fine-tuned and prompted, LLMs can be very effective in well-defined, non-physical domains like logic, writing, math and code but making things function in the real-world quickly spirals into combinatorial complexity.
So are we and our sensors can be pretty vague in comparison, I can only imagine human error correction is pretty next level.
That's true. Of course we can hook an llm up to Motors and sensors. That's a robot. Or a self-driving car. So would you say that those devices can confirm the ground truth about physical reality, and therefore are capable of creativity?
Within the limits of resolution, range, location and veracity of sensors, a machine can register their reported state and use it as a variable. That's substantially different than the level of knowledge and understanding implied when we say a human "confirms ground truth about physical reality." The first ~half of the difference is humans have a deep world model about the planet which surrounds any sensor and human sensations are pre-processed and filtered by a highly evolved bio-chemical substrate before ever reaching the higher-order cognitive processes which assign words, meaning and qualia to them.
> therefore are capable of creativity?
I never mentioned creativity, nor would I in relation to LLMs. Like "Intelligence", "Creativity" is far too vague to be of any use in assessing the capabilities, limitations or utility of LLMs.
On HN, posts like the OP tend to attract POVs at polar extremes from "LLMs are nothing more than stochastic parrots" to "LLMs are (or can be) as intelligent, creative, innovative (etc) as any human or all humans combined." I've researched and thought a lot about these and related topics for a very long time, Neither POV is going to find any quick agreement or easy answers from me.
There's a tiny germ of truth somewhere in both extremes that's drowning in an ocean of confusion ranging from "definitionally or categorically muddled" to "mostly incorrect" to "not even wrong". But neither POV seems interested in anything more than drive-by hot takes, debating over-simplistic strawmen or trading 'gotcha' hypotheticals.
Great, but we also build robots with world models and sensory pre-processing. So nothing you've described is unique to humans.
> I never mentioned creativity, nor would I in relation to LLMs
The topic of this thread is whether LLMs are capable of genuinely contributing "new" ideas. So if that's not the point you're trying to make, I'm not sure of the purpose of your comment.
> But neither POV seems interested in anything more than drive-by hot takes, debating over-simplistic strawmen or trading 'gotcha' hypotheticals.
I think you've oversimplfying the discussion here.
In my view, at least, both "LLMs are nothing more than stochastic parrots" and "LLMs can be as intelligent, creative, innovative (etc) as any human or all humans combined." are simultaneously true, for the simple reason that humans are stochastic parrots. Our brains learn patterns and respond to stimuli.