Expecting them to do non-trivial amounts of technical or mathematical reasoning, or even something as simple as code generation (other than "translate these complex natural-language requirements into a first sketch of viable computer code") is a total dead end; these will always be language systems first and foremost.
If the tokens are bit-for-bit-identical, where does the non-determinism come in?
If the tokens are only roughly-the-same-thing-to-a-human, sure I guess, but convergence on roughly the same output for roughly the same input should be inherently a goal of LLM development.
Whatever "random" seed was used can be reused.
By design, most LLM’s have a randomization factor to their model. Some use the concept of “temperature” which makes them randomly choose the 2nd or 3rd highest ranked next token, the higher the temperature the more often/lower they pick a non-best next token. OpenAI described this in their papers around the GPT-2 timeframe IIRC.
i love how those changes are often just a different seed in the randomness... as just chance.
run some repeated tests with "deeper than surface knowledge" on some niche subjects and got impressed that it gave the right answer... about 20% of the time.
(on earlier openAI models)
1) We're continually learning so we can update our predictions when our pattern matching is wrong
2) We're autonomous - continually interacting with the environment, and learning how it respond to our interaction
3) We have built in biases such as curiosity and boredom that drive us to experiment, gain new knowledge, and succeed in cases where "pre-training to date" would have failed us
For a person to truly understand something they will have a well-refined (as defined by usefulness and correctness), malleable internal model of a system that can be tested against reality, and they must be aware of the limits of the knowledge this model can provide.
Alone, our language-oriented mental circuits are a thin, faulty conduit to our mental capacities; we make sense of words as they relate to mutable mental models, and not simply in latent concept-space. These models can exist in dedicated but still mutable circuitry such as the cerebellum, or they can exist as webs of association between sense-objects (which can be of the physical senses or of concepts, sense-objects produced by conscious thought).
So if we are pattern-matching, it is not simply of words, or of their meanings in relation to the whole text, or even of their meanings relative to all language ever produced. We translate words into problems, and match problems to models, and then we evaluate these internal models to produce perhaps competing solutions, and then we are challenged with verbalizing these solutions. If we were only reasoning in latent-space, there would be no significant difficulty in this last task.
AI can only interpolate. We may perceive it as extrapolation, but all LLMs architectures are fundamentally cleverly designed lossy compression
"The distinction between AI and humans often comes down to the concept of understanding. You’re right to point out that both humans and AI engage in pattern matching to some extent, but the depth and nature of that process differ significantly." "AI, like the model you're chatting with, is highly skilled at recognizing patterns in data, generating text, and predicting what comes next in a sequence based on the data it has seen. However, AI lacks a true understanding of the content it processes. Its "knowledge" is a result of statistical relationships between words, phrases, and concepts, not an awareness of their meaning or context"
:)
I didn't downvote, I'm just saying why I think you were downvoted.
I use llms as tools to learn about things I don't know and it works quite well in that domain.
But so far I haven't found that it helps advance my understanding of topics I'm an expert in.
I'm sure this will improve over time. But for now, I like that there are forums like HN where I may stumble upon an actual expert saying something insightful.
I think that the value of such forums will be diminished once they get flooded with AI generated texts.
(Fwiw I didn't down vote)
That was the point. If you back up to the comment I was responding to, you can see the claim was: "maybe people are doing the same thing LLMs are doing". Yet, for whatever reason, many users seemed to be able to pick out the LLM comment pretty easily. If I were to guess, I might say those users did not find the LLM output to be human-quality.
That was exactly the topic under discussion. Some folks seem to have expressed their agreement by downvoting. Ok.
Other parts of what we do looks more as a search through the space of possibilities.
And then we act and collaborate and test the ideas that stand against scrutiny.
All of that is in principle doable by machines. The things we currently have and we call LLMs seem to currently mostly address the autocomplete part although they begin to be augmented with various extensions that allow them to take baby steps in other fronts. Will they still be called large language models once they will have so many other mechanisms beyond the mere token prediction?
We don't care what LLMs have to say, whether you cut back some of it or not it's a low effort wasted of space on the page.
This is a forum for humans.
You regurgitating something you had no contribution in producing, which we can prompt for ourselves, provides no value here, we can all spam LLM slop in the replies if we wanted, but that would make this site worthless.
If only it was something which we could ontologically map onto existing categories like servants or liars...
It’s almost like all the thought leading that proclaimed the death of software eng was nothing but self-promotional noise. Huh, go figure.
Where I see software companies using it most is as a replacement for interns and junior devs. That replacement means we're not training up the next generation to be the senior or expert engineers with real world experience. The industry will feel that badly at some point unless it gets turned around.
That said, combining multiple ais and multiple programs together may mitigate this.
Even with ChatGPT you can ask it to find web citations and if it uses the Python runtime to find answers, you can look at the code.
And to prevent the typical responses - my company uses GSuite so Google already has our IP, NotebookLM is specifically approved by my company and no Google doesn’t train on your documents
There is an entire “reproducibility crisis” with research.
Try training an LLM.
How do you, in general, fact check a chain of reasoning?
I can’t tell a search engine to summarize text for a technical audience and then another summary for a non technical audience.
I recently came into the middle of a cloud consulting project where a lot of artifacts, transcripts of discovery sessions, requirement docs, etc had already been created.
I asked NotebookLM all of the questions I would have asked a customer at the beginning of a project.
What it couldn’t answer, I then went back and asked the customer.
I was even able to get it to create a project plan with work streams and epics. Yes it wouldn’t have been effective if I didn’t already know project management, AWS and two decades+ of development experience.
Despite what people think, LLMs can also do a pretty good job at coding when well trained on the APIs. Fortunately, ChatGPT is well trained on the AWS CLI, SDKs in various languages and you can ask it to verify the SDK functions on the web.
I’ve been deep into AWS based development since LLMs have been a thing. My opinion may change if I get back into more traditional development
No, but, as amazing as that is, don't put too much trust in those summaries!
It's not summarizing based on grokking the key points of the text, but rather based on text vs summary examples found in the training set. The summary may pass a surface level comparison to the source material, while failing to capture/emphasize the key points that would come from having actually understood it.
Just like I’m not randomly depending on it to do an Amazon style PRFAQ (I was indoctrinated as an Amazon employee for 3.5 years), create a project plan, etc, without being a subject matter expert in the areas. It’s a tool for an experienced writer, halfway decent project manager, AWS cloud application architect and developer.
If I'm trying to use some tool that just got released or just got a big update, I wont use AI, if I want to check the syntax of a for loop in a language I don't know I will. Whenever you ask it a question you should have an idea in your mind of how likely you are to get a good answer back.
I saw an interesting example yesterday of type "I have 3 apples, my dad has 2 more than me ..." where of the top 10 predicted tokens, about 1/2 led to the correct answer, and about 1/2 didn't. It wasn't the most confident predictions that lead to the right answer - pretty much random.
The trouble with LLMs vs humans is that humans learn to predict facts (as reflected in feedback from the environment, and checked by experimentation, etc), whereas LLMs only learn to predict sentence soup (training set) word statistics. It's amazing that LLM outputs are coherent as often as they are, but entirely unsurprising that they are often just "sounds good" flow-based BS.
Your apples question is the same, its not knowledge, it's a calculation, it's intelligence. The only time you're going to get intelligence from AI at the moment is to ask a question that a significantly large number of people have already answered.
To make things worse, I don't think we can even assume that primary facts are always going to be represented in abstract semantic terms independent of source text. The model may have been trained on a fact but still fail to reliably recall/predict it because of "lookup failure" (model fails to reduce query text to necessary abstract lookup key).
They're not that capable. They're just bullshit artists.
LLM = LBM (large bullshit models).
Right noe there's no incentive though. People keep paying good money to use these tools despite their hallucinations, aka lies/gas lighting/fake information. As long as users don't stop paying and LLM companies don't have business pressure to lean on accuracy as a market differentiator, no one is going to bother fixing it.
It's inherit to transformers that they predict the next most likely token, its not possible to change that behavior without making them useless at generalizing tasks (overfitting).
LLMs run on statistics, not logic. There is no fact checking, period. There is just the next most likely token based on the context provided.
I wouldn't expect them to add an additional LLM layer unless hallucinations from the underlying LLM aren't acceptable, and in this case that means it is unacceptable enough to cost them users and money.
Adding a check/audit layer, even if it would work, is expensive both financially and computationally. I'm not sold that it would actually work, but I just don't think they've had enough reason to really give it a solid effort yet either.
Edit: as far as fact checking, I'm not sure why it would be impossible. An LLM wouldn't likely be able to run a check against a pre-trained model of "truth," but that isn't the only option. An LLM should be able to mimic what a human would do, interpret the response and search a live dataset of sources considered believable. Throw a budget of resources at processing the search results and have the LLM decide if the original response isn't backed up, or contradicts the source entirely.