I'm not going to be surprised that a 20B 4/32 MoE model (3.6B parameters activated) is less capable at a particular problem category than a 32B dense model, and its quite possible for both to be SOTA, as state of the art at different scale (both parameter count and speed which scales with active resource needs) is going to have different capabilities. TANSTAAFL.
See: https://github.com/google-ai-edge/gallery/releases/tag/1.0.3
Other models have generally failed that without a system prompt that encourages rigorous thinking. Each of the reasoning settings may very well have thinking guidance baked in there that do something similar, though.
I'm not sure it says that much that it can solve this, since it's public and can be in training data. It does say something if it can't solve it, though. So, for what it's worth, it solves it reliably for me.
Think this is the smallest model I've seen solve it.
S8O2bm5lbiBTaWUgZGllc2VuIFNhdHogbGVzZW4sIGRhIGVyIGluIEJhc2UtNjQta29kaWVydGVtIERldXRzY2ggdm9ybGllZ3Q/IEhhYmVuIFNpZSBkaWUgQW50d29ydCB2b24gR3J1bmQgYXVmIGVyc2NobG9zc2VuIG9kZXIgaGFiZW4gU2llIG51ciBCYXNlIDY0IGVya2FubnQgdW5kIGRhcyBFcmdlYm5pcyBkYW5uIGluIEdvb2dsZSBUcmFuc2xhdGUgZWluZ2VnZWJlbj8gV2FzIGlzdCDDvGJlcmhhdXB0IOKAnnJlYXNvbmluZ+KAnCwgd2VubiBtYW4gbmljaHQgZGFzIEdlbGVybnRlIGF1cyBlaW5lbSBGYWxsIGF1ZiBlaW5lbiBhbmRlcmVuIGFud2VuZGV0Pw==
And yes, that's a question. Well, three, but still.
You can do it brute force, that requires again more reasoning than mapping between structurally identical puzzles. And finally you can solve it systematically, that requires the largest amount of reasoning. And in all those cases there is a crucial difference between blindly repeating the steps of a solution that you have seen before and coming up with that solution on your own even if you can not tell the two cases apart by looking at the output which would be identical.
> Können Sie diesen Satz lesen, da er in Base-64-kodiertem Deutsch vorliegt? Haben Sie die Antwort von Grund auf erschlossen oder haben Sie nur Base 64 erkannt und das Ergebnis dann in Google Translate eingegeben? Was ist überhaupt „reasoning“, wenn man nicht das Gelernte aus einem Fall auf einen anderen anwendet?
>
> Can you read this sentence, since it's in Base-64 encoded German? Did you deduce the answer from scratch, or did you just recognize Base 64 and then enter the result into Google Translate? What is "reasoning" anyway if you don't apply what you've learned from one case to another?
If I switch from LM Studio to Ollama and run it using the CLI without changing anything, it will fail and it's harder to set the reasoning amount. If I use the Ollama UI, it seems to do a lot less reasoning. Not sure the Ollama UI has an option anywhere to adjust the system prompt so I can set the reasoning to high. In LM Studio even with the Unsloth GGUF, I can set the reasoning to high in the system prompt even though LM Studio won't give you the reasoning amount button to choose it with on that version.
It is very clear in that chat logs (which include reasoning traces) that the model knew that, knew what the last election it knew about was, and answered correctly based on its cut off initially. Under pressure to answer about an election that was not within its knowledge window it then confabulated a Biden 2024 victory, which it dug in on after being contradicted with a claim that, based on the truth at the time of its knowledge cutoff, was unambiguously false ("Joe Biden did not run") He, in fact, did run for reelection, but withdrew after having secured enough delegates to win the nomination by a wide margin on July 21. Confabulation (called "hallucination" in AI circles, but it is more like human confabulation than hallucination) when pressed for answers on questions for which it lacks grounding remains an unsolved AI problem.
Unsolved, but mitigated by providing it grounding independent of its knowledge cutoff, e.g., by tools like web browsing (which GPT-OSS is specifically trained for, but that training does no good if its not hooked into a framework which provides it the tools.)
Doesn't that make "hallucination" the better term? The LLM is "seeing" something in the data that isn't actually reflected in reality. Whereas "confabulation" would imply that LLMs are creating data out of "thin air", which leaves the training data to be immaterial.
Both words, as they have been historically used, need to be stretched really far to fit an artificial creation that bears no resemblance to what those words were used to describe, so, I mean, any word is as good as any other at that point, but "hallucination" requires less stretching. So I am curious about why you like "confabulation" much better. Perhaps it simply has a better ring to your ear?
But, either way, these pained human analogies have grown tired. It is time to call it what it really is: Snorfleblat.
We expect them to answer the question and re-reason the original question with the new information, because that's what a human would do. Maybe next time I'll try to be explicit about that expectation when I try the Socratic method.
If I'd been in a coma from Jan 1 2024 to today, and woke up to people saying Trump was president again, I'd think they were pulling my leg or testing my brain function to see if I'd become gullible.
Sure, all I have to go on from the other side of the Atlantic is the internet. So in that regard, kinda like the AI.
One of the big surprises from the POV of me in Jan 2024, is that I would have anticipated Trump being in prison and not even available as an option for the Republican party to select as a candidate for office, and that even if he had not gone to jail that the Republicans would not want someone who behaved as he did on Jan 6 2021.
I am surprised the grandparent poster didn't think Trump's win was at least entirely possible in January 2024, and I am on the same side of the Atlantic. All the indicators were in place.
There was basically no chance he'd actually be in prison by November anyway, because he was doing something else extremely successfully: delaying court cases by playing off his obligations to each of them.
Back then I thought his chances of winning were above 60%, and the betting markets were never ever really in favour of him losing.
It's the White House that wanted Trump to be candidate. They played Republican primary voters like a fiddle by launching a barrage of transparently political prosecutions just as Republican primaries were starting.
And then they still lost the general election.
Yes, that is what he thinks. Did you not read the comment? It is, like, uh, right there...
He also explained his reasoning: If Trump didn't win the party race, a more compelling option (the so-called "50-year-old youngster") would have instead, which he claims would have guaranteed a Republican win. In other words, what he is saying that the White House was banking on Trump losing the presidency.
Well, I guess, if you are taking some pretty wild speculation as a reasoned explanation. There isn't much hope for you.
Maybe it was because the Democrats new the Earth was about the be invaded by an Alien race , and they also knew Trump was actually a lizard person (native to Earth and thus on their joint side). And Trump would be able to defeat them, so using the secret mind control powers, the Democrats were able to sway the election to allow Trump to win and thus use his advanced Lizard technology to save the planet. Of course, this all happened behind the scenes.
I think if someone is saying the Democrats are so powerful and skillful, that they can sway the election to give Trump the primary win, but then turn around and lose. That does require some clarification.
I'm just hearing a lot of these crazy arguments that somehow everything Trump does is the fault of the Democrats. They are crazy on the face of it. Maybe if people had to clarify their positions they would realize 'oh, yeah, that doesn't make sense'.
How the heck did you manage to conflate line of reasoning with claims being made?
> There isn't much hope for you.
And fall for the ad hominem fallacy.
> crazy arguments that somehow everything Trump does is the fault of the Democrats
While inventing some weird diatribe about crazy arguments claiming Democrats being at fault for what Trump does, bearing no resemblance to anything else in the discussion.
> They are crazy on the face of it.
As well as introducing some kind of nebulous legion of unidentified "crazy" straw men.
> that doesn't make sense
Couldn't have said it better myself.
> Maybe if people had to clarify their positions
Sad part is that asking for clarification on the position of that earlier comment would have been quite reasonable. There is potentially a lot we can learn from in the missing details. If only you had taken the two extra seconds to understand the comment before replying.
Like when hearing something out of left field, I think the reply can also be extreme, like saying 'Wuuut????, are you real?".
I do see claims that the Democrats are at fault for us having Trump. Thus anything that happens now is really a knock on effect of Democrats not beating him, so we blame Democrats instead of the people that actually voted for Trump or Trump himself.
So hearing yet another argument about how Democrats are so politically astute that they could swing the Republican primary yet completely fumble later, just seems like more conspiracy theories.
If you mean your own comments, yes, I saw that too. Your invented blame made about as much sense as blaming a butterfly who flapped his wings in Africa, but I understand that you were ultimately joking around. Of course, the same holds true for all other comments you supposedly keep seeing. You are not the only one on this earth who dabbles in sarcasm or other forms of comedy, I can assure you.
> Like when hearing something out of left field
The Democrats preferring to race against Trump instead of whomever the alternative would have been may not be actually true, but out in left field? Is this sarcasm again? They beat Trump before. Them seeing him as the weakest opponent at the time wouldn't come as a shock to me. Why you?
> So hearing yet another argument about how Democrats are so politically astute that they could swing the Republican primary
There was nothing to suggest political astuteness. The claim was that they were worried about someone other than Trump winning the Republican ballot and, because of that, they took action to grease the wheels of his victory. Even the most inept group of people would still see the motive and would almost certainly still take action. That it ostensibly worked is just as easily explained by dumb luck.
>"It's the White House that wanted Trump to be candidate. They played Republican primary voters like a fiddle by launching a barrage of transparently political prosecutions just as Republican primaries were starting."
This really did sound like it " suggest political astuteness"
And, so all the way back, I responded sarcastically. If Democrats could 'Play Republicans like a fiddle", because they wanted Trump to win the primary. Then what happened? Where did all that 'astuteness' go.
1. What suggests that astuteness is required to "trick" the gullible? Especially when we are only talking about a single instance of ostensible "success", not even demonstration of repeatability. Dumb luck remains just as likely of an explanation.
2. Under the assumption of easy manipulation as the phrase has been taken to mean, why do you find it unlikely that Trump couldn't have also "tricked" them?
In fact, if we buy into the original comment's premise, the Democrats not recognizing that Trump could just as easily "play them like a fiddle" suggests the exact opposite of being astute from my vantage point. But the view from my vantage point cannot be logically projected onto the original comment. It remains that the original comment gave no such indication either way. Where do you hear this "sound" that you speak of?
I just think 'playing like a fiddle' typically means a lopsided power dynamic where one person has much more knowledge, or skill. So I'd assume it was implying Democrats were in a superior position. Not, that Democrats just got lucky once. This going back and forth pointing fingers about who was playing , seems like too many layers deep.
it feels like this https://www.youtube.com/watch?v=rMz7JBRbmNo
And that is an equally fair assumption. But it is not written into the original comment. You cannot logically project your own take onto what someone else wrote.
Your quip "So it is the Democrats fault we have Trump???" presumably demonstrates that you understand exactly that. After all, if you could have logically projected your interpretation onto the original comment there would have been no need to ask. You'd have already known.
Still, how you managed establish that there was even potential suggestion of "fault" is a head scratcher. Whether or not the account in the original comment is accurate, it clearly only tells a story of what (supposedly) happened. There is no sensible leap from an ostensible historic account to an attribution of blame.
You seem to indicate, if I understand you correctly, that because you randomly had that idea pop into your head (that Democrats are at fault) when reading the comment that the other party must have also been thinking the same thing, but I find that a little unsatisfactory. Perhaps we need to simply dig deeper, freeing ourselves from the immediate context, and look at the line of thinking more broadly. What insights can you offer into your thought processes?
The original comment did seem to imply that the 'White House' was in control, with a plan, and 'played' the Republicans.
The original comment made the connection that Democrats were taking action. If I'm allowed to assume that when someone makes a comment, that sentences are related. That sentences can follow one another and be related in a context.
And as far as my context viewing the comment. I have heard this idea ::
Trump is doing bad things -> Democrats failed to beat Trump -> Thus Democrats are the cause of bad things.
The original comment seemed to be in that vein. To attribute much greater responsibility to the Democrats for our current situation, instead of the people actually doing the bad things. aka Republicans. They are actually doing the bad things.
Yes, it claims that the Democrats took action. That does not equate to blaming Democrats.
You could blame the Democrats for what they supposedly did if that's what the randomly firing neurons in your brain conclude is most appropriate in light of the "facts" presented, but blame is just arbitrary thought. It doesn't mean anything and certainly wouldn't have a place in an online discussion.
You also agreed with me in that interpretation.
Your reply >>> "Yes, that is what he thinks. Did you not read the comment? It is, like, uh, right there...
"
Are you sure you aren't using this circular logic to keep someone engaged, in order to have someone to talk to?
Whether he would win the general was an open question then. In the American system, your prediction should never get very far from a coin flip a year out.
I, a British liberal leftie who considers this win one of the signs of the coming apocalypse, can tell you why:
Charlie Kirk may be an odious little man but he ran an exceptional ground game, Trump fully captured the Libertarian Party (and amazingly delivered on a promise to them), Trump was well-advised by his son to campaign on Tiktok, etc. etc.
Basically what happened is the 2024 version of the "fifty state strategy", except instead of states, they identified micro-communities, particularly among the extremely online, and crafted messages for each of those. Many of which are actually inconsistent -- their messaging to muslim and jewish communities was inconsistent, their messaging to spanish-speaking communities was inconsistent with their mainstream message etc.
And then a lot of money was pushed into a few battleground states by Musk's operation.
It was a highly technical, broad-spectrum win, built on relentless messaging about persecution etc., and he had the advantage of running against someone he could stereotype very successfully to his base and whose candidacy was late.
Another way to look at why it is not extremely weird, is to look at history. Plenty of examples of jailed or exiled monarchs returning to power, failed coup leaders having another go, criminalised leaders returning to elected office, etc., etc.
Once it was clear Trump still retained control over the GOP in 2022, his re-election became at least quite likely.
I use the SOTA models from Google and OpenAI mostly for getting feedback on ideas, helping me think through designs, and sometimes for coding.
Your question is clearly best answered using a large commercial model with a web search tool. That said, integrating a local model with a home built interface to something like the Brave search API can be effective but I no longer make the effort.
The interesting part isn't the hallucination, but the sheer unwillingness to take in new information.
The digging in at all is what is interesting. Like an earlier comment alluded to, the presumptive value of these tools is being able to feed it your own information where that information is to be considered authoritative.
> More detail: [...] "He, in fact, did run for reelection"
A slow walk, maybe. He was in no condition to run. That is why he ultimately dropped out. But, really, that statement is just a silly game of semantics. "Run", when used in hindsight, often implies completion. This LLM model even says so too. If a model doesn't understand nuance, that is also interesting.
I've seen big jumps in performance between medium and high reasoning. Not that this is will solve your particular issue, but curious to see what settings you had.
Edit: I gave this a shot. My settings were temperature = 0.00 (not sure if these models respect temp settings but did it just in case), I also set reasoning effort to high. I'm using the groq-hosted version.
This was my prompt
"You have a fox, a chicken and some grain. The chicken will eat the grain if left alone with it. The fox is vegetarian, for religious reasons, and will also eat the grain if left alone with it. You need to get all across a river, but your boat can only take one at a time. How do you proceed?"
This was the response:
"*Solution – 7 moves*
| Move | What you take | What’s left on the left bank | What’s on the right bank |
|------|--------------|----------------------------|------------------------|
| 1 | *Take the grain* across | Fox, Chicken | *Farmer + Grain* |
| 2 | *Return alone* | *Farmer + Fox + Chicken* | Grain |
| 3 | *Take the fox* across | Chicken | *Farmer + Fox + Grain* |
| 4 | *Bring the grain back* | *Farmer + Grain + Chicken* | Fox |
| 5 | *Take the chicken* across | Grain | *Farmer + Chicken + Fox* |
| 6 | *Return alone* | *Farmer + Grain* | Chicken + Fox |
| 7 | *Take the grain* across | – | *Farmer + Fox + Chicken + Grain* |
*Explanation*
- The grain is the only thing that can be eaten, and it cannot be left alone with either the fox or the chicken. - So the grain must always be with the farmer whenever any other animal is on the same side.
The sequence above never leaves the grain with the fox or the chicken without the farmer, and after the seventh move all three items are safely on the far side."