Reflection 70B, the top open-source model
twitter.com
twitter.com
Under the hood Reflection 70B seems to be a Llama-3.1 finetune that encourages the model to add <think>, <reflection> and <output> tokens and corresponding phases. This is an evolution of Chain-of-Thought's "think step by step" -- but instead of being a prompting technique, this fine-tune bakes examples of these phases more directly into the model. So the model starts with an initial draft and 'reflects' on it before issuing a final output.
The extra effort spent on tokens, which effectively let the model 'think more' appears to let it defeat prompts which other strong models (4o, 3.5 Sonnet) appear to fumble. So for example, when asked "which is greater 9.11 or 9.9" the Reflection 70b model initially gets the wrong answer, then <reflects> on it, then spits the right output.
Personally, the comparison to Claude and 4o doesn't quite seem apples-to-apples. If you were to have 4o/Claude take multiple rounds to review and reflect on their initial drafts, would we see similar gains? I suspect they would improve massively as well.
You are a world-class AI system, capable of complex reasoning and reflection. Reason through the query inside <thinking> tags, and then provide your final response inside <output> tags. If you detect that you made a mistake in your reasoning at any point, correct yourself inside <reflection> tags.
Also, only "smarter" models can use this flow, according to https://x.com/mattshumer_/status/1831775436420083753They may already implement this technique, we can't know.
But they can do a little.
I have been testing the model for the last few hours and it does seem to be an improvement on LLAMA 3.1 upon which it is based. I have not tried to compare it to Claude or GPT4o because I don't expect a 70b model to outperform models of that class no matter how good it is. I would happy to be wrong though...
[0]: https://github.com/ggerganov/llama.cpp/blob/master/grammars/...
You can somewhat recreate the essence of this using a system prompt with any sufficiently sized model. Here's the prompt I tried for anybody who's interested:
You are an AI assistant designed to provide detailed, step-by-step responses. Your outputs should follow this structure:
1. Begin with a <thinking> section. Everything in this section is invisible to the user.
2. Inside the thinking section:
a. Briefly analyze the question and outline your approach.
b. Present a clear plan of steps to solve the problem.
c. Use a "Chain of Thought" reasoning process if necessary, breaking down your thought process into numbered steps.
3. Include a <reflection> section for each idea where you:
a. Review your reasoning.
b. Check for potential errors or oversights.
c. Confirm or adjust your conclusion if necessary.
4. Be sure to close all reflection sections.
5. Close the thinking section with </thinking>.
6. Provide your final answer in an <output> section.
Always use these tags in your responses. Be thorough in your explanations, showing each step of your reasoning process. Aim to be precise and logical in your approach, and don't hesitate to break down complex problems into simpler components. Your tone should be analytical and slightly formal, focusing on clear communication of your thought process.
Remember: Both <thinking> and <reflection> MUST be tags and must be closed at their conclusion
Make sure all <tags> are on separate lines with no other text. Do not include other text on a line containing a tag.The system prompt used for training this model is:
You are a world-class AI system, capable of complex reasoning and reflection. Reason through the query inside <thinking> tags, and then provide your final response inside <output> tags. If you detect that you made a mistake in your reasoning at any point, correct yourself inside <reflection> tags.
from: https://huggingface.co/mattshumer/Reflection-70BThe only real difference between artificial intelligence and artificial consciousness is self-awareness through self-supervision. Basically the more transparent that AI becomes, and the more able it is to analyze its thoughts and iterate until arriving at a solution, the more it will become like us.
Although we're still left with the problem that the only observer we can prove exists is ourself, if we can even do that. Which is only a trap within a single time/reality ethos.
We could have AGI right now today by building a swarm of LLMs learning from each other's outputs and evolving together. Roughly the scale of a small mammalian brain running a minimalist LLM per cell. Right now I feel that too much GPU power is spent on training. Had we gone with a different architecture (like the one I've wanted since the 90s and went to college for but never manifested) with highly multicore (1000 to 1 million+) CPUs with local memories running the dozen major AI models including genetic algorithms, I believe that AGI would have already come about organically. Because if we had thousands of hobbyists running that architecture in their parents' basements, something like SETI@home, the overwhelming computer power would have made space for Ray Kurzweil's predictions.
Instead we got billionaires and the coming corporate AI tech dystopia:
https://www.pcmag.com/news/musks-xai-supercomputer-goes-onli...
Promoting self-actualization and UBI to overcome wealth inequality and deliver the age of spiritual machines and the New Age are all aspects of the same challenge, and I believe that it will be solved by 2030, certainly no later than 2040. What derails it won't be a technological hurdle, but the political coopting of the human spirit through othering, artificial scarcity, perpetual war, etc.
Also could replace "invisible" with wrap section with "---IGNORE---" or with "```IGNORE" markdown tags and then filter it out after
Personally I strongly disapprove of the first/second person pronouns and allowing them [encouraging, even] to output 'we' when talking about humans.
[0] https://arstechnica.com/information-technology/2023/09/telli...
That said, I'm withholding judgment on how likely the claims are. A friend who developed NoCha [1] is running the model on that benchmark, which will really stress test its ability to reason over full novels. I'll reserve judgement until then.
And with fine tuning, there's zero math needed, it's a bit of common sense, and a lot's of data optimization.
Please do update us on the result.
Here's the updates to the model config on huggingface:
https://huggingface.co/mattshumer/Reflection-Llama-3.1-70B/c...
https://www.wolfram.com/llm-benchmarking-project/
https://help.kagi.com/kagi/ai/llm-benchmark.html
Edit : There are few other benchmarks that give pretty low scores (<20%) to top LLMs. Can't find them atm. There was a benchmark with common sense easy looking questions.
Edit: found two more papers
https://arxiv.org/html/2405.19616
https://arxiv.org/html/2406.02061v1
Edit: How about Wordle?
Only LLAMA 3 makes the justification that only 2 horses can be raced at a time, but then gets its modified question wrong by racing three horses. I personally would consider an answer that presumes some restriction to how the horses can be raced to be valid if it answers the restricted version correctly.
I think this benchmark would really only tell me whether Wolframs book was in the training data.
https://www.wolfram.com/language/elementary-introduction/3rd...
I gave it a medium-complexity design problem: Design the typescript interface for the state of a react app that manages a tree of chat turns/responses and displays the current path through the tree. (In other words, the kind of state that sits logically behind the ChatGPT or Claude Web UI, where previous conversation turns can be edited and used as a branching off point for new turns.)
Reflection-70B suffered from a bad initial idea, just as Llama 70B generally does (proposing to duplicate state between the "tree of all messages" and the "path to currently displayed message"), which is a very common error. The automated reflection process identified a whole bunch of nitpicks but missed the glaring logical bug. Furthermore the final output was missing many of the details included in the initial reflection / chain-of-thought scratchpad, even though the UI hides the scratchpad as though it's unimportant for the user to read.
Playground: https://reflection-playground-production.up.railway.app/
Still impressive that it can beat top models with fine-tuning, but now I’m mostly impressed by the fact that the 70b model was so good to begin with.
Note that there's a threshold for how smart the model has to be to take advantage of this flow (https://x.com/mattshumer_/status/1831775436420083753) - 8B is too dumb.
In which case, what happens if you apply this to a GPT-4o finetune, or to Claude 3.5 Sonnet?
What happens if you combine it with variants of tree-based reasoning? With AlphaProof (https://www.nature.com/articles/s41586-023-06747-5#Sec3)? With MCTSr (https://arxiv.org/abs/2406.07394)?
I'm guessing most of the training data was single-turn, instead of multi-turn, but that should be relatively easy to iterate on.
Alternatively, Twitter links could be rewritten to redirect to one of the few Nitter instances that are still functional.
That limit actually doesn't apply to premium users/bluechecks, and he's using the other features like bold text.
The problem with long posts like that is one, they're annoying to read because when you open one up you don't know how much of a time commitment they will be, and two, you can't reply to just part of them.
I can't keep track of the flailing over at Twitter, especially because I don't have an account. Regardless, it's not all that relevant to what I was saying; maybe I got the reason wrong, but the fact remains that the vast majority of people who I see trying to post longer content on Twitter do it via multiple posts.
As a related aside, it baffles me why people still use the site when many superior alternatives are available.
> The problem with long posts like that is one, they're annoying to read because when you open one up you don't know how much of a time commitment they will be, and two, you can't reply to just part of them.
Those don't actually seem like problems to me.
HN allows, and has always allowed, links to paywalled sources, sources with geographic restrictions that refuse to display the content for some readers, and won't modify a posts URL due to the site being slashdotted / suffering from an HN hug of death. Twitter is no different, except maybe by being more ideologically polarizing.
The place for alternative URLs is, and has always been, the comments.
Seems like most others disagree with me though, so I guess I’ll just skip over anything posted on Twitter.