HN
Hacker News
Top
New
Best
Ask
Show
Jobs
Comment by sauwan | Hacker News Reader
Parent
Full thread
sauwan
·
I'm assuming higher is better? How does this compare to GPT-3.5/4?
View on HN
qwerty3344
·
I think it's still significantly behind GPT 3.5/4, both of which can get 67% on HumanEval, and 88% with Reflexion
enum
·
Keep in mind that StarCoder(Base) is just a pretrained LM. The extra stuff that makes 3.5/4 like RLHF gets built on this.
manojlds
·
Aren't GPT-3 etc base LM and ChatGPT the instruction tuned? Or am I wrong?
dpf
·
code-davinci-002 is a base LM, and the other 3.5 models (text-davinci-{002,003}, gpt-3.5-turbo, and ChatGPT) use instruction tuning and/or RLHF. Source:
https://platform.openai.com/docs/model-index-for-researchers
qwerty3344
·
https://newatlas.com/technology/gpt-4-reflexion/
Reply on news.ycombinator.com