PaLM 2 on HumanEval coding benchmark (0 shot):
37.6% success
GPT-4:
67% success
Not even close, gpt4 miles ahead
37.6% success
GPT-4:
67% success
Not even close, gpt4 miles ahead
[0] https://platform.openai.com/docs/model-index-for-researchers
In the GPT-4 technical report, they reported contamination of humaneval data in the training data.
They did measure against a "non-contaminated" training set but no idea if that can still be trusted.
HellaSwag: GPT-4: 95.3%, PaLM 2-L: 86.8%
MMLU: GPT-4: 86.4%, Flan-PaLM 2-L: 81.2%
ARC: GPT-4: 96.3%, PaLM 2-L: 89.7%
(from: GPT-4 paper: https://arxiv.org/pdf/2303.08774.pdf)