APPS has 3 subsets by difficulty level: introductory, interview, and competition. It isn't clear which subset Claude 3 was benchmarked on. Even if it is just "introductory" it is still pretty good, but it would be good to know.
(There are 1000/3000/1000 problems in the test set in each level).
It’d be great if someone from Anthropic provides an answer though.