Did you use the same pass@1 generation method as in the code llama paper (greedy decoding)? I couldn't find this in the blog post.
Edit: it could also be misleading to directly compare humaneval pass@1 against codellama without the same generation methodology. (possibly against GPT-4, also, but I don't know their methodology).