What's the relevance of the pelican benchmark when models probably saw it during training? Didn't OpenAI stop testing against SWE-Something because it was tainted?
That aside, the relevance these days is in comparing models and effort levels within the same model families - hence the comparison grids.
If the dialogue is slop and not like the old memes then it fails.