What are the typical context lengths in SWE-bench problems? Does it partly measure performance in the 64-128k context range?
https://huggingface.co/datasets/princeton-nlp/SWE-bench_Veri...
Its up to your retrieval system/model to selectively hunt for relevant context. Here's a few critiques of the benchy: