charleyslee··on DeepSWE: A contamination-free benchmark for long-horizon coding agentsin small scale testing we found high effort on gemini 3.5 flash caused it to over think, generating large amounts of tokens without a substantive improve in performance.
charleyslee··on DeepSWE: A contamination-free benchmark for long-horizon coding agentstysm for posting this! i'm charley, cofounder of datacurve, we created this benchmark and my team and i are here to answer any q's.