What are the chances that CursorBench ranks Cursor's model highest, and Cognition's bench ranks Cognition's model highest? Both are to be RL'd from Kimi as a base model, BTW.
I'd posit that it's not deliberate deception, but for both companies their training data and benchmarks come from the same dataset (Devin/Cursor interaction logs) so they naturally overfit.