Step 1. Model can’t do something challenging Step 2. You try a bunch and fail Step 3. Anthropic trains on your usage data. Your current code base and current problem are now in domain Step 4. Model comes out and you’re shocked when it can tackle the thing you were stuck on Step 4. Codebase drifts significantly and you try new problems you thought were a similar level. Your code is less familiar and the problem doesn’t have a bunch of failure cases in the train set. Feels of it being worse on similar problems