I think that's probably why everyone in this thread has such different experiences - someone whose workflow is mostly asking a model for code and pasting it in would have seen modest improvement and would reasonably wonder what the fuss is about whereas someone who was already running agents on 20-step loops would have felt a much bigger shift, because the thing that used to kill those runs was the failure at step 12 cascading into garbage by step 20, and that got a lot better.
The local model story Simon kind of glosses over is interesting for the same reason - a 20GB model drawing a decent pelican on a laptop is a cute data point in isolation. The thing worth noticing is that a competent local model inside a good harness now gets you closer to frontier performance than running the frontier model without a harness does.