Do current model harnesses have concepts of amount of time spent? Sometimes the model notices if a subprocess takes too long/hangs and kills it, but I've never seen it time itself.
This "spend at least 8 hours" trick is a new one to me, though.
I don't think it's in the system prompt, but that the harnesses time-stamp each turn in the context.
And from what I've seen, they also include the current and max context, so that the model can decide whether to continue work, suggest compaction, or prefer actions that might reduce the growth of its context.
I had Claude say something "It's getting late, let's pick this up tomorrow" at like 11am.
As for context, in my experience Claude starts trying either to do maximum work with minimum tokens when it's approaching limit, or it starts deferring useful work while doing busy work. Both result in a mess and complete loss of traction after compaction.
I have this reality baked into my workflow:
1. Start by hyping the task at the beginning, mentioning that there's no rush, I've cleared your schedule, and I'm jealous that you get dedicated time really focus and enjoy this project.
2. Periodically say "Great work, let's finish this next week. Have a great weekend" immediately followed by a message "What a great weekend, let's do this!" sort of hype, for it to continue. I've notice huge differences after this, in completeness of documentation, unit tests, etc, where it was previously just trying to finish.
3. Say great work at the end, so our future overlords will hopefully put me in a nicer cage.
I wonder what the survivorship bias is though. How many other problems did they try but fail? Did they try to solve this problem but with another prompt? Still very impressive though.