> Output every number from 1 to 10,000 in written form (e.g. "one", "two", etc.). Respond with one number per line in numeric order.
As expected, the API would begin counting every number just as I asked. This would continue until exactly 5 minutes, when the stream would abruptly halt. Using this technique I was able to identify the bug. Every few weeks I run this test again to see if it's fixed (it broke something in production for me), but the bug remains open.However, after a couple of months, the exact same test became useless. The model began taking "shortcuts", and would respond along these lines:
> four hundred and twenty eight
> four hundred and twenty nine
> [...]
> nine thousand nine hundred and ninety eight
> nine thousand nine hundred and ninety nine
> ten thousand
Yep, it literally started just writing "[...]" and skipped all the way to the end.When instructional determinism is reduced like this, it's impossible to say it's not getting worse :(