That's max output tokens per response limit, separate from context length
It's the output token limit, which has been 128,000 for Claude models for quite a while note
Pretty crazy that the model doesn't know that it needs to stop before it hits 128k output tokens. I guess it has no sense of how many tokens in it is? Wouldn't this be possible to work into the architecture?
I think this is a bug. I've not seen this problem from any of the other frontier models.
Do other models put a hard cap on the output tokens it can generate?