Context went from 8,192 tokens on GPT 4 to 1M tokens currently with zero improvement. The latest models got even worse.
> The latest models got even worse.
Which models? This is one of those things that likely has both model and domain specific aspects that impact your experience. In my experience with OpenaAI models predominantly (I previously worked there), they've improved significantly over the last 6-12 months. My experience with Claude is worse, but I haven't spent as much time getting into a mechanical sympathy there. They're still not perfect though and I have many steering docs that help avoid the biggest problems in the models I use when generating docs.