I won't be surprised, but that's less than "can expect", and I disagree that this is straightforward to forecast… unless you're currently playing with another similar model that hasn't been published and which can do this. As the saying goes, "forecasting is hard, especially if it's about the future".
AI progress has always been this weird combination of two sides, one saying for every breakthrough "this is just around the corner", the other saying "this is impossible".
This even happens anachronistically, with some people convinced AI can already do things they can't, and others that they could never do things they already do.
What's more common though is that even though it gets good on that task, it may not translate to practical real world application.
The best example I can think of is object detection vs self driving: lot of people though that the improvements in object detection on images will easily translate to great self driving, and here we are, still with cars not stopping when a car is blinking in front of it.
> hyperparameter optimization and prompt engineering
Prompt engineering seems a lot like "tweaking the question format until the AI gets the answer" which is the first lesson in ANNs; don't train on your test set. If you can't put the actual question with the only context being "this question concerns US law" then there's an awful lot of reasoning and thinking that the human is doing which the AI cannot.
Let's not have another Moore's law fallacy, it's not reasonable to extrapolate progress based on existing results. It's like building a car which can go at 200mph then saying 300mph is just around the corner.