> Progress is still tracking the steep part of the S-curve, and there's no indication that they're near the top yet.
If I understand it correctly, that seems to be exactly what this paper is suggesting.
Scoring high on the bar exam is pretty trivial for AI - the data needed for that is fairly generic and widely available on the internet. It requires you to demonstrate a relatively basic understanding of the concepts by answering a bunch of multiple-choice answers. If anything, I'd expect AI to have a perfect score.
Like I said, such an AI is not entirely useless. You can replace quite a few legal assistants with that, and I bet it could be used to create first drafts or to expand a core concept into a full legal argument. There is plenty of money to be made there, and it's going to make an awful lot of people jobless. But that's just replacing more-trivial jobs with automation, it doesn't add anything novel to society.
On the other hand, the actual difficult work involves being able to come up with completely novel concepts, and being able to expand upon some obscure but crucial stuff few people have ever heard about. Current models simply aren't capable of that, and the results achieved here with multimodal models suggests that they never will. We risk getting stuck with models which can do some trivial work, but silently produce complete garbage when you ask them to do anything providing substantial value.