A top expert in US Trust & Estate Tax law whom I know well tells me that although their firm is pushing use of LLMs, and they are useful for some things, there are serious limitations.
In the world of T&E law, there are a lot of mediocre (to be kind) attorneys who claim expertise but are very bad at it (causing a lot of work for the more serious firms and a lot of costs & losses for the intended heirs). They often write papers for marketing themselves as experts, so the internet is flooded with many papers giving advice that is exactly wrong and much more that is wrong in more subtle ways that will blow up decades later.
If an LLM could reason, it would be able to sort out the wrong nonsense from the real expertise by applying reason, e.g., comparing the advice to the actual legal code and precedent-setting rulings, and by comparing it to results, and be able to identify the real experts, and generate output based on the writings of the real experts only.
However, LLMs show zero sign of any similar reasoning. They simply output something resembling the average of all the dreck of the mediocre-minus attorneys posting blogs.
I'm not saying this could not be fixed by Altman et. al. applying a large amount of computer power to exactly the loops I described above (check legal advice against the actual code and judges' rulings, check against actual results, select only the credible sources and retrain), but it is obviously no where near that yet.
The big problem, is that this is only obvious to a top expert in the field who deeply knows from training and experience the difference between the top experts and the dreck.
To the rest of us who actually need the advice, the LLMs sound great.
Very smart parrot, but still dumbly averaging and stochastic.