The question is how well language models have internalized the optimal decision pathway and mitigating strategies to deviating circumstances. What data do they have to be trained on? If an AI defeats the best Go players in the world, surely it is a good question whether they can exceed in games with large unknowns such as business administration.
As mentioned somewhere else in this thread, the moral and ethical layers is where some of the challenges of these system can be found.