We know what calculations it does. We have some hazy idea of some bits of how those calculations lead to something that at least somewhat resembles intelligent behaviour. But that's a far cry from actually knowing how it works.
For instance, suppose you give one of today's frontier models some of those chain-of-cubes rotation puzzles (the sort that infamously men are about 1sd better at than women, statistically speaking). How well will it do? I have absolutely no idea and I'm quite sure that a more detailed understanding of the transformer architecture would not make my guesses any better. (Actually, I do kinda have some guesses but they're based on a vague notion about how the models might be partitioned between vision-y bits and language-y bits, and it's very possible that that notion is out of date.)
> it does exactly what we expect it to do
Were you, let's say 6 months ago, expecting it to resolve one of the Millennium Prize problems?
(I do agree that it is more productive to ask "what can and can't they do?" than "should we classify that as intelligent or not?".)
> for Navier-Stokes they spent in 3 days more money than the whole mathematical community over the last 20 years easily.
Are you sure?
(The numbers I've heard, which I admittedly have no very strong reason to trust, don't seem that way to me.)