>They use inscrutable internal language of embeddings
No different than the electrical signals in the intermediate neurons in your brain that comprises the latent space where all the processing happens
>They communicate in natural language which is itself ambiguous
They lack one-shot precision, sure, but it doesn't matter. They are precise enough with refinement over multiple prompts.
>the weights and training inputs are being hidden as a "trade secret"
For frontier models that make the company money through api pricing, sure. There are plenty of open source models that can be used for the same tasks, which have open weights.
>This doesn't really mean much unless we understand what is the quality and relevance of these sources for the problem at hand.
All of the modern models are RL trained on specific tasks when it comes to coding. I.e the initial training run learns to predict the next token based on context, from all the available texts, but then the RL runs specifically train the model in a harness where it produces code and RLed to produce correct code with specific formatting.