"understanding" is overstating it. Correlation between tokens embedded in the weights via training, yes.
Agent reads a skill file about how to use a CLI tool. It tries to use the tool but gets an error about the input format. It tries again with a different format based on the error message, and sees that command succeeded. It compares what worked to what was in the skill file and notes the difference. On future invocations it continues to use the new format.
Is that not "understanding" how to use the tool?
They train on a billion "jobs". Which is not terribly efficient but oh man they do train.
This fact is currently the most limiting factor for LLMs.