How is it improving, that would require rearranging its weights and biases which it cannot do easily or quickly.
With the level of compute they have they aren't stuck with frozen models like you are.
The llms are supervising rlhf and creating synthetic data (to an extent) but they're nowhere close to being able to operate the full training stack end to end. This is a fantasy being sold to investors to create fomo.
Remember they're also limited by an effective memory of like 500k words a turn. Memory systems are lossy, so are swarm/sub agent mechanism. Im not worried about llms becoming self powered super entities anytime soon.