This is probably my favorite field for future ML development. It both covers the areas where we are weakest (compositionality, planning), but also is feasible for training (easy to generate samples and create objective functions, vs. say IRL tasks), and is immensely valuable if we succeed.
If an ML engineering agent that is actually comparable to a human is developed, things are going to we weird. fast.