ParentFull threadheavenlyblue·They don't do global optimisation of all layers at the same time, instead training all layers independently of each other.View on HN