The goal of most/all machine learning is hunting down performance maxima in a super high dimensional space. In this case the space is the positions and angles etc of the atoms that comprise a protein and the "performance" is the stability of the protein.
For the number of atoms that comprise most proteins, iterating through all the possible positions would take an unimaginable amount of time, so you have to have some kind of search method to identify good position-space-areas to investigate more closely.
My guess as to where people help in is getting away from bad local maxima. In my experience playing foldit, sometimes you can see pretty clearly that the stability is not good and it's not going to get much better with small changes - the algorithm has found a bad local maxima of performance - so you can manually move big chunks of the protein around to explore a new part of the position-space. This kind of evaluation, knowing when to stop climbing a small hill and instead go looking for bigger hills, seems to be something that humans are pretty decent at. Of course there's also a million algorithms to do the same thing without humans.