Author here. We did not pre-programm basic control or feedback parameters. The trained policy (fully connected neural network) takes the state estimate (position, orientation, linear and angular velocity) and a history of the previous actions as the input and directly outputs the PWM/RPM setpoint (uint16 that sets the pwm interval).
The code is open source, you can verify it yourself ;)