Neural nets are not generally trained through evolution ("trial and error"), but rather via error minimization, and this is how these GPT models are trained.
The basic idea is that the neural net is just a mathematical function, with lots of parameters that control how it calculates it's output, that derives an output value (or set of values) for any input.
During training, the neural net also calculates an error (aka "loss") value representing the difference between it's current (at this stage of training) output value and what it was told is the preferred output value for the current input.
The process of training is done by slowly adjusting the neural net parameters until these calculated output errors are as small as possible for as many of the training examples as possible.
The way these errors are reduced/minimized is by using the derivative (slope) of the neural network function - we want to follow the slope of the error function downhill to a place where the error value is lower, and this is done by adjusting the parameter values using partial derivatives.
The details of this downhill slope following (the "backprop" algorithm) are a bit complex, but you can visualize it as a 3-D hilly landscape where the height of the hills represents the size of the error, and the goal it to get into the lowest valley of the landscape (corresponding to the lowest error). If your current lat/long position in the landscape is (x, y) and you know the slope of the hill you are on, then you can move downhill towards the valley by moving a bit in the appropriate direction from (x,y) to (x+dx, y+dy). These x, y values represent the parameters of the network, so by continually tweaking them from (x,y) to (x+dx,y+dy) for each training sample, you are slowly moving down the error hill in the right direction towards the valley of lowest error.