>Since you have a literal dependency between one data point and the next, you can't train your model using randomized data
This is why the train/test split was not randomized, but sequential.
This is an out of sample result. 4 formulas were compared. Sure it may be an overfit, as any machine learning model can be, but the procedure that was used to generate this formula cannot be so easily dismissed.
As a rule of thumb, you cannot. I have tried to fit the Mona Lisa with the software (brighness as a function of x and y) and could not find anything. It performs poorly when no structure exists.
This is a Pareto optimization with a limit on the size of the formulas. Sure, with a formula of arbitrary size you can fit anything with 100% accuracy, but the task is much harder if what you are looking for are short models.
This is out of sample performance. Only 4 models were compared outside the train domain. Also I don't see where you see 22 parameters, the complexity of a model is not defined as its number of parameters.