Sure, you can evaluate them purely as optimisation algorithms, but does it follow that the better optimisation algorithm is necessarily better at picking hyperparameters that generalise to unseen data?
One way that hyperparameter optimisation can overfit that people don't always think about, is by repeatedly evaluating high-variance metrics and picking the best of N tries. This has burned me when it comes to optimising settings for stochastic optimisation algorithms for example. An algorithm that was very aggressive in doing this might reach a better maximum on the validation set but wouldn't do any better on held-out data.
There are things you can do to compensate for that of course (variance estimates for metrics is a good idea!), but evaluating on a test set data usually doesn't hurt and seems like the safest option.