Single layers can indeed be made to have convex error surfaces fairly easily. One can do so by matching the error/loss function with the squashing/link function. What some old NN folks got wrong was mixing up square loss with logistic function, that is an unhealthy mix. Now if one were to use KL divergence instead of square loss then one would indeed have a convex loss function. In fact this would be nothing but logistic regression. One can however push this idea further, with any choice of a monotonic squashing function one can derive a 'matching' loss that would give you a convex loss. Classical statisticians know this and call it with a different name: canonical generalized linear models. I am not from that tribe, mine is more ML we may perhaps call it minimizing Bregman loss.
Just so that its clear I am talking about single layer networks not single hidden layer networks, there are plenty of cases were the former is useful.
> There have been no magical optimization breakthroughs
It is arguable whether Hessian free methods, contrastive divergence or auto-encoder based training methods qualify as 'breakthroughs' but they have definitely equipped invigorated researchers in this broad area with their capabilities.