The reason for exp(x) is that its derivative is exp(x), which makes it possible to express the gradient of s(x) in terms of s(x), or both in terms of exp(x). This simplifies the computation of backward pass.
Luckily, there are more axiomatic reasons for why softmax is the preferred way to map inputs to a probability distribution.