> A matrix can represent a linear map (or, if you prefer, a module homomorphism) but if you think of it as a map from V to W* instead of to W, it represents instead a bilinear form on VxW.
If you mean that it's also a bunch of other things like an element of V* tensor W or a linear form on V tensor W*, and none of these are the most "natural"/better than the others, then I guess that's true. Matrices still seem like the least natural to me though.
For your least squares example, on the face of it it's not linear until you choose to use a linear least squares model :)
It's been a decade or so since I've thought much about this stuff, so I'm not highly confident in this answer, but I think the reasoning goes something like this:
You're trying to solve f(x)=y, but the problem is y is outside of the image of f. The next best thing is to say you at least want <f(x'),f(x)> = <f(x'),y> for all x'. i.e. f(x)=y up to dot products with all possible test vectors in the image of f, which forces their projections onto im(f) to be equal. Then <x',ft(f(x))> = <x',ft(y)> for all x', so ft(f(x))=ft(y), or x = (ft f)^-1 ft y.
It's not obvious to me that this minimizes the square norm of f(x)-y, (or that (ft f) is necessarily invertible) but I guess the point is that that the only nonzero part is in the orthogonal complement of im(f), so you can't do any better.