It's a bit difficult to explain the visualization without pictures, but I'll give it a shot.
The transpose is really about converting the matrix to operate on a different vector space, namely, the dual space. In particular, the dual space of a vector space V is the vector space of "linear functionals", which are linear functions
\phi: V -> R
A linear functional on R^2 looks like a gradient (the "gradient fill" gradient, not a calculus gradient). These gradients are in one-to-one correspondence to vectors in R^2. In particular, given a vector w \in R^2, the direction of the gradient is along the direction of v, and the speed with which the gradient is changing corresponds to the magnitude of w.
The precise mathematical correspondence is that (i) given a vector w \in R^2, the function
f_w(v) = <w, v>
is a linear function (here, <,> is the inner/dot product), and (ii) every linear function has this form. Now, note that f_w is exactly multiplication by the transpose w^T of w! In particular,
f_w(v) = w^T v
More generally, for any linear map A : V -> W, the adjoint A* of A is defined to be the linear map from the dual space of W to the dual space of V that satisfies
<w, A v> = <A* w, v>
The transpose A^T is the adjoint of A when V is a finite-dimensional real vector space:
<A^T w, v> = (A^T w)^T v = w^T A v = <w, A v>
In summary, you can try to visualize A^T as a linear map "acting" on dual vectors. For example, let v \in R^2 and let w be a dual vector (i.e., a gradient), and suppose that A rotates v clockwise by 90 degrees. To preserve the inner product <w, A v>, A^T rotates w counter-clockwise by 90 degrees.