Consider a linear transformation from Rn to Rm. i.e. T[x1,...,xm] = [y1,...,yn].
Let's simplify the left-hand side with what we know about linear transformations: T[x1,...,xn] = T[x1,0,...0] + ... + T[0,...0,xn] (by 1). And T[x1,...,xn] = x1T[1,0,...,0] + ... + xnT[0,...,0,1] (by 2)
We don't know what T[1,0,...,0],...,T[0,...,0,1] will be (cause we're generalizing for any linear transformation) so let's just assume the values will be [c11,c12,...,c1n],...,[cm1,cm2,...,cmn].
So now we have: [y1,...,yn] = x1[c11,c12,...,c1n] + ... + xn[cm1,cm2,...,cmn]
And using the properties of linear transformations we can revert it back to the original form: [y1, ..., yn] = [x1c11 + x2c12 + ... + xnc1n, ..., x1cm1 + x2cm2 + ... + xncmn]
So no matter the linear transformation, all entries yi of the vector you output will be equal to a linear combination of x1,...,xn. So if all linear transformations satisfy this condition, what makes a linear transformation different from the others? It will be the set of coefficients [c11,c12,...,c1n],...,[cm1,cm2,...,cmn] that multiply with the x values.
This set of coefficients is unique to every linear transformation from m to n (it's what defines it), this means that we can set up a one-to-one correspondence between these coefficients [c11,c12,...,c1n],...,[cm1,cm2,...,cmn] and any linear transformation (from m to n)!
So let's make it easier for ourselves to describe every single linear transformation from m to n by setting up a rectangular array of the coefficients with m rows and n columns: [[c11,c12,...,c1n],...,[cm1,cm2,...,cmn]].
And so if we find a matrix [[A,B],[C,D]] then we immediately know it corresponds to the linear transformation [y1, y2] = [Ax + By, Cx + Dy]
Note: You won't find a one-to-one correspondence between a matrix and a non-linear transformation with this method because non-linear transformations could have terms like xy, x^a, sin(x), etc... that can't be captured by a linear combination.