In a nutshell, you learn a vector of real-valued parameters for each word in your vocabulary. To train a network on sequences of words, you represent said sequence as a concatenation of the vectors of these words, and feed it as an input to the network.
To learn these vectors, you define the problem of "language modeling" as that of discriminating between two sets of sequences: S1 and S2. S1 is the set of sequences that occur in Wikipedia (of which there are many) and S2 is the same as S1, but where you replace a word in each sequence with a randomly chosen word from your vocabulary (which makes it, with very high probability, an invalid sequence of words).
Basically, by learning to discriminative between "good" and "bad" English word sequences, you can learn a language model of sorts. The model is represented by those vectors for each word.
You can then project those vectors into 2D, as bravura did a while ago, and look at what is close to each other: http://www.cs.toronto.edu/~hinton/turian.png