Basically what you end up doing is discretizing spacetime on some grid, maybe e.g. 100x100x100x100 (not much larger than that last I checked or the problem becomes intractable even on very large computers) and then sample field configurations monte-carlo style (the way to do it is actually really clever - it turns out that there is a mathematical trick that can take you from the description of your theory right to a distribution from which you can draw your field configurations). The number of fields is decently large (some discretization introduce extra auxiliary fields for every physical one) and in general you need to consider the entire field configuration to compute an observable. Then do that a few billion times (to bring down statistical error) at a couple different lattice spacings (to be able to extrapolate to the continuum limit) and you quickly end up with a very large computations.
There is a number of clever ways to reduce the amount of computation required, but in general it's a very different problem.
Hope that makes some sense. I had looked into this in detail about two years ago, but looking back, I feel like I only know every other word in my write-up ;).
A few minor things:
People typically use lattice sizes that have lots of factors of 2, because they fit on the supercomputers in a better way. So rather than 100 people do 96 or 128. Only some calculations require lattices this large.
The field content is usually a few fermions, and the gluons. On each site there is (usually, depends on the discretization as you say) a Dirac spinor (2 spins * 2 particle/antiparticle * 3 spins = 12 numbers) for each fermion and on each link (ie. the edges) there is an SU(3) matrix (3x3 complex doubles) for the glue.
The field configuration is simply the glue. But this already is a large amount of data. volume * (9 complex doubles per link) * (1 link per spacetime direction) * (4D spacetime) ~= 9 gigs for a (32^3 * 64 spacetime volume). This is effectively a giant (sparse) matrix, which we want to invert.
The mathematical trick is called importance sampling, https://en.wikipedia.org/wiki/Importance_sampling which people accomplish with the metropolis algorithm https://en.wikipedia.org/wiki/Metropolis%E2%80%93Hastings_al... . If only we could do a few billion samples! People usually make a few thousand to a few hundred thousand samples.
Gauge symmetry is a local symmetry that dictates a lot of the structure of QCD. If your discretization breaks it (and say, only recovers it in the continuum limit) you will have very bad renormalization and fine-tuning problems. Putting the gauge on the links automatically guarantees that gauge symmetry is respected, even at the discretized level, and jibes very well with the fact that the gauge field / gauge connection describes, in some way, how to perform parallel transport. So, it's very natural for the gauge to connect the sites together in that you need to know the value of the link in order to compare quantities on two adjacent spacetime sites.
https://en.wikipedia.org/wiki/Gauge_theory https://en.wikipedia.org/wiki/Connection_form https://en.wikipedia.org/wiki/Parallel_transport
Feynman actually wrote this pretty amazing paper on the topic of simulating quantum mechanical systems with computers: http://link.springer.com/article/10.1007/BF02650179. You should have a look, it's quite readable!
(If the paywall is a problem for you, you should be able to find the paper elsewhere)
There are only 3 variables in a discrete logarithm problem but it takes super computers to solve. Same with prime number factorization.
Number of inputs is completely unrelated to the time to solve depending on the algorithm being used to solve the problem.