Maybe the fact that moment conditions are often first order conditions for some agent's optimization problem is why rationality was mentioned, but I'm just speculating.
Micro models estimated from agent's optimization problems are more frequently estimated using MLE (John Rust's GMC bus paper being a canonical example, though still true in the current literature e.g. Nevo, or Bajari and Hong).
Micro models estimated from aggregated data frequently lack agent-level optimization, and those are the models more frequently estimated with gmm (examples here would include the hundreds of papers based on Berry, Levinsohn and Pakes).
So, this explanation doesn't seem to hold in micro contexts.
Though, as the previous commenter pointed out, macro is quite a bit different, and your explanation is probably correct there.