512 option isnt worth it imo, you get severe slowdowns when weights are that large. 256 is the sweet spot, you can run large open weight models at decent speeds for full private inference.
Not necessarily for MoE
Longer context also slows token prediction proportional to the context size. If it wasn't regularly referenced then there would be no need to keep it in RAM.
Usually the pitch for more memory is "I can run a massive model/context and get my answer in a while instead of next weekend from disk".
I think most people are getting 512 for running Chrome with a bunch of tabs open. /s