If we had unlimited memory, compute and data we'd use a rank N tensor for an input of length N and call it a day.
Unfortunately N^N grows rather fast and we have to do all sorts of interesting engineering to make ML calculations complete before the heat death of the universe.
You are assuming you can match Gemini's performance, Google's engineering resources and costs being constant in to the future.
I'm not assuming. We already did, 18 months ago with better performance than the current generation of Gemini for our use case.
You're falling into the usual trap of thinking that because big tech spends big money it gets big results. It doesn't. To quote a friend who was a manager at google "If only I could get my team of 100 to be as productive as my first team of three.".
That is you'd need 5 exa yotta bytes to solve it.
Currently the whole world has around 200 zetabytes of storage.
I short for the next 120 years mnist will need mathematical tricks to be solved.
Its more about the information about the specific problem you are solving having less impact than techniques that target the compute. So in this case, breaking down how to parse a PDF in stages for your domain is involving specific expert knowledge of the domain, but training with attention is about efficient use of compute in general; with no domain expertise.