I know people complain about hardware and compute resources. But, this is like complaining about Python resource usage in the early 90's. The development complexity & resources is far more expensive than chips on the long run. I am personally re-organizing my AI organization to move away from complex RAG setups and get comfortable with long-context workflows.
Just to be clear - I also think that inference-optimized chips are the next frontier - Nvidia GPUs were designed & built in a different age than what’s going on now.