It's the economically viable part that I think is currently hard, but agreed this is the right approach.
The basic problem is that scaling up understanding over a large dataset requires scaling the application of an LLM and tokens are expensive.
The basic problem is that scaling up understanding over a large dataset requires scaling the application of an LLM and tokens are expensive.