Yes but you can use an llm to label data and then train a bert model which then costs a small fraction of time and money to run than the original llm.
Is this because the final represention in bert style models more globally focused, rather than being optimized for next token prediction?