This discussion is so dumb - finetuning a base model costs ~$1 with LORA/QLORA and can yield same performance as gpt-4, but at 1/100 of the cost per token.
What Bloomberg did for $10M was not finetuning..
What Bloomberg did for $10M was not finetuning..
That's a big claim - can you back that up with any examples?