Why so many updating in LLM so-called SOTA let remind of iPhone 4-x
medium.com
medium.com
HumanEval done, SWEBench done... Since it will lasting on and on, but task completion of LLM remain hard to be achieved esp. for high complexity task?
Just investor and Big Three of LLM fool of ppl?