64 karma · joined April 4, 2025
Even if LLMs aren’t currently up to snuff, I don’t see any deceleration yet in the s-curve. Speed, intelligence, and persistence are all improving every month. I agree that end-to-end prompting “e.g. build me app X” is not production worthy, but i think it’s pretty bold to say coding isn’t solved. AI can code almost anything you adequately specify. The specification is the engineering (it’s a process, not one prompt).
Most errors are probably responses that didn’t finish before their 3K token limit. They’ve measured how well RL is able to shorten the response to their limit.
In a similar way that Google Maps shows eco routes, it’d be fun for them to show “safest” routes which avoid areas with common crashes. (Not always possible, but valuable knowledge when it is.)
For example, only 7% of pharmaceutical research is publicly accessible without paying. See https://pmc.ncbi.nlm.nih.gov/articles/PMC7048123/
The progression from basic arithmetic, to complex ratios and basic algebra, graphing, geometry, trig, calculus, linear algebra, differential equations… all along the way, there are calculators that can help students (wolfram alpha basically). When they get to theory, proofs, etc… historically, thats where the calculator ended, but now there’s LLMs… it feels like the levels of abstractions without a “calculator” are running out.
The compiler was the “calculator” abstraction of programming, and it seems like the high-level languages now have LLMs to convert NLP to code as a sort of compiler. Especially with the explicitly stated goal of LLM companies to create the “software singularity”, I’d be interested to hear the rationale for abstractions in CS which will remain off limits to LLMs.
- Bad Mental Health: At the start of the war in Ukraine I read/listened to the news every day. I’d frequently cry, hearing an interview from someone who was trapped in a bombed out building etc. After a few weeks I realized I was being emotionally exhausted by something around the world about which I had no control.
- Enshitification: Working for a b2b “tech/coding education web app” company as a data scientist and realizing the perversity of incentives which were ruining the product.
- Increasing Opportunity Cost: Working on LLMs and realizing that the possibilities for what I could do were expanding because I would now have more answers and information at my fingertips than ever before.
Great post… it was useful to be able to reflect and understand my experiences with these abstractions.
An interesting problem since the creators of OLMO have mentioned that throughout training, they use 1/3 or their compute just doing evaluations.
Edit:
One nice thing about the “critic” approach is that the restaurant (or model provider) doesn’t have access to the benchmark to quasi-directly optimize against.