But overall in my opinion if devs are able to rebuild it from scratch with a predefined outcome, and even know how to improve the system to improve certain aspects of it, we do understand how it works.
But overall in my opinion if devs are able to rebuild it from scratch with a predefined outcome, and even know how to improve the system to improve certain aspects of it, we do understand how it works.
Regarding the car, if you know how to build a car, you understand how a car works. A driver is more like someone using and llm, not a developer able to create an llm.
Sorry to inform everybody doing their Ph.D. on LLM interpretability that they’re just wasting their time.
Yes! loads! (: I want to be able to say statements like "this model will never ask the user to kill themselves" and be confident, but I can't do that today, and we don't know how. Note that we do know how to prove similar statements for regular software.
Common misconception, MoEs do have different "experts", but the model learns when to send input to different experts, and the model does not cleanly send coding tasks to the coding agent, physics tasks to the physics agent, etc. It's quite messy, and not nearly as intepretable as we'd want it to be.