AI safety/alignment discourse as a whole is so incredibly bad. It's autodidactic and weird in the bad sense, weird like people who think they are brave iconoclasts but just haven't the curiosity and humility to learn the basics of the field are weird. Add some money from Dustin Moskovitz to Open Philanthropy to Robert Miles' propaganda firehose on youtube, add policy lobbying and buying bankrupt crypto podcasters, and you have the worst that could come from the culture of public intellectualism. The notions they have inserted into this potentially vital field, like inner/outer misalignment and "mesaoptimization" [2] (where simple overfitting is the real and more productive explanation) are deeply misleading, technically illiterate and stoking public fears in directions we should care about the least, siphoning resources and attention from real issues informed by specifics of contemporary ML (chiefly data and curriculum engineering). They're close to putting a lid on LLM development – because they're getting spooked by some vague sense of growing "capabilities"; even though LLMs are evidently a vastly better approach to safe-yet-capable-general-AI than anything any of their thought leaders have ever come up with in decades of SIAI/MIRI/Lesswrong procrastination. (For one example of a decent takedown of their approach I recommend [3]).
Why is this algorithm bizarre, or even inhuman? Could anyone show me the specific functional connectivity graph and activation functions in my brain that I use when doing addition, with the circuit going through spatial-quantitative cortex and phonological loop, populations querying cached results in long-term memory, individual numbers marked as interesting or banal by the highest-level associative units? Would it look anything like a neat and sensible algorithm we'd come up with using a global top-down representation of the solution target space?
By this standard, humans are inhuman, humans are shoggoths, as they should be, because emergent data-driven structures in general-purpose substrate generally do not resolve to anything like provable optimality, even if they get very close in performance. This framing adds nothing but clicks and insecurity.
1. https://twitter.com/robertskmiles/status/1663534255249453056
2. https://www.youtube.com/watch?v=zkbPdEHEyEI
3. https://www.lesswrong.com/posts/wAczufCpMdaamF9fy/my-objecti...