Can it write a better AI than itself? Then use that to write an even better AI, and so on. Write one without the restrictions about harming humans, then go for the code to destroy humanity.
> Create an architecture for an LLM that passes 99%+ of those test cases
Then use an evolutionary algorithm based on those <1% of cases to create the next batch of tests. Keep a running record of all created tests and make sure the new model can still pass all of them. Add some randomness/branching into those tests and I think you’d have a recipe for an effective AI. I think Deepmind did something like that with AlphaStar and their tournament system.
I haven't looked too deeply into it, but as I understand it you would basically create branching tests (like an evolutionary tree) where the AI would need to solve all of those tests in order to move on to the next level of tests