Sol is closer to Fable than Opus - I like SlopCodeBench the most as a test - https://github.com/humanlayer/advanced-context-engineering-f...
You add requirements and make previous tests invisible to see how pigeon brained the model is - Sol and Fable seem to rank the same as Opus tends to fall behind