Results for GPT-6 Astra and Gemini Flash 3.8 are just in!
GPT-6 claimed the first spot with a score of 69.3.
Gemini 3.8 Flash on a very solid 5th spot with 55.4.
In the benchmark, have you considered instructing the models to build their own SPICE simulations to test their work? Simply asking them to write and run simulations could improve performance, even without telling them what to simulate.