I ran into the exact same wall when doing some SIMD + OpenMP optimisations.
Me: "Parallelise this loop without changing the results."
AI: "I can't do that, because of floating point accumulation."
Me: "So spin out a temporary fixed-length array, accumulate into that thread-locally, and then do an ordered summation of that at the end."
AI: "Oh, you're right!"
It's a fantastic tool for automating the implementation grind, but you still have to hand-feed it the strategy. I guess it will get there at some point.