At scale, if something can go wrong with a model, it will. Earlier this year, we ran into a strange issue. On certain inputs, Gemini 3 Flash would loop in its own reasoning, burn 16k tokens each turn, and produce nothing. The failure was deterministic, so retries just multiplied the cost. We got 92% recovery by telling the model what happened and asking it to try again.