The boundaries of LLMs’ capabilities are really weird and unpredictable. I was doing a very basic info extraction task a few days ago and none of Gemini 2.0 Flash, Llama 3.3 70B, or GPT-4o could do it reliably without going off the rails. With the same prompt, I switched to the open-weights Gemma 2 27B - released last spring - and it nailed it.