My current AI chat benchmark is to ask about LLVM APIs. I’ve asked to both Bing Chat and Bard three questions:
- What info can the LLVM lazy-value-info analysis pass analyze?
- Can you give me code samples to use the info?
- Can you give me a list on important LVI methods?
Both fared well on the first question, but on the second question, Bard either didn’t produce LLVM code at all (showing examples of LLVM IR), or hallucinate non-existent methods not even close to what existed.
Bing chat, in comparison produced correct answers that are very helpful. In my experience Bing chat almost always produces something that is both coherent and useful; Asking LLVM questions, Bing chat can search the Doxygen documentation, find out common LLVM patterns on its own (like using a Worklist while iterating), and I’ve succeeded in writing whole LLVM passes that compiled in one go, from a very light description of an algorithm.
Maybe I’m spoiled with Bing chat, but I don’t think I’ll be using Bard that much if it’s quality doesn’t improve. Very disappointed personally. (I’d have expected Bard to work much better mostly just because Google searches better than Bing.)