> To understand symbolic programs, LLMs may need to possess the ability to imagine how the corresponding graphics content would look without directly accessing the rendered visual content
To my understanding, LLMs are predictive engines based upon their tokens and embeddings without any ability to "imagine" things.
As such, an LLM might be able to tell you that the following SVG is a black circle because it is in Mozilla documentation[0]:
<svg viewBox="0 0 100 100" xmlns="http://www.w3.org/2000/svg">
<circle cx="50" cy="50" r="50" />
</svg>
However, I highly doubt any LLM could tell you the following is a "Hidden Mickey" or "Mickey Mouse Head Silhouette": <svg viewBox="0 0 175 175" xmlns="http://www.w3.org/2000/svg">
<circle cx="100" cy="100" r="50" />
<circle cx="50" cy="50" r="40" />
<circle cx="150" cy="50" r="40" />
</svg>
- [0] https://developer.mozilla.org/en-US/docs/Web/SVG/Element/cir...