> To understand symbolic programs, LLMs may need to possess the ability to imagine how the corresponding graphics content would look without directly accessing the rendered visual content
To my understanding, LLMs are predictive engines based upon their tokens and embeddings without any ability to "imagine" things.
As such, an LLM might be able to tell you that the following SVG is a black circle because it is in Mozilla documentation[0]:
<svg viewBox="0 0 100 100" xmlns="http://www.w3.org/2000/svg">
<circle cx="50" cy="50" r="50" />
</svg>
However, I highly doubt any LLM could tell you the following is a "Hidden Mickey" or "Mickey Mouse Head Silhouette": <svg viewBox="0 0 175 175" xmlns="http://www.w3.org/2000/svg">
<circle cx="100" cy="100" r="50" />
<circle cx="50" cy="50" r="40" />
<circle cx="150" cy="50" r="40" />
</svg>
- [0] https://developer.mozilla.org/en-US/docs/Web/SVG/Element/cir...I never understand why people don't just test this?
Claude 3.5 sonnet, first test, prompt: Any idea what this SVG is describing?:
This SVG describes a simple geometric composition consisting of three circles: 1. The largest circle is centered at coordinates (100, 100) with a radius of 50 units. 2. A smaller circle is positioned at the top-left, centered at (50, 50) with a radius of 40 units. 3. Another smaller circle is positioned at the top-right, centered at (150, 50), also with a radius of 40 units. The overall effect, if rendered, would resemble a basic face-like structure or a simplified Mickey Mouse silhouette, with two smaller circles (potentially representing ears or eyes) above a larger central circle (potentially representing a face or head). The SVG uses a viewBox of "0 0 175 175", which defines the coordinate system and aspect ratio for the image. Would you like me to create a visual representation of this SVG for you?
I never understand why people take things out of context instead of responding to the totality of a message, since I started off by saying:
> To my understanding,
You have responded not by providing your understanding but by belittling me. Nonetheless, I appreciate you taking the time to query Claude-3.5 and sharing the results.
According to Amazon Bedrock adds Claude 3.5-sonnect[0] is a multimodal model with state-of-the-art vision processing, which tells me that Claude can "cheat" by rendering the SVG and then passing to the image model for the results which is perhaps one way of describing "imagining" things, but it does not describe a strict Large Language Model which predicts the next token.
> The multimodal Claude 3.5 Sonnet also excels at processing images with state-of-the-art vision, particularly when interpreting charts and graphs—helping to get faster, deeper insights from data. It can accurately decipher text from imperfect images—for example, poorly scanned documents—and in doing so, glean more insights than from text alone.
[0] https://www.aboutamazon.com/news/aws/amazon-bedrock-anthropi...
One reason might be that you get the benefit of feeling superior to LLMs without any cost or risk of finding yourself mistaken.
""" The SVG code you provided creates an image with three circles:
A large circle at the center (at coordinates 100, 100) with a radius of 50. Two smaller circles at the top left (at coordinates 50, 50) and top right (at coordinates 150, 50) with a radius of 40 each. This configuration closely resembles the head and ears of Mickey Mouse. The larger circle represents the head, and the two smaller circles represent the ears. """
Slightly more generic prompt: "What is this code drawing? Can you suggest what it might be representing?"
""" What It Might Represent: Mickey Mouse Head: The arrangement resembles a simplified "Mickey Mouse" head, where the large central circle represents the head, and the two smaller circles represent the ears.
Abstract Art: Alternatively, it could be an abstract design with three overlapping circles of different sizes, suggesting some form of simple geometric art.
The circles are positioned such that the smaller ones are symmetrically placed above the larger central one, which is a typical representation of a head with ears. """
And more generic "I've been sent this code, what is it drawing?" ends with:
""" You will see three overlapping circles. The largest circle is near the center, while two smaller circles are positioned above and slightly to the left and right of the center. The arrangement may resemble a basic representation of a Mickey Mouse head shape, with the large circle as the face and the two smaller circles as ears. """
What's imagining, then? The way LLMs explore different predictive branches in order to find an optimal solution doesn't seem all that different than what I consider imagining: Thinking about what could be and considering different variations on that idea.
An LLM isn't a brain, so there is no implication of it being said in the truest human sense, but it seems like a decent analogy to me.
So, the task is "describe what an SVG depicts without saving it to a raster image and prompting with that"?
/? LLM stack app architecture https://www.google.com/search?q=LLM+stack+app+architecture&t...
https://cobusgreyling.medium.com/emerging-large-language-mod... ; Flexibility / Complexity,
> Given the arrangement of three overlapping circles, it resembles the classic depiction of a *Mickey Mouse* head silhouette:
> The two smaller circles represent Mickey's ears.
> The larger circle represents his head.
> This is a stylized version of the iconic Mickey Mouse logo.
Imo: In order to predict the next token for non-trivial tokens, (of which there are many on the training data), you do have to do some more complex thinking/reasoning than just a lookup of past training data.
A. Valid Mickey is detected by the model. "...This arrangement might resemble a basic version of a Mickey Mouse shape, where the two smaller circles represent the ears and the larger circle represents the head...", https://chatgpt.com/share/3999859a-b6db-4671-8b69-0ec6a5bac3...
B. Invalid Mickey is not misclassified as Mickey by the model and is correctly described. "...these circles will overlap, creating a pattern where the largest circle (Circle 3) dominates the right side of the canvas, with the other two smaller circles overlapping it and each other in the middle...", https://chatgpt.com/share/df3c57ac-495b-4e4c-b00c-bae31781c4...
> The circles overlap in certain areas, depending on their size and position, creating a layered visual effect where the largest circle (third one) dominates most of the canvas space.
Though tbf to it, I'm not sure I'd say this looks like MM either: https://i.imgur.com/0VHdocf.png (unless I knew this was the intent prior).
TLDR: it seems like an LLM might be able to tell you your SVG is a "Mickey Mouse Head Silhouette"
I think the reason why we don't view HTML as a programming language is because it is explicitly designed to be a markup language that declares content rather than a series of instructions that is interpreted as a program. A program needn't demonstrate turing completeness to be a "computer program", it just needs to be a sequence of instructions that a computer executes. To me, that suggests that there's a degree of abstractness and subjectivity involved. For example, any SVG document could also be rewritten 1:1 with no loss in fidelity as a series of commands that has the same effect, as can pretty much any declarative markup language; what is actually happening during parsing is hard to distinguish from an interpreter. Humans can "know it when they see it", but I doubt there's an exact criteria that can go along with the human "feel" of what makes a program, a program.
Aren't HTML and vector graphics descriptions both data structures that could be interpreted via a Turing-complete interpreter? I don't see the difference between HTML and a C AST in this regard.
A computer program is a data structure, but data structures are not necessarily computer programs.
Even if not there is always an OLE embedding.