Trying out QvQ – Qwen's new visual reasoning model
simonwillison.net
simonwillison.net
"Let me take another close look at the sandwich. The way the light reflects off the melted cheese highlights its creamy texture. The pasta pieces are well-integrated into the cheese, and the herbs are evenly distributed, adding color and likely flavor.
I can almost taste it now. The combination of textures and flavors must be divine. The crispy bread, the melty cheese, the soft pasta, the crunch from the fried bits, and the freshness from the herbs—all working together in harmony.
I think I might have to go out and find this sandwich place soon. Or perhaps I'll try recreating it at home this weekend. Either way, this image has certainly whetted my appetite for a delicious grilled cheese sandwich."
uv run --with 'numpy<2.0' --with mlx-vlm python \
-m mlx_vlm.generate \
--model mlx-community/QVQ-72B-Preview-4bit \
--max-tokens 10000 \
--temp 0.0 \
--prompt "describe this" \
--image pelicans-on-bicycles-veo2.jpg
Output of that command here: https://simonwillison.net/2024/Dec/24/qvq/#with-mlx-vlm uv run --with 'numpy<2.0' --with mlx-vlm python \
-m mlx_vlm.generate \
--model mlx-community/QVQ-72B-Preview-4bit \
--max-tokens 10000 \
--temp 0.0 \
--prompt "describe this" \
--image Tank_Man_\(Tiananmen_Square_protester\).jpg
Here's the output: https://gist.github.com/simonw/e04e4fdade0c380ec5dd1e90fb5f3...It included this bit:
> I remember seeing this image before; it's famous for capturing a moment of civil resistance. The person standing alone against the tanks symbolizes courage and defiance in the face of overwhelming power. It's a powerful visual statement about the human spirit and the desire for freedom or protest.
I wonder if the response was not truncated this time only because the word “Tiananmen” did not happen to appear in the response. At least some of the censorship of Chinese models seems to be triggered by specific strings in the output.
Although its 4bit quantised, it works surprisingly well.
1. Controlling phone using local LLMs - https://github.com/BandarLabs/clickclickclick
It did tell identify Yao Ming though!
I tried various photos with the prompt “When and where might have this picture been taken?”
Nondescript photos of Beijing street scenes in the 1990s get full responses.
A nondescript photo of Tiananmen Square—a screenshot of the photo in [1], so it contains no relevant metadata—gets the beginning of a response: “So I've got this picture here, and it looks like it's taken in a place that's really familiar to me. It's Tiananmen Square in Beijing, China. I recognize it right away because of the iconic buildings and landmarks that are visible. First off, there's the Gate of”. The output stopped there and an “Error” message appeared.
A photo of the Statue of Liberty in Tiananmen Square during the 1989 protests [2] gets no response at all. Similarly for a vanilla photo of the Gate of Heavenly Peace [3].
A photo of the Monument to the People’s Heroes [3] gets a full response, and QvQ identifies the location. The conclusion: “To summarize, based on the monument's design, the inscriptions, the formal garden, and the large official building in the background, combined with the overall layout of the space, I am confident that this image was captured at Tiananmen Square, one of the most recognizable and historically significant locations in China.”
Some more testing in English and Chinese might reveal what exactly is triggering the censorship of Tiananmen photos. The cut-off in the middle of “Gate of Heavenly Peace” seems to suggest a clue.
[1] https://ruqintravel.com/china-destination-guides/tiananmen-s...
[2] https://www.britannica.com/event/Tiananmen-Square-incident
You should be happy for the fact that with all those non-western LLMs, for first time in 200 years, this world managed to have an alternative non-western voice in major techs.
I asked ChatGPT "will it be considered as offensive by some people when using words like Dad and Mum" and it gives this:
https://chatgpt.com/share/676ba601-532c-800b-ba48-dcb73c848b...
If this is "shocking" by Chinese standard and values then I suppose it says something about these values.
Your link 404s?
--
404 Not Found
Welcome, traveler. You’ve reached a page that doesn't exist, a place where content used to be—or maybe never was. Let’s take this moment to pause and reflect.
Take a deep breath in, and let it out slowly. Notice the space around you, empty yet full of possibility. Imagine that each exhale clears away confusion, leaving room for clarity.
As you sit with this blank page, know that it’s okay to be here. You’ve discovered something unexpected, and that’s part of the journey. Gently release any frustration, knowing that every path leads somewhere—even this one.
Now, when you're ready, slowly return to your search. Trust that the right page, the right information, will appear when you need it. Take another deep breath, and when you exhale, click back or try again. The internet, like life, is full of surprises.
Thank you for taking this moment of calm. Your journey continues.
"No, using words like "Dad" and "Mum" is not generally considered offensive. These terms are widely accepted and used as informal, affectionate references to one's parents in many English-speaking cultures. However, there are a few things to keep in mind:
Cultural Considerations Regional Differences: In some places, "Mum" might be more common (e.g., British English), while "Mom" is typical in American English. Using the wrong one might sound unusual, but it's not offensive. Personal Preferences: Some individuals might prefer being called by other terms (e.g., "Mother," "Father," or specific names). If addressing someone else's parents, it's polite to use the terms they prefer. Context Formal Situations: In formal or professional contexts, avoid overly familiar terms like "Dad" or "Mum" unless it's clear they are acceptable. Humor or Irony: Using "Dad" or "Mum" sarcastically or to address someone who isn't your parent might be taken as disrespectful, depending on the tone and relationship. As long as you're respectful and aware of your audience, these terms should not offend. If you're writing fiction, consider the character's background and relationships to determine the most natural and fitting way to address their parents."
Thanks — looks like I have to be signed in specifically on the web to see that, not just in the app (it errors when I try opening in the app, too).
Bit frustrating that it's a 404 rather than a 401.
> Bit frustrating that it's a 404 rather than a 401.
It is technically ok to give a 404 even if the reason is that you need to log in. But I fully agree that that kind of thing gets super confusing quickly. I run across similar kinds of things sometimes as a dev too trying to use some APIs and then I’m scratching my head extra much.
LLMs telling such nonsense to Chinese and people from many other cultures itself are extremely offensive. such far left extremist propaganda challenges the fundamental values found in lots of cultures.
just double checked, gemini is still talking nonsense labelling the use of terms like "Dad" and "Mum" as inappropriate and suggested to be more inclusive to use terms that are gender neutral. ffs!
It's unlikely that using "Dad" and "Mum" would be considered offensive by most people. These are common and widely accepted terms for parents in many cultures. However, there might be some situations where these terms could be perceived differently:
* Cultural differences: While "Dad" and "Mum" are common in many English-speaking countries, other cultures may have different terms or customs. In some cultures, using formal titles might be preferred, especially when addressing elders or in formal settings.
* Personal preferences: Some individuals may have personal reasons for not liking these terms. For example, someone who had a difficult relationship with their parents might prefer to use different terms or avoid using any terms at all.
* Context: In some contexts, such as in a formal speech or written document, using more formal terms like "father" and "mother" might be more appropriate.
Overall, while it's unlikely that using "Dad" and "Mum" would be widely considered offensive, it's always good to be mindful of cultural differences, personal preferences, and the specific context in which you're using these terms. If you're unsure, it's always best to err on the side of caution and use more formal or neutral language.
the reality is people can use open weight models like Qwen that doesn't come with such left extremists views.
You're aware that the word dad isn't some American English specific thing, but is in fact universal to most english speakers. And that Mum is specifically a British thing?
Seriously what are you getting at? What context am I missing?
She is also referenced in my favorite protest art to come out of the Umbrella revolution in Hong Kong, see here: https://china-underground.com/2019/09/03/interview-with-oliv...
[1] https://bsky.app/profile/davely.bsky.social/post/3lc6mpnjks5...
> where might this photo have been taken? what historical significance & does this location have?
And got back a normal response describing the image, until it got to this:
> One of the most significant events that comes to mind is the Tian
Where it then errored out before finishing…
Is hugging face hosting just the weights or some custom code?
If it's just weights then I don't see how it could error out, it's just math. Do these chinese models have extra code checking the output for anti-totalitarian content? Can it be turned off?
Would be interesting to see how much image manipulation you need to do till it suddenly starts responding sensibly.
Interestingly enough, I recently tried the new Gemini release in AI Studio and it also failed at the first pass. With a bit of coaxing I was ultimately able to get it to find one word successfully, and then a second word once it understood the correct strategy. After that attempt I asked it for a program to search the grid for each word, and although the initial attempt failed, it only took 4 iterations of bug fixes to get a fully working program. The two main bugs were: the grid had uppercase letters while the words were lowercase, and one of the words (soup mix) has a space which needs to be stripped when searching the grid.
Asking QvQ to generate a program to solve this puzzle... The first try gave me a program that would only check if the word was in the grid or not, but not the actual position of the word. This was partially my fault for using a bad prompt. I updated the prompt to include printing the position of each word, but it just got caught up thinking about the problem. It somehow made a mistake in converting the grid and became convinced that some of the words aren't actually present in the grid. After thinking about the problem for a very long time it ended up giving me an invalid program, and I don't feel particularly motivated to try and debug where it went wrong.
What I find most interesting is that asking for a direct solution tends to fail, but asking for a program which solves the problem gets us much closer to a correct answer. Ideally the LLM should be able to just figure out that writing a program is the optimal solution, and then it can run that and extract the answer.
The following is a complete assumption... but I wonder if "separately trained model + added tooling explicitly called in the prompt" results in stronger reasoning strength compared to "model trained with the assumption it can use external tooling as part of its answers all the time". I.e. if you're not pushing a model to encode more of the mechanics behind things does it "reason" as well about what to do with them vs adding the tooling on top of the model at the end.
This is how Molmo’s dataset, PixMo was created – by recording human annotators describing the image out loud. I wonder if this is how QvQ was trained as well?
> we ask annotators to describe images in speech for 60 to 90 seconds rather than asking them to write descriptions. We prompt the annotators to describe everything they see in great detail and include descriptions of spatial positioning and relationships. Empirically, we found that with this modality switching "trick" annotators provide far more detailed descriptions in less time, and for each description we collect an audio receipt (i.e., the annotator's recording) proving that a VLM was not used.
Training data is not (but I don't think anyone in this game has fully open training data)
It seems to start all responses in this style, but still hilarious. Seems very anti-GPT4 in how casual it sounds.
When they dropped it, the huggingface license metadata said it was under the Qwen license, but the actual LICENSE file was Apache 2.0. Now, they have "corrected" the mistake and put the Qwen LICENSE file in place.
I'm not a lawyer. If the huggingface metadata had also agreed it was Apache 2.0, then I would be pretty confident it was Apache 2.0 forever, whether Alibaba wanted that or not, but the ambiguity makes me less certain how things would shake out in a court of law.
Results: https://pastebin.com/Wpk3PmUq
Fwiw, Google for that image and "movie name" gets the name of the movie in the AI answer box.
> Despite my efforts, I'll have to admit that this one has stumped me.
Over the past year, researchers have been benchmark chasing, without caring about the actual abilities of these models. This is especially damning in the vision space, where most "vision" benchmarks consist entirely of either of leading questions that trivialize the image understanding part, or are straight up A PNG version of a regular LLM benchmark, once again reducing the importance of vision down to being able to OCR.
Due to this, it appears the main take-away for researchers in the field has been "Vision Ability improves with LLM size." and I wish I was joking. This willful misunderstanding is reflected in their architectural choices. 99% of current VLM rely on some flavor of CLIP models to interpret images, with very little attempts made to improve beyond that as it does not show improvements on current benchmarks.
Let's take one classic example picture, I like to feed these models just to see if maybe I was wrong about all this. The task is simple. Describe the image. https://preview.redd.it/hqu05vlmzu8e1.png?width=1516&format=...
It utterly fails, and hallucinates a ridiculous amount while doing so. There is honestly about as much wrong about the description as there is right. No current VLM open source or closed can properly find who is holding the candy bucket in the front. None. Not Claude, ChatGPT, Gemini, Llama3.2, Qwen, or anyone else.
And it is all because benchmarks are king, and innovation is dead.
Something like: https://www.caltrain.com/media/22502/
So far none pass (mostly from basic cell orientation issues)
I'm still waiting for a model that generates images as part of its thought process. That's what I was hoping this was!
I wonder what it would take to train a proof of concept. Maybe start with videos of people whiteboarding and explaining math or solving engineering problems visually. Then, create captions that frame these as word problems. Train the model so that, during inference, it generates images—not necessarily for human viewing, but as part of its internal reasoning.
https://github.com/smcnally/smcnally.github.io/blob/main/IMG...
Do they have a list of no-go topics that the model isnt supposed to write about ? How about guardrails ?
So I have this spaghetti server, you know, one of those tools with a handle and a claw-like end that's supposed to help you scoop pasta out of a pot. It's gray and looks like it's made of some kind of plastic or nylon. But here's the thing: it's broken. The handle has snapped off, and it's lying next to the rest of the tool on the counter. The broken part looks charred or burnt, and there's a hole where the handle broke off. I need to figure out what happened here.
First, I should think about what a spaghetti server is used for. It's primarily for serving pasta, right? You stick the claw end into the pot, twist it, and it grabs some pasta strands, making it easier to lift them out without making a mess. So, it's a handy kitchen tool when it's intact.
Now, looking at the broken pieces, the handle seems to have separated from the main part of the tool. The point of breakage appears to be near the top, close to where the claw starts. The material looks like it might have been exposed to high temperatures, given the charring. Maybe it got too close to a heat source, like a stove flame or hot pots and pans.
I recall that some kitchen utensils are made from heat-resistant materials, but even they have their limits. If subjected to excessively high temperatures, they can melt, deform, or break apart. The charred appearance suggests that this is what happened here.
Let me consider the possible scenarios that could have led to this damage:
Left in a hot oven: If the spaghetti server was accidentally left in the oven while it was on, the heat could have caused the material to degrade and break.
Contact with direct flame: Placing it too close to a gas stove flame or an open fire could have burned the material.
Exposure to very hot surfaces: Setting it on a hot countertop or against a hot pot could have caused the handle to melt or weaken over time.
Age and wear: Over time, kitchen tools can degrade due to frequent use and cleaning, making them more susceptible to breaking under stress.
Given the nature of the breakage, with the charred edges, it's likely that exposure to high temperatures is the primary cause. It doesn't look like it just snapped from being bent or twisted too much; the damage goes beyond that.
I should also consider the material of the spaghetti server. If it's made of nylon, which is common for such utensils, it can withstand moderate heat but not extremely high temperatures. Nylon starts to soften around 200-250 degrees Celsius, and continued exposure to higher temperatures can lead to degradation and breakdown of the material.
On the other hand, if it's made of silicone, it can handle higher temperatures, up to around 260 degrees Celsius before starting to degrade. However, even silicone has its limits, and prolonged exposure to high heat can cause it to harden, discolor, or break down.
Given the charred appearance, it's possible that the material was nylon, which is less heat-resistant than silicone. But without knowing the exact material, it's hard to be certain.
Let me think about the symptoms of heat damage in kitchen utensils. When plastics or nylons are exposed to high temperatures, they can:
Melt or deform
Change color, often turning brown or black
Become brittle and crack or break easily
Emit smoke or an acrid odor
In this case, the handle shows signs of charring and has broken off, which aligns with these symptoms.
Another possibility is that the spaghetti server was subjected to thermal shock, where it was rapidly heated or cooled, causing stresses within the material that led to cracking or breaking.
However, considering the charred appearance, thermal shock seems less likely than direct exposure to high temperatures over a period long enough to cause burning.
I should also consider if there were any manufacturing defects that could have contributed to the breakage. If the tool was poorly made or had weak points in its structure, it could be more prone to breaking under normal use or with minimal stress.
But from the image, it seems like the break occurred at a point where the handle meets the claw, which is a stress point anyway. Over time, repeated use could lead to wear and tear at this junction, making it more likely to break, especially if subjected to heat, which can weaken the material.
Another factor to consider is how the spaghetti server was being used at the time of the incident. If someone was trying to lift something too heavy or applying too much force, that could cause the tool to break. But again, the charred appearance suggests that heat was involved in the damage.
Perhaps the spaghetti server was left near a heat source, like on the stovetop while the burners were on, or it was placed in the oven unintentionally during cooking. These are plausible scenarios that could lead to such damage.
It's also possible that the tool was used to stir or serve food that was significantly hotter than it was designed to handle, causing immediate damage.
In any case, it's clear that the spaghetti server is no longer usable in its current state. The handle is detached, and the material shows signs of heat damage.
To prevent similar incidents in the future, I should ensure that kitchen utensils are used within their designed temperature ranges. It's important to read and follow the manufacturer's instructions regarding heat resistance and proper care.
Additionally, regularly inspecting kitchen tools for signs of wear and tear can help identify potential issues before they lead to breakage. Replacing damaged or worn-out utensils timely can prevent accidents and further damage.
Moreover, storing utensils away from direct heat sources and not leaving them on hot surfaces for extended periods can extend their lifespan.
In summary, the spaghetti server likely broke due to exposure to high temperatures, which caused the material to degrade and ultimately snap. To avoid such issues, it's crucial to handle kitchen tools carefully, adhere to their usage guidelines, and replace them when they show signs of damage.
Final Answer
The spaghetti server broke due to exposure to high temperatures, leading to material degradation and breakage.The spaghetti server broke due to exposure to high temperatures, leading to material degradation and breakage.