Is the locally run model censored in the same way as the model hosted on HuggingFace? If so, I wonder how that censorship is baked into the weights—particularly the errors thrown when it starts to output specific strings.
uv run --with 'numpy<2.0' --with mlx-vlm python \
-m mlx_vlm.generate \
--model mlx-community/QVQ-72B-Preview-4bit \
--max-tokens 10000 \
--temp 0.0 \
--prompt "describe this" \
--image Tank_Man_\(Tiananmen_Square_protester\).jpg
Here's the output: https://gist.github.com/simonw/e04e4fdade0c380ec5dd1e90fb5f3...It included this bit:
> I remember seeing this image before; it's famous for capturing a moment of civil resistance. The person standing alone against the tanks symbolizes courage and defiance in the face of overwhelming power. It's a powerful visual statement about the human spirit and the desire for freedom or protest.
I wonder if the response was not truncated this time only because the word “Tiananmen” did not happen to appear in the response. At least some of the censorship of Chinese models seems to be triggered by specific strings in the output.