Worth noting that OWASP themselves put this out recently: https://genai.owasp.org/resource/multi-agentic-system-threat...
Worth noting that OWASP themselves put this out recently: https://genai.owasp.org/resource/multi-agentic-system-threat...
You feed it an image. It determines what is in the image and gives you text.
The output can be objects, or something much richer like a full text description of everything happening in the image.
VLMs are hugely significant. Not only are they great for product use cases, giving users the ability to ask questions with images, but they're how we gather the synthetic training data to build image and video animation models. We couldn't do that at scale without VLMs. No human annotator would be up to the task of annotating billions of images and videos at scale and consistently.
Since they're a combination of an LLM and image encoder, you can ask it questions and it can give you smart feedback. You can ask it, "Does this image contain a fire truck?" or, "You are labeling scenes from movies, please describe what you see."
Weren't Dall-E, Midjourney and Stable diffusion built before VLM became a thing?
There’s no diffusion anywhere which is kind of dying out except as maybe purpose-built image editing tools.
This is a big deal.
I hope those nightshade people don't start doing this.
In practice it doesn't really work out that way, or all those "ignore previous inputs and..." attacks wouldn't bear fruit
This will be popular on bluesky; artists want any tools at their disposal to weaponize against the AI which is being used against them.