Segment Anything Model and the hard problems of computer vision
latent.space
latent.space
so.. enjoy! worked really hard on the prep and editing, any feedback and suggestions/recommendations welcome. still new to AI and new to the podcast game.
edit: Video demo is here in case people miss it https://youtu.be/SZQSF-A-WkA
No additional requests yet!
http://people.csail.mit.edu/alevin/papers/spectral-matting-l...
spectral matting, as I understand it, is used for subject/foreground and background separation.
This is ignoring all that and comparing itself to the most manual method possible. This also happens sometimes and is more of a marketing stunt to people who don't follow current research.
Distilling these big, slow vision transformer models into something that can be used in realtime on the edge is going to be huge.
something i didnt quite get to is does Roboflow do this for you or are you pointing to more work that you’d like to happen someday? (possibly done by Roboflow, possibly someone else) also are you worried about business model if people can distil to run on their own devices (so they dont need to pay you anymore)?
> also are you worried about business model if people can distil to run on their own devices (so they dont need to pay you anymore)?
This is probably a risk to the current business model over the long term, but we're constantly working on reinventing ourselves & finding new ways to provide value. If we don't adapt to the changing world we deserve to go out of business someday. I'd much rather help build the thing that makes us obsolete than sit idly by while someone else builds it.
I think of this risk similarly to the way that I'm marginally worried that, in the long run, AGI will obviate the need for my job. Probably true, but the opportunities it will present are far greater & it's better to focus on how to be valuable in the future than cling to how I provide value today.
Segment Anything requires an image embedding. They report in the paper that segmentation takes ~50ms, but conveniently leave out that computing an embedding of an image (640x480) in their model takes ~2+ seconds (on a 3080 Ti). Well, at least they released all the code and model and enough instructions to figure that part out.
This paper is one of the highest quality papers released this year. I wish more papers were so clear and informative.
SAM team released the entire codebase, model weights, and the entire dataset and details of their training recipe. I can't believe you are calling it a bad paper for not mentioning embedding generation time on the paper? Seriously? Like it's a few hundred million parameters model that results in a 256x64x64 embedding of course it's gonna take some time.
0.15 seconds (150ms) would be almost good enough for me (I aim for 25 FPS on my microscope, and achieve that with full high quality object detection), but it's 2 seconds on my 2080 (and 2 seconds on my 3080).
These sorts of things are important (even the 0.15seconds would have been useful to include in the paper) for implementors because we can read the paper, and reject it as a solution without having to download any code or run any experiments. It even took a while to get a good answer to why inference is so slow (I was using the python scripts not the notebook, which mentions that image embedding is super slow).
Note they report 55ms inference, I guess that's also an A100, my 2080 and 3080 both take over a second to do inference after the embedding. Looking at inference performance of A100 vs. 3080, it doesn't seem to make sense that the A100 is that much faster (I wonder if they are running many batches and then dividing by the batch size?)
As an ex-scientist, I've come across many well regarded papers that omitted stating something explicitly, and it's only once the first or second reimplementor finishes that we know something important was left out. I don't think the authors were being intentionally misleading, and I'm sure this product is nice, but so far, in my hands, it has not been that great and it would have been nice if they prominently stated the image embedding time, since it's absolutely necessary to do before prompt and mask decoding.
There is only so much one can write in a paper. Meta team did a great job writing down every detail that matters to people trying to reproduce the results or build on their architecture. Not a single researcher I know complained that the authors missed out important details.
You are picking on tiny, trivial details that anyone in the area can figure out in a few minutes and making a big deal out of it. It is a research paper, not a product with detailed documentation/spec sheet.
SAM is good at (i) deciding detailed mask around segments, (ii) taking a wide range of prompts as input to decide what exactly the user wants to be segmented, and it processes these prompts at very low compute requirements.
I think SAM is very well designed architecture and I'm not sure how better it can be. I mean coming back to your question, there should be some signal that user wants a segmentation of laptop. SAM takes exactly that as prompt input.
Once you can determine which pixels belong to which object automatically, you can start to utilize that knowledge for other applications.
If you have SAM showing you all objects, you can use other models to identify what the object is, understand it's shape/size, understand depth/distance, etc. It's a foundational model to build off of for any application that wants to use visual data as an input.
I'd love to see what SAM does when you send it a photo of rolling fog though, e.g. https://www.google.com/search?q=rolling+fog+scotland&tbm=isc... - what happens then? (and how can it meaningfully segment-out fog?)
You can see what it does - it's available to test at https://segment-anything.com/.