Segment Anything Model (SAM) can "cut out" any object in an image
segment-anything.com
segment-anything.com
Really impressive stuff! Congrats to the team that achieved it
[0] https://segment-anything.com/demo
[1] https://segment-anything.com/model/interactive_module_quanti...
[2] https://segment-anything.com/model/interactive_module_quanti...
Also read the "Efficient & flexible model design" section on the page.
Also in the FAQ: "How big is the model?
The image encoder has 632M parameters. The prompt encoder and mask decoder have 4M parameters."
> What platforms does the model use? > The image encoder is implemented in PyTorch and requires a GPU for efficient inference. > The prompt encoder and mask decoder can run directly with PyTroch or converted to ONNX and run efficiently on CPU or GPU across a variety of platforms that support ONNX runtime.
WebGPU just shipped today in Chrome, there are no people reporting that the demo doesn't work with their days-old browser, so it doesn't use WebGPU.
While it's possible, without WebGPU it's really tedious to run NN in the browser.
Also, the model is implemented in PyTorch and wasn't converted to other model format for other runtime. While technically you can compile CPython and PyTorch to WASM and run the duo in browser, there are definitely no GPU access.
Given that they explicitly mentioned the decoder was converted to ONNX, it's obvious this isn't done for the encoder and they really mean PyTorch, running with Python, on a server.
Okay, so you browser can't run the encoder, yet the web demo works, it's quite obvious on which server the encoder run.
I downloaded the code from their repo, exported their pytorch model to onnx, and ran a prediction against it. Everything ran locally on my system (cpu, no cuda cores) and a prediction for the item to be annotated was made.
The small ONNX model just decodes the output of the larger model into the masks etc. But the bulk of the "computation" is done by a much larger vision transformer somewhere else. It really needs a GPU with a fair amount of memory to run anywhere close to real-time.
We have a similar "Smart Polygon" tool[2] built into Roboflow but this is next level. Having the model running in the browser makes it so much more fun to use. Stoked it's open source; we're going to work on adding it to our annotation tool ASAP.
[1] Some examples from Open Flamingo last week https://news.ycombinator.com/item?id=35348500
[2] https://blog.roboflow.com/automated-polygon-labeling-compute...
Paper: https://scontent-sea1-1.xx.fbcdn.net/v/t39.2365-6/10000000_6...
Announcement: https://ai.facebook.com/blog/segment-anything-foundation-mod...
Code & Model Weights: https://github.com/facebookresearch/segment-anything
-Emily
-Emily
We do on some platforms like Discord.
> What if you get logged out and the other identity is the one that knows the password?
Before we moved to a password manager we used the same password everywhere. Now, obviously, we have a password manager.
-Emily
No, because systems don't necessarily work that way. For us, the boundaries between our members aren't totally uncrossable. Information has gotten through in the past when it would be especially important or needed.
Though I guess it's funny that this topic comes up under a post called "segment anything". I guess our brain did that~
-Emily
Abstract: https://ai.facebook.com/research/publications/segment-anythi...
> "Prompt encoder. We consider two sets of prompts: sparse (points, boxes, text) and dense (masks). We represent points and boxes by positional encodings [95] summed with learned embeddings for each prompt type and free-form text with an off-the-shelf text encoder from CLIP [82]. Dense prompts (i.e., masks) are embedded using convolutions and summed element-wise with the image embedding."
Edit: Found the answer myself: https://github.com/facebookresearch/segment-anything/issues/...
I think a more BSD would be better, or LGPL. Either would be more business friendly.
With some caveats, software licenses from most to least business friendly roughly go:
Apache > BSD > MIT > MPL > LGPL > GPL > AGPL
You can use LGPL in commercial, closed-source projects as long as you keep the LGPL code in a separate dynamically linked library, e.g. a DLL, and provide a way for users to swap it out for their own patched DLL if they wish. (Plus some other license terms.)
Also, you can always use LGPL code under the terms of the GPL, so there's no way LGPL is more restrictive than GPL.
* Do what you want
* Don't sue us
* You license any patents you control and used in this work. If you sue someone for patent violation for using this then other entities can counter sue you for violating any of their patents used in this work.
There is no viral nature, and it is older than GPLv3.
It's most simialr to BSD of the licenses you list.
On the other hand, this is one of those things (like VR) that is a distinctly non-Facebook project. It makes no sense to position or market this as "Facebook" research. The Homepod isn't called the iPod Home for obvious reasons, so it stands to reason that Facebook execs realized selling someone a "Facebook Quest" sounds like a metaphor for ayahuasca. It's not entirely stupid to rebrand, especially considering how diverse (and undeniably advanced) they've become in fields like AI and VR.
But yeah if you do open source adding an element of corporate branding is a sure way to kill the project. That's why it's not called "Apple Swift" or "Microsoft TypeScript".
To me, this feels like they are not avoiding the "Meta" brand at all.
Relative to the rest of FAANG (or even Fortune 500), Facebook might have the least blood on their hands when everything is said and done.
Whereas, if one engineer spins up some random static open source documentation website on AWS, it really can't go wrong in a way that causes trouble for the rest of the company.
My IT experiences elsewhere have left me a little jaded. :)
Beginning in August 2017, the Myanmar security forces undertook a brutal campaign of ethnic cleansing against Rohingya Muslims. This report is based on an in-depth investigation into Meta (formerly Facebook)’s role in the serious human rights violations perpetrated against the Rohingya. Meta’s algorithms proactively amplified and promoted content which incited violence, hatred, and discrimination against the Rohingya – pouring fuel on the fire of long-standing discrimination and substantially increasing the risk of an outbreak of mass violence. The report concludes that Meta substantially contributed to adverse human rights impacts suffered by the Rohingya and has a responsibility to provide survivors with an effective remedy.
https://www.amnesty.org/en/documents/ASA16/5933/2022/en/
See also
https://en.wikipedia.org/wiki/Rohingya_genocide
https://en.wikipedia.org/wiki/Rohingya_genocide#Facebook_con...
Of course being on github and permissively license is huge.
Right now my home server w/ an RTX 2080 is able to do mask prediction in about ~4 seconds (I'm running the sample script in "directory" mode) per image (640x480).
I'd love to be able to get the first mask back in 0.1ms, so I can do 10FPS on the scope. Is there a practical way to speed things up (my guess would be buying the absolute fastest GPU I can afford)? I can run the object detection to get a reasonable seed location, is that what the paper means by prompting?
For the second part of the question, a 2080 should get you close to 10FPS operation. For a ballpark estimate, using an off-the-shelf repo like Ultralytics's YOLOv5 lets you run object detection (not masking) at something like 100FPS. Masking should not add that much overhead.
w.r.t. GPUs yes, these days more money equals more speed for GPU NN inference, though there are diminishing returns. A 3090 might get you the best bang for your buck these days while still having enough VRAM to run fancier models which may need more than the 12 GiB many other GPUs have.
Finally, I haven't read the paper too carefully but I believe that by prompting they mean that you have the option of describing in human language what you want the model to select, rather than the model being "hardwired" to do this. In other words, you could prompt the model to "segment the red car only" and it would do it, rather than just having the model blindly segment every object in the image, and then relying on custom scripting to potentially post-process these segments.
If you want to try it on one reach out to me (email in profile). We rent those out in the cloud. Would allow you to confirm performance before buying one for local use.
It's definitely using the GPU- I'm running nvidia-smi and I see near 100% utilization on the GPU while the CPU is using 1 core. If I run the script with --device=cpu then I see my server using 4 cpu cores and no GPU and it takes tens o seconds per image.
I'm trying to check with people who have experience with this specific model.
I have bounding box already, so I could prompt the model with that, but all of this runs counter to the published performance numbers.
however, CLIP/BLIP and boxes should be much faster than 4 seconds, even on a 2080. I had a python CLIP CSV tag script running in a directory with thousands of images and it was taking <=2 seconds per image on a Geforce GTX 1070 TI, with 8GB of memory - an old card without any tensors. CLIP is much slower than some other mechanisms, for instance a deepbooru classifier is about 4x faster than CLIP/BLIP on my RTX 3060 12GB. CLIP of a random image around your dimensions takes ~4-5 seconds, and deepbooru takes about 1.5 seconds. edit: the additional time is the overhead of the webUI, i am guessing
What will probably have to happen is some sort of auto-crop that only forces the model to view a very tiny section of the image. You mentioned you already had a model, was it trained from scratch, or using an existing model?
The main issue I have with DIS is that creating the labels of my own dataset is super expensive (I think it might be easier to generate the training data using stable diffusion rather than human labelling)
BTW, I used DIS to create the labels of a batch of 20 images, I manually corrected the labels and used them to fine tune a new model. That worked well but still it took me several hours to edit labels.
I tried using stable diffusion generated labels several weeks ago but I think with controlnet and other advances I should try again.
(My dataset is about 100k images. I probably only need to label about 10k to fine tune DIS).
You can long press on an image and it cuts out whatever thing it thinks you're pressing on
They also use it in interesting ways, like making stuff in the photo slightly overlap the clock on the lockscreen
Does anyone know if that works the same way as this?
What's preventing us from taking something like convnext or a hybrid conv/attention model and hooking that up to a decoder stack? I fee like the results would be similar if not better.
EDIT: Clarifying that encoder/decoder refers to the transformer stack, not an autoencoder.
You mean like in an U-Net architecture?
> Here we highlight some results of ViT-22B. Note that in the paper we also explore several other problem domains, like video classification, depth estimation, and semantic segmentation.
https://ai.googleblog.com/2023/03/scaling-vision-transformer...
Still uses patches though, but they're mixing data between patches.
Finally, an actually open model
The image encoder has 632M parameters.
The prompt encoder and mask decoder have 4M parameters.https://www.youtube.com/watch?v=hHwjceFcF2Q
All it takes is a couple of tools glued together and you're getting there.
I don't read issues for major repositories, so perhaps this is standard? There are a ton of one line "issues", no clear example, test case, attempt to debug, not even a pull request for the one that points out a typo in the README.
Gross. It seems like none of these issues will never be read because they are going to drown in garbage.
[1] https://github.com/facebookresearch/segment-anything/issues
Microsoft knows this. Get on the ball, guys.
Meta New Segmentation Model - https://news.ycombinator.com/item?id=35453625 - April 2023 (7 comments)
(It does have difficulty finding the smallest possible area, but it's a significant advance over most existing options since in my brief test, it can usually spot the entire silhouette of figures, which is where painting a boundary is most tedious).
<human></human> <ad></ad>
If Meta frequently shares their "algorithms" they take the blame out of its usage. After all, who is to blame when everybody does "it" and you are very open about it.
Use cases, talent visibility as well as attraction also plays a role. After all, Google was so fancied, due to its many open source projects. "Show, don't tell".
Probably too cynical, but you can potentially view it as a weak form of collusion under the guise of open research.
For competitive LLMs targeting text generation, especially for training, a compute-based barrier is more significant.
I think that argument falters when the weights are released, which lowers the barrier by a lot as training of large models is much more expensive than inferences. A weak form of collusion would be publishing papers that explain enough for the practitioners to fill in the gaps (so casuals are left out) and not publishing the weights so only other large companies can afford to implement and train their versions of models.
My own view is that open-publishing in AI is mostly bottom-up, and the executives tolerate open publishing for the reasons you gave.
Incidentally most companies won't publish their crown jewels i.e. Camera apps on Google and Apple phones had great segmentation on the usual photography subjects, would rather not publish them. I'm not holding my breath for video Recommendation models from TikTok or Facebook either
For some reason this tool makes a slightly smaller mask than is found in the original image. So when you copy the masked area back into Photoshop, it doesn't match. Almost there, but not quite.
As a polished product & ease of use, still has ways to go.
Apple’s cut&paste of subjects in photos, while lacking SAM’s generality, is masterfully executed - even my 70yr old mom can use it.
You could bring the cut-outs back into Photoshop. I tried that, but this SAM tool reduces the size of the cut-outs slightly, so the cut-out won't match original image dimensions.
Thanks - like I wrote, the demo was DOA for me.
>and has nothing to do with Adobe.
I know. Masking and selecting is something you do in Adobe products and Adobe will be coming out with their own version of this (if I were a betting man)