Background Matting: The World is Your Green Screen
grail.cs.washington.edu
grail.cs.washington.edu
- people who are using occupying a wider z-axis (for example leaning forwards in the camera or who have arms in front of them)
- people holding objects like cups
How well do your method handle those kind of situations?
The neural network used seems mostly for allowing for variations in lighting and dealing with the fuzzy effects you'll usually get around the subject. Depth of the subject doesn't appear to be relevant here.
(wouldn't it be nice if every time a research topic pops up here that there would be a small list of the essential keywords to look for more background information?)
I thought GP was asking for an implementation of background removal without a green screen.
If you have a background picture, you have all the info you need to identify your subject - just plain subtraction. I think this is what the Photo Booth app on my circa-2012 MacBook does, quite effectively.
A good intuition is that if it were easy to do it already with any background, professional studios wouldn’t be spending so much money on green screens. Background subtraction is pretty poor in general without very constrained setups. Our goal is really to provide professional quality without any of the equipment.
And yeah, it requires constrained setup and a lot of additional work, because even before you "subtract" the background you have to think about lighting (your demo video might have very nice background matting, but the lighting is off, so it's relatively useless except for toy applications (which there are a lot).
Also: did you compare somewhere withe the very basic fixed-exposure method? Beause for fixed exposure, background and camera placement I suppose this should work just as well... Still I think this is a really cool project, I didn't get disappointed like with the last link of that sort, where someone tried the same with horrible artifacting.
Green screens are crap with hair, because its translucent the green/blue bleeds through which means that it has to be cleaned up by hand.
Then there are the situations where there isn't a green screen. Again manual cleanup is required. Each frame needs to be cut out by hand. 24 times a second.
The same with a difference matte. Cameras are noisy, so there is constant noise in the alpha channel. This makes the effect look wobbly and cheap.
What this method does is pull a key from a difference matte, and makes it look good.
I can assure you they read these papers. If it is good, it will be part of a future version of the software.
The foundry have been trying to get this into Nuke for years. The problem is that normally you get flickering, as you'll have seen.
It's not really "just plain subtraction", it's keying. Which AIUI basically means setting the alpha according to the difference between the image and the reference.
Green screen works well for this because, excepting Zoe Soldana, people tend to hang out around the opposite side of the colour wheel, so there tends to be a good distance between foreground colour and background colour. If you're trying to do this against arbitrary backgrounds, you seemingly need to augment keying with additional techniques like image segmentation to get good results.
Background subtraction methods on the other hand usually fail if the camera moves even a tiny bit or the lighting changes slightly. More advanced methods can recover eventually, but you still get a few frames with improperly removed background.
I mean, it's trying to do the same kind of thing, but this looks to be a lot better at it.
https://github.com/senguptaumd/Background-Matting, which points to
https://grail.cs.washington.edu/projects/background-matting, which points to
https://arxiv.org/abs/2004.00626, which points to
https://arxiv.org/pdf/2004.00626.pdf, which is inlined at the originally submitted URL. I'm not sure what's going on here, but on HN the convention is probably to link to the project home page first, and after that maybe the Github page and if neither of those exist, to the arxiv.org homepage (but not the pdf since those change with each revision). So I've changed to the project home page for now.
The link share here was likely with keeping the relevance of this project to HN in mind, and that easy access to the code and authors would be valuable for anyone here looking to take it further.
Thanks for clarifying the convention here on HN, being transparent, and for updating accordingly. Much appreciated.
Always open to feedback if you have any as well! :)
Also, the submitted title ('Zoom’s virtual background swap but better. DL+GANs for background replacement') was too promotey.
FYI your first sentence effectively says > I wasn't suggesting anything, but I was suggesting exactly that.
I had a meeting where one participant uses an actual green screen, and the difference was remarkable, with none of the issues above.
I've seen someone advised that the background on their webcam makes them "look poor", where the concern was that looking poor is a (perverse) impediment to getting paid work, but they can't exactly move, especially under lockdown. It may be better to use a calm artificial background in that case.
See also people doing online-conference presentations and Youtube videos. I've seen quite a few of those are using virtual backgrounds.
Perhaps for the same reason - thousands of people may see the video, and some people, having made the effort to put on a nice suit/makeup/etc, get a haircut, and look their best, don't want thousands of people to see what their not so nice home looks like behind it.