Open-Source Virtual Background
elder.dev
elder.dev
It's got the option for "real" chromakey, but like the author, I don't have a green screen, Amazon isn't scheduling any deliveries for another month, and I don't feel like a trip to the fabric store would count as "essential" travel (especially if it's only so I can screw around with stupid backgrounds).
Tried a few different sheets/blankets I had at home, but none are a suitable color or uniform/matte enough to work well, even with proper lighting. I admit this is such a non-issue and only something I want to play around with, but it would be fun nonetheless.
I, too, have been playing with the chromakey in OBS, and like you tried a bunch of different sheets and blankets. The one that worked the best I discovered by accident; my wife was pulling out a puzzle, and with it she pulled out a green felt puzzle mat (really just a large rectangle of green felt). It was perfect. But I had to get it from her so I could use it on my meeting before she started her puzzle. After some brief, tense negotiating, I was able to snag it. All I have to do is the dishes for another month!
It works really well, but this would be way better and easier to set up.
Likewise there are loads of cheap ones on Amazon or I could even get some bright paint and do a wall (or maybe dye my current non-chromakey white sheet background). It's just a matter of suspended/delayed deliveries and state orders to avoid unnecessary travel--including shopping for things like home furnishings.
It makes sense. There are plenty of important reasons why I should stay home and Amazon should focus on delivering essentials. I guess it's the definition of being spoiled/privileged that this is what I'm musing about.
I was fortunate enough that the previous owners left behind a god-awful teal paint that works amazingly as a ‘blue’ chromakey.
For a huge upgrade in camera I’m using the OBS Camera app on iOS which wirelessly beams an NDI stream to my desktop.
For lighting I’ve just played around with various lamps from my house.
The one thing I did order was a USB microphone because I couldn’t stand the thought of wearing a headset or adding latency with wireless earbuds.
Cam is the basic 1280x720 Microsoft Lifecam USB webcam I've had for years but is still much better than any laptop cam due to placement flexibility and better optics. OBS lets me color correct and exposes the manual controls for focus, exposure, etc.
I did grab a USB headset from work before work-from-home started, but I found it had a frayed wire and isn't any good. Instead I am using my (wired) cell headphones with their integrated mic instead of the USB headset.
I dug out an old USB audio interface I had packed away and tried it with a Shure vocal mic, but that led me down another hole of messing with Voicemeeter to tweak EQ and noise reduction because it's really meant to be held right up to your mouth (for singing, etc) and picks up background noise if I have the levels up high enough to use as a desk/stand mic. For the time being I am sticking with the headphone/mic for simplicity's sake.
None of this is what I'd use if I had been handed a couple hundred bucks to put together a VTC setup but it's all stuff I had laying around and looks/sounds so much better than everyone who's just using built-in laptop cam/mic/speakers.
You can still try. I ordered some ink (which wasn't deemed essential by Amazon sadly, but in my case, it wasn't that essential, it could wait a month) 3 days ago and I just got it today (I saw yesterday that it was updated for tuesday, but got it way earlier, most probably because of easter holidays). They clearly give priority to essential stuff because it was still longer than usual, but they doesn't seems to have a huge backlog of essentials stuff either.
Injecting this into a web client seems like the sweet spot effort wise.
I've always been disappointed that BodyPix and some similar models are mobile-first and mobile-only on TensorFlow (https://blog.tensorflow.org/2019/11/updated-bodypix-2.html) -- are these models just not used much on server-side settings? There seems to be very little documentation on doing this server side.
But i am really impressed by the detection of Microsoft Teams, i would love having this quality and speed available in a browser.
Chroma keying is the stable version of this idea: the single background color of a separately and well lit background removes these issues.
I did actually tinker with this and other approaches a little.
Also the things in the background get moved around pretty frequently, e.g. what is on the couch. I don't have a dedicated office room
If you have the space, and good lighting it's a much simpler approach.
It's also less fun though, and I could do this with what I had on hand pretty quickly.
Random example clip: https://youtu.be/T5uqB1Kqukw
Do you have a github/gitlab/something else link?
I'm also using this with other apps for fun though (duo, hangouts, etc.)
[1]: https://support.zoom.us/hc/en-us/articles/210707503-Virtual-...
There's one thing I'd love to achieve though, which seems not possible on Linux desktop (specifically kubuntu)....
I want to be able to use the loopback as the source for screen share instead of webcam; i.e. to use the loopback as the conference presentation.
Has anyone got any ideas how to achieve this? Given most conference solutions on Linux do not seem to support either 'share this window' or 'share this screen region'. It seems to be the whole desktop or nothing.
Yes. Thanks for giving me a reason to write this up.
1. Download and install OBS. OBS will be your video processor; among other things it's super easy to make it capture the whole screen or individual windows.
2. Install the v4l2 loopback kernel module[1]. This makes it possible to have a virtual webcam. On Ubuntu 19.10, this was as easy as apt install v4l2loopback-dkms and then modprobe v4l2loopback.
3. Install the OBS plugin obs-v4l2sink[5]. This exports the OBS output to the new virtual webcam device. I just installed the deb file provided by the project[2]. In OBS, under Tools, select v4l2sink and Start.
That's all I had to do. Surprisingly straightforward. At least Chrome and Firefox[3] will now pick up a "Dummy Video Device" webcam that streams the window, or whatever scene I set up in OBS.
In my case, the primary advantage was that this virtual webcam is streamed in Jitsi Meet at a higher quality/framerate than the regular desktop share feature. It's also much lower latency than both Twitch and Youtube Live streaming (Jitsi Meet/WebRTC: <1s, Twitch: 5s, Youtube: 15s[4]; YMMV).
You also get to enjoy the rich feature set OBS provides for Twitch streams; for one thing, you can include the real webcam video.
Bonus: Desktop audio "just worked" in Firefox, which offers the pulseaudio monitor (loopback) device as an input. Chrome doesn't -- probably the intended behaviour. I'm sure there's a workaround.
[1] https://github.com/umlaeute/v4l2loopback
[2] https://github.com/CatxFish/obs-v4l2sink/releases
[3] For some reason, Gnome's Cheese won't
[4] Microsoft's Mixer allegedly has super-low-latency streaming (FTL protocol), but new account are cleared manually and I haven't had the chance to try
[5] For Windows, you can use OBS Virtualcam https://obsproject.com/forum/resources/obs-virtualcam.539/
The problem is that v4l2loopback only provides a virtual _webcam_ (video source), not a virtual _screen_ - the two are different and are handled differently both by browsers (webrtc) and desktop conference apps (Slack, Teams, etc).
I guess the other issue is that conference apps treat webcam and screen capture differently; usually if someone is sharing a screen, then that feed takes over the full view for all participants so they you can actually read the content.
I don't want my screen recording only to show up in my _webcam_ view, which is usually just a tiny thumbnail.
In particular, one I've found very useful with Zoom is being able to zoom in to a small region and scroll around. I also suspect Zoom prioritized resolution (for content clarity) over frame rate for screen sharing, which probably doesn't apply when it's just a "webcam" in the eyes of the client. I'm guessing your window capture would get decimated in terms of quality.
It seems like moving all that data backwards and forwards between Python and Node might be a bottleneck, no?
The inference / ml is expensive (which I did profile initially...), and I suspect not really optimized on this backend. It appears to be faster with webGL in the browser.
I sorta stopped worrying about it once it was "good enough" to show up to a few meetings with, but with all the attention I'll probably take another look.
It does look like someone ported bodypix to python, I'll probably try that next.
You can usually alter performance (with Bodypix that's an accuracy/speed tradeoff) or do something silly like downscale, run, and upscale the mask. I'd like to try this.
Amusingly I did some hacking on this and the current bottleneck is actually reading from the webcam which is capped at <10fps without doing anything else. Switching the capture to MJPG helps.
from keras.models import load_model
model = load_model('models/transpose_seg/deconv_bnoptimized_munet.h5', compile=False)
def get_mask(frame):
# Preprocess
frame = cv2.cvtColor(frame, cv2.COLOR_BGR2RGB)
simg = cv2.resize(frame, (128, 128), interpolation=cv2.INTER_AREA)
simg = simg.reshape((1, 128, 128, 3)) / 255.0
# Predict
out = model.predict(simg)
# Postprocess
msk = out.reshape((128, 128, 1))
mask = cv2.resize(msk, (frame.shape[1], frame.shape[0]))
return mask
The model file I got from:
https://github.com/anilsathyan7/Portrait-SegmentationIf you just want an easy greenscreen https://obsproject.com/ has a very good chromakey filter and a V4L2loopback plugin https://github.com/CatxFish/obs-v4l2sink
It's a couple of lines to use: https://pytorch.org/hub/pytorch_vision_deeplabv3_resnet101/
Do you have a live video recorded showing how quickly it can process a stream?
IIRC it's something like 10FPS currently which is sufficient enough for meetings so far (about 1/3 what you might get with sufficient bandwidth in most video conference tools).
There's definitely room to improve it.
P.S. What happens when they do e2ee on the webcam stream?
Obviously Zoom isn't end to end encrypted, it's client-server encrypted.
The Zoom client also requires minimal hardware requirements from the processor iirc.
Now, I know that there are many companies that force people to be on with a live video feed, and that many don't really like it.
How about recording a 3-min clip and playing that in an infinite loop - creating a fake feed (remember Keanu Reeves' Speed?) - so that people can avoid not being seen, but still get things done better? A mask on the face is a simple addition to avoid detection. As the saying goes, modern problems require modern solutions!
The virtualvideo readme has an example of looping a single frame. https://github.com/Flashs/virtualvideo#errorhandling
because for a scripting/backend focused language, react/vue is not even something you want to aspire towards.
It also can't make efficient use of resources, which explains it's lack of prevalence in embedded/back-end apps
I'm iterating on it again tonight, and tentatively virtualvideo [1][2] gives better results by just piping frames into ffmpeg vs pyfakewebcam
[1]: https://github.com/Flashs/virtualvideo/ [2]: https://pypi.org/project/virtualvideo/
I'll have to update the post or do a follow-up at some point.
The magic is pyfakewebcam and v4l2loopback, I was looking foe a way to turn myself into a potato on Teams. The bit I was missing was how to create a virtual webcam.
The web requests are just an easy mode of IPC to pass around some bags of bytes, "high frame rate" is at most 30 qps ... that part isn't really interesting performance wise and this isn't a production tool :-)
I'm not sure I'd be so confident about tensorflow.js being so fast on the CPU ... you can see a marked difference in the backends httpss://www.tensorflow.org/js/guide/platform_environment
Any follow up on making this more generic, e.g. with AMD/Intel setups?