I recall that I had some issues. After fiddling a while, I disabled all the 4K options, etc. from camera and selected 1080p output. This is cropped down to 1280x720 output resolution. It's enough.
Generally speaking, video resolution is not important for the sense of quality. Just have decent lighting and don't frame your face like an idiot. To get sense of quality, I recommend investing a little bit to a microphone. People don't realize it but audio quality is the king.
Also, it's likely that his camera is mostly at fault; OBS with my Logitech camera doesn't have any significant lag. OBS itself has very little lag.
Or does one of you mean that even the constant delay is significant enough to be noticed on top of the video call latency?
I just realized that I don't know how the missed/skipped counters are actually defined--presumably frames are skipped for encoding if the encoder can't keep up with the framerate and missed for encoding if if any other part of the pipeline preceding it cannot.
That's what I meant. Video call latency is noticeable but in most cases is not detrimental to conversation (or I have gotten used to mentally correcting for it). Colleagues who have used virtual webcams often have a noticeable additional 100-300ms steady-state latency which goes over the threshold for easy conversation, for me at least.
My experience on an older Mac (late 2014) was that even at 320x240 I'd get non-negligible delay. I ended up getting a Magewell Capture. As far as I know Magewell and Blackmagic are the only HDMI to USB devices that perform hardware h.264 encoding. It works perfectly and even has configurable video processing options, but I can't currently use it with OBS without introducing a delay.
Regardless of whether newer machine might make OBS encoding delay negligible, I wonder if a layered approach might work: I'm not well versed in video streaming technologies, but perhaps _layered streams_ approach which I think is supported in some video formats could allow OBS to work faster as a tool that combine multiple encoded video streams into a single video. One such stream for example might be the main video camera + additional say keyboard facing camera and finally a text layer which should be much faster to encode than having to re-encode the entire output.