Salsify – A New Architecture for Real-time Internet Video
snr.stanford.edu
snr.stanford.edu
The Salsify implementation in the paper has slightly more accurate way of producing one single frame as it encodes two frames with different quality targets and takes the largest one below the frame size target.
Longer answer: The codec needs to support it. Codecs actually allow prediction from multiple reference frames, and maintain a buffer of them (2 to 16, depending on the codec, profile, and level). An individual frame may refer to several (potentially all) of those. So re-scaling up to 16 frames for every frame you decode will get quite expensive, not to mention the generational losses of doing this repeatedly for every resolution change. In practice what happens is you scale individual blocks when they get referenced by the current frame. But that has to be integrated into the motion compensation routines of the codec.
Both VP9 and AV1 support this, for example.
Very interested to see what they cook up (and kinda envious I didn't have the idea / don't have the space in my life to have a crack at it myself---it sounds very interesting).
Another funny thing about Fox Matchpass streams is that when watching it live, somewhere around the 85th minute, the stream magically jumps back to the very beginning of the broadcast well before kick-off. I have to manually click the 'Live' button to get back to it. This one is consistent, and odd. It's almost like some test code got left in, and nobody has noticed/reported/etc.
"6.1 Limitations of Salsify
No audio. Salsify does not encode or transmit audio."
Claiming that you beat a bunch of codecs that have synchronized audio (even though they disable it) is kind of misleading ...
If you wanted to add audio to Salsify, you would want to control a receiver-side video and audio buffer to reduce audio gaps and keep a/v in sync during periods of happy network, but this is unlikely to affect the system's ability to recover more quickly from glitches or to avoid building up in-network queues that delay audio and video alike. If you watch the video (or see Figure 6(f), Figure 7, and Figure 8), I don't think there's much reason to think audio can justify what the Chrome/webrtc.org codebase is doing -- WebRTC's frame delays are distributed over a broad range (so it's not like they're synchronized to some fixed timebase either) and are very high, especially in the seconds after a network glitch.
More to the point for our academic work, it would have been trivial to add shitty audio that made no difference to the metrics. The hard-but-necessary part is in designing an evaluation metric to assess (1) the qualify of the reconstructed audio (including how many gaps/rebuffering delays were there when the de-jitter buffer went dry), (2) the delay of the reconstructed audio, keeping in mind this is not constant over time, (3) the quality of the audio/video synchronization, which also will not be constant over time. Then measuring that in a fair way across Skype/Facetime/Hangouts/WebRTC/Salsify, and then trying to decide which compromise on those three axes is desirable. Somebody should do all that work at some point, but it's a major piece of work to bite off and pretty far from anything we've done so far.
Any reason you didn’t choose to start from VP9? Is the encoder still too slow overall?
From the FAQ:
> Why the name “Salsify”?
> It's not a very interesting reason. Salsify comes from an older project called “ALFalfa,” for the use of Application Layer Framing in video. Alfalfa gave way to Sprout, a congestion-control scheme intended for real-time applications, and now Salsify, a new design where congestion-control (in the transport protocol) and rate control (in the video codec) are jointly controlled. Alfalfa, Sprout, and Salsify are all plantish foods.
The company you linked seems to be meant as "sales-ify".
Salsify is led by Sadjad Fouladi, a doctoral student in computer science at Stanford University, along with fellow Stanford students John Emmons, Emre Orbay, and Riad S. Wahby, as well as Catherine Wu, a junior at Saratoga High School in Saratoga, California. The project is advised by Keith Winstein, an assistant professor of computer science.
Salsify was funded the National Science Foundation and the Defense Advanced Research Projects Agency (DARPA). Salsify has also received support from Google, Huawei, VMware, Dropbox, Facebook, and the Stanford Platform Lab.
Financially supported by the government, tech juggernauts, and executed by top tier doctoral students + a high school student + a top tier university professor.
Assuming this could be game-changing innovation to further advance worldwide communication, it's refreshing to see the positive externalities of a combination of capitalistic (F500 tech co's) and socialistic (university, government) systems executed by a seemingly diverse set of actors.
Really exciting work.
Encoding multiple versions of a video and picking a smaller one in response to congestion already happens for video-on-demand (think YouTube and Netflix videos) in DASH. That said, with VOD you can encode the video slower than real-time.
I can't imagine this ever making it into Skype/FaceTime/Hangouts/Duo. The big corps will probably continue to focus on "more internet" (fiber optic, zero rating, wi-fi hotspots, and internet traffic management practices).
The real trick is balancing all such product concerns in any next gen.
> Standardize an interface to export and import the encoder’s and decoder’s internal state between frames!
Can't this be achieved using sandboxing/emulation/VM techniques?
The idea of using layers, though, is much older (I remember reading papers about this already back in 2001 or so)
No.
Are you sure? Your website looks like a startup company’s.
It's just the HTML template! They all look like this. [...]"
Brilliant
Oh