Quite a few. We're often affected with VLC, and code execution is easy to get to. But with VLC, you're "only" in userland.
Quite a few. We're often affected with VLC, and code execution is easy to get to. But with VLC, you're "only" in userland.
(I know you know that, I'm just spelling it out for people).
Especially since you can see often fake videos circulate around, that tell you to download a special software or codec to play them (with malware, of course). WMP also had a scripting/runtime system that could be abused greatly.
Be careful what you download :)
A sandboxing system provided by either ffmpeg or VLC would be a very good idea, though it would be some work... encoded data in, decoded frames out via shared memory. Negligible performance impact.
Given that there are apparently thousands of bugs in the video parsing code, it seems like a no-brainer.
Section 5.2 in this DJB paper talks about (portable) isolation of plain transformations. Video playing is already close to "pure" or could be made pure pretty easily.
Not sure if serious or sarcasm, to be honest, since this seems very far from what we see.
A media player is not simple to sandbox, (as the MacOS X sandbox showed us for example), because:
- you need to open files by yourself, without user interaction, to support playlists,
- you need to open connections by yourself to support video protocol like RTSP, RTMP, RTP,
- you need raw device access to support Webcams, Capture devices, DVDs, DVB tuners,
- you need to access GPU buffers for direct rendering, and/or shaders to do fast filtering or just plain chroma-conversions,
- you need to be able to access the audio output, at low-level, for libsync which is not always doable with the simplified APIs,
- and I don't understand what you mean by "very limited gui input"; how is that less than other programs?
Sure, it can be done, with performance costs but it's clearly not a "no-brainer".
http://www.chromium.org/developers/design-documents/multi-pr...
The entire app isn't sandboxed -- just the code that does video parsing, i.e. with the thousands of bugs and hundreds of remote code execution exploits (!).
See my other comment on this topic. ffmpeg is already very modular, and used in many video players (user interfaces), so this separation is more than natural -- it already exists in the codebase.
BTW, some people seem to be unfamiliar with the multiprocess/Unix design approach (usually people with a Windows background, which I came from as well). I recommend http://www.catb.org/esr/writings/taoup/ for a great intro to this design philosophy.
The exploits are usually on the protocol (access) level, the demuxer (format) level but also the decoder level.
While in theory the first 2 are what you call parsing, many security issues appear also at the decoder level.
And if you want to split the video decoder from the rendering, you need to introduce an additional memcpy (or two) of full decoded frames, which has an important impact.
I agree it would be nice to try, with a correct, separated processes architecture, but it means changing also a bit the usual architecture, where there is a direct-rendering between the decoder and the output.
You can open your X11 socket (or whatever your GUI uses) in the I/O process and hand it off to the renderer before dropping privileges, and hand it off without any loss of performance. Another thought -- although I don't generally touch 3d rendering, or know enough to know if it's possible off the top of my head -- is something like cross-process pbuffers: you render into one, and composite the result to the screen from another process. Since your OS already does this sort of thing (X11 compositors, for example, do this sort of cross-process compositing), this can't be too expensive.
You might need to use the XACE extension to restrict the renderer to just the rendering output window in case it gets exploited, so that you don't have access to the rest of the UI.
The places where vulnerabilities can do damage are when communicating with the rest of the OS: File I/O allows you to get user data, permanently install malware, and delete things. The window system allows you to snoop keystroke/mouse data, synthesize input, and watch what the user does. Drop privs on those, and your app is fairly effectively sandboxed.
Off the top of my head:
>video output
Meaning: RW access to /dev/dri/cardX on desktop. In an embedded system, replace that with /dev/fbX. In most cases, we're rendering to GL buffers. Do you trust the GL drivers to be completely bug free? (Hint: from some embedded drivers I've dealt with I find it remarkable they work at all.)
>doesn't need to [...] open network connections
How about RTMP streaming? Or DLNA discovery and playback?
Some of the cases may be protected with sandboxing, but I have serious doubts about getting anything near complete coverage.
I think you could also provide less than full access to the graphics driver. You could have a file descriptor or shared memory protocol, with the sandboxed process outputting video frames, and the parent process actually communicating with the driver.
In addition to being more secure, it's also better software architecture. VLC/ffmpeg are already way more modular than say the Microsoft equivalents.
Memcpying complete video frame in HD at 30 or 60 fps is not negligible performance impact. But I agree it would be a good idea.
It might get hairier if the platform graphics api wants to supply you a buffer to render into, and you can't pass it a shared memory buffer yourself (that is accessible by that sandboxed process).