The hidden cycle-eating demon: L1 cache misses
x264dev.multimedia.cx
x264dev.multimedia.cx
It required changes to processor microarchitecture, though.
Hence, I really appreciate it when a language offers some way to drop down to byte-level constructs and build space-efficient implementations where they're necessary. Improvements in memory compactness can often save you the effort of dropping all the way down to C.
Plus, a CPU can be used for things other than encoding video, while a hardware device is of course useless for anything else, so it's easier to justify spending money on a fast CPU than on a task-specific piece of hardware.
Also, it isn't really an edge case; it will occur in any application which has a working set that is unavoidably larger than the L1 cache. Video encoders are just one of many cases where this occurs.
(Note: added that last point to the blog post after I posted this.)
Thank you also for informing me that hardware encoders aren't very good. I had actually considered purchasing one of the ~$100 ones, but now I'll steer clear.
I must be doing something wrong, though, because I can't seem to get much better than 2x real-time on my Q6600 w/ 4GB memory when using x264.
x264 now has encoding presets you can use to trade off speed for compression: they go from "ultrafast" to "placebo" (full list in the help). Grab the latest from x264.nl; we've also had a lot of speed improvements lately ;).
Do note that at high speeds, it's easy to get bottlenecked: the most common case is the decoding of the input, which is often only single-threaded. If your input is uncompressed, then you can get bottlenecked by reading it off the disk. And if you're running filters on the input (e.g. resizing it), that can also serve as a bottleneck.
I think that quite a few "ordinary consumers" would share my opinion. If they don't mind spending hours or days compiling and tweaking a delicate chain of open source software, they're not "ordinary" - they're power users.
This was a while ago, but I have no reason to think the software encoders haven't improved faster than the hardware.
edit: rereading your post I see your main problem with the software solution is "spending hours or days compiling and tweaking a delicate chain of open source software". I can't help but note that if that is the case then you're doing it wrong™.
Most likely, but since ffmpeg has the least helpful documentation I've seen (for a newbie at least), I'm not surprised.