If the fast path is small enough, you probably want to inline the function, but the slow path makes it large enough that the compiler doesn't inline it (for good reasons).
If you force the slow path into a separate function, then your function becomes fast path + a call instruction, and it can be inlined.
I've often done this manually, and this is a very cool trick that I'm going to adopt immediately.
EDIT: plus, it's also beneficial for instruction cache/iTLB, as cornstalks pointed out.
Here it is on godbolt: http://goo.gl/iPCiSy
In practice, I'd be surprised if you can come up with a case where this makes a significant difference. If you are on a hot enough path for this to matter, on a modern processor you probably are running out of the even lower level decoded µop cache, which doesn't cache µops for branches that are not taken. If you aren't in this cache, your efforts are probably better spent making this happen.
Edit: ot's comment about how this affects the size of the parent function and whether it will be inlined is a good point, and might well make a measurable difference in the cases where it is true.
Whether or not this will impact performance depends, and will need some careful profiling. But I can imagine some situations where keeping the "hot" instructions in the cache and the "cold" (error-handling) instructions out of the cache could be beneficial.
Surely that's a decision for the optimizer to make, in the case of __builtin_expect?