Yes that's the idea.
> Which would be slow I would assume.
How expensive do you think a trap is? It takes about the order of 10 billionths of a second.
> If you never switched back, wouldn't any process using advanced vectorized instructions (like anything using a decent libc) be permanently pinned to the large core?
I think you can switch back next time you schedule.
Ok, yes, then we're on the same page. I would still think that would be slow? You'd need a full transition-to-kernel and context switch before you could execute again, which AFAIK would take at least microseconds…unless you think there would be a faster path to resume execution?
No that's around 30 ns on modern hardware I believe.
are you sure about that? I would expect at least a couple of orders of magnitude more just for the userspace->kernel transition.
edit: for what is worth, a syscall it takes 250ns on my (admittedly vintage) machine. That's using the lowlatency sysenter path. An interrupt is probably going to cost more.
Anyway the cost of scheduling on another core is going to dwarf that.
edit2: for reference, this was a Sandy Bridge turboing at 3.5 Ghz during the test. With spectre mitigations on (which is going to be a good chunk of that overhead).