So absent a microcode update that outright fixes Meltdown, there will always be some level of slow-down for vulnerable devices. System calls now jump from user mode code to a stub kernel in "supervisor memory". The stub kernel then does a full context switch (touching %cr3 paging register and wiping a good portion of the TLB), and once the real kernel finishes, it does a full context switch back to the stub kernel. It's all terribly inefficient, and realistically it's unlikely that there will only be negligible performance impacts. It should also be noted that this "work-around" doesn't fix processor, it just makes it so that that there's nothing juicy in the supervisor memory.
You may have to learn to live with this for a while. Even if it takes Intel a month to design and validate a fix for Meltdown, prototype and mass production turn around times mean that no customer will have a processor that isn't vulnerable to Meltdown until April-June 2019.
The performance loss comes from extra overhead on syscalls, so it could be sidestepped by allowing programs to do more work per syscall.
At its simplest that could mean adding more syscalls that perform the same operation over an arbitrarily long list of inputs (like linux's sendmmsg) but I would like to see kernels take inspiration from modern graphics APIs that allow arbitrarily long lists of arbitrary operations to be batched and executed with a single syscall. GPUs had this stuff figured out years ago.
2) SQL transactions as prior art? If they’re successful w their patent, I sure hope there’s a way for BSDs to try it on.