Some small-ish items:
-- Virtualization of rr's required PMU features on Azure and GCP
-- Virtualization of Intel's CPUID faulting on AWS and other cloud providers (Linux KVM has it, but it's not enabled in AWS)
-- Fix the reliability bug(s) in the retired-conditional-branches counter on AMD Ryzen so rr works there
-- Implement CPUID faulting in AMD Ryzen
-- Reliable instructions-retired counter on AMD/Intel (for more reliable/simpler operation)
-- Trap on LL/SC failures on ARM (AArch64 I guess) and other archs (so we can port rr there), with Linux kernel APIs to expose that to userspace
-- Similar traps for Intel HTM (XBEGIN/XEND) (less important now that's mostly been disabled for security reasons)
Bigger items:
-- Invest in kernel-integrated record-and-replay for Linux and other OSes
-- Implement QuickRec or some other hardware support for multicore record-replay on at least some CPU SKUs
(If anyone thinks they can help with any of that, talk to me!)
> what do you think is keeping the overhead so high in the average case described in the post (50%)?
I assume you've read https://arxiv.org/abs/1705.05937 ? For many workloads the major unavoidable cost is context switching. With rr's approach every regular tracee context switch (intra or inter process) turns into at least 2 inter-process context switches (tracee to rr, rr to tracee). This is bad for thread ping/pong-like behaviour. For many other workloads the cost is simply the loss of parallelism as we force everything onto a single core. For other workloads there is a high cost due to system calls that require context switches to rr because we don't currently accelerate them with "syscall buffering". The last one can be mostly engineered around, the former two are inherent to rr's approach.
Who do you work for? :-)