Neat! When you say that "CPU is very slow", do you mean that using the CPU is (and always will be) slower than using the GPU, or that you could still optimize it further?
The current implementation is very slow with a CPU. Probably, it could be optimized with parallelism for batches. I'm not sure how it could be improved for single embeddings.