In my opinion in most cases where you might want to write a project in two languages with FFI, it's usually better not to and just use one language even if that language isn't optimal. In this case, just write the whole thing in C++ (or Rust).
There are some exceptions but generally FFI is a huge cost and Python doesn't bring enough to the table to justify its use if you are already using C++.
if you want multiprocessing, use the multiprocessing library, scatter and gather type computation, etc
Typically Python is just the entry and exit point (with a little bit of massaging), right?
And then the overwhelming majority of the business logic is done in Rust/C++/Fortran, no?
That is probably why his demo was Sobel edge detection with Numpy. Sobel can run fast enough at standard resolution on a CPU, but once that huge buffer needs to be read or written outside of your fast language, things will get tricky.
This also comes up in Tauri, since you have to bridge between Rust and JS. I'm not sure if Electron apps have the same problem or not.
https://pythonspeed.com/articles/python-extension-performanc...
You can avoid that problem to some extent by implementing your own data container as part of your C extension (the article's solution #1); frobbing that from a Python loop can still be significantly faster than allocating and deallocating boxed integers all the time, with dynamic dispatch and reference counting. But, yes, to really get reasonable performance you want to not be running bytecodes in the Python interpreter loop at all (the article's solution #2).
But that's not because of serialization or other kinds of data format translation.
For 99.99% of the programs that people write, the modern M.2 NVME hard drives are plenty fast, and thats the laziest way to load data into a C extension or process.
Then there is unix pipes which are sufficiently fast.
Then there is shared memory, which basically involves no loading.
As with Python, all depends on the setup.