Interestingly enough, it is possible to do private inference in theory, e.g. via oblivious inference protocols but prohibitively slow in practice.
You can also throw a model into a trusted execution environment. But again, too slow.
https://pasteboard.co/k1hjwT7pWI6x.png
reach out if interested in collab.