Then your (client's) problem wasn't reference counting but premature optimization.
Are there situations where you'd like to have code run fifty times faster than native Python. You bet there are, lots and lots of them - for example, in a Unix-clone Kernel. Sorry if somehow you didn't find one of them.
Hopefully you used Valgrind and Formal language specification to reduce the work required.
And to avoid premature optimization, use gprof to find the bottleneck(s) rather than just diving into what seems to need optimization.