An interesting approach: giving each thread its own "young generation" sub-heap, so transient objects can be disposed of without coordination from other threads / CPUs and their cache pages.
A group at Intel independently came up with a similar approach as well: http://portal.acm.org/citation.cfm?id=1133967
E.g. - I have a test-driver here: http://github.com/roboprog/buzzard/blob/master/test/src/main... (although I have barely started the library I was tinkering on)
Any example client program for your allocator? I'd like to see what use cases you are handling.