This doesn't seem to test the GC, nor important operations such as dict lookup/iteration, set add/membership, list iteration and updating.
Is really a recursive implementation of Fib an enlightening benchmark? I don't think so. And bubblesort is just messing around with two pointers.
Why not, if you're first doing benchmarking, find out what the most time consuming popular operations/algorithms are, and then use them?
Now it seems that you are just testing two obscure algorithms that are not really representative for Python coding in general.