I'm coming from the scientific computing side of python, so beyond the BDFL's suggestions, I use the following protocol to make python fast. Note, that I'm doing numerical analysis and manipulations, so this isn't a general workflow for all python applications.
1. Write algorithms using numpy, attempting to not use too many crazy indexing/array tricks. You can generally avoid costly for loops using broadcasting.
2. Profile code using cProfile and Robert Kern's line_profiler
3. Judiciously optimize by re-writing computationally intensive methods in Cython. Numba also looks like a contender in this arena as well, but I haven't done a benchmarked comparison.