Just as an example of the difference - I wrote a naive implementation of ChaCha20 in C with zero optimization effort and it does 5cpb out of the gate (Sandy Bridge). Just using vector-types and letting GCC/Clang vectorize brings that down to ~3cpb on Sandy Bridge - no effort. The Krovetz implementation of ChaCha20 is closer to 1.2cpb on my machine, with AES-256 doing 1.0cpb using AESNI (again, my own naive implementation). All software.
Even the most hand optimized, secure AES software implementations are still in the realm of ~15-20cpb (IIRC), three-to-four times worse than the unoptimized competitor. As linked elsewhere in this thread by me, non-scientific tests show it 3x faster in software on some mobile phones. That's a lot of extra cycles-per-byte for your battery to chew through using AES-256, and I'd guess I easily churn through a low number of gigs of HTTPS data every month..