The problem wrt open source / free solutions is data. Kaldi is open source and gets state of the art results -- but the data costs a lot of money. Training the models is doable on a commodity GPU although it takes quite a while.
https://github.com/kaldi-asr/kaldi/blob/master/egs/fisher_sw...
Kaldi hasn't been in first place on that dataset recently, but it was a few years ago.
On other more researchy datasets (eg. for distant speakers or languages other than English), the best system is often based on Kaldi.