Thank you. The engine is definitely domain-specific. This would apply to vocabulary and also inference. For example, if you want to use the tech for smart lighting then you would need a different model/context.
The demo is done with some noise and there is some reverberation as well. The speakers are also somewhat accented. That being said we would like to open-source a benchmark for this (similar to other products we have). The comment on accuracy is a bit tricky as it would depend on parameters you mentioned and specific task. I will provide more information when we open source the benchmark in Q1 2019.