Inferencing can be done in software entirely (e.g. INT8) but it's very slow compared to GPU or APU. nVidia cornered the market because everything (tensorflow and everything after) is optimized for it, but you can get good results on AMD now, and on ARC too in some cases. And slow results entirely in software (CPU-RAM), which for personal and non-constant use may be just fine too.