The process of cooking it down that way is creating a derivative work and copyright still applies, it just gets
more complicated. (Questions of "transformative" versus "fair use" versus license complications such as "no derivatives" and "share-alike" and much more.)
It's going take to years, if not decades, to sort out the copyright mess of machine learning. I take an extremely cautious approach to it myself, not because I think that the laws or the authorities will "crack down" on it any time soon, but because I ethically want to do the right thing (and took ethics oaths in the past, unlike so many in this industry; it's not fear of future enforcement, it's promises to myself that I'd try to do the right thing).
At the end of the day, it's about respect: I can't tell you what you are doing is illegal or not (and probably not even a lawyer can tell you for sure until far more court cases have tried or worst case you are at the end of your own personal trial on your specific use). I can point out that blog articles come from writers. Even writers not making money on their writing have implied copyrights. Many also have explicit copyright statements on their websites they blog from and sometimes also in the metadata of their RSS feeds (and you can't assume the lack of one implies the lack of the other). Respect the writers as best you can. Don't just treat them as faceless metrics to be compiled into your models. That's your responsibility in training an ML model. I would suggest doing it cautiously, but that's my sense of ethics talking, I cannot tell you what your ethics should be here in this moment, only suggest what I find ethical or not and that I'd probably be a lot more cautious.