That being said, I like the use-case and it seems like nice work in a short amount of time. I have a couple of questions that I am curious about
1. Can you elaborate a bit more on your fine-tuning process? Did you "just" feed the model a bunch of regular emojis? Have you considered using any RLHF/DPO approaches?
2. You mention you generate a vector emoji. As far as I know the flux model just generates bitmaps, how do you handle that conversion?