How to Train an AI Image Model on Yourself
coryzue.com
coryzue.com
You should do the same with your training images. Caption everything you do not want the model to remember as "you" (what you're doing, wearing, accompanied by, accessories, etc).
That said I usually train with about 100 really varied images of the person so it tends not to overlearn any other particular thing.
Possible? Yes. Convincing results? Probably. Good idea? I doubt it.
Extra challenge: extend that to produce videos (e.g. via "live portrait" nodes/models), to implement the digital version of the magic paintings (and newspaper photos) from Harry Potter.
EDIT:
I'm not joking. This feels like a weekend challenge today; "live portraits" in particular work fast today on a half-decent consumer GPU, like my RTX 4070 Ti (the old one, not Super), and I believe (but haven't tested yet) even training a LoRA from a couple dozen images is reasonably doable locally too.
In general, my experience with Stable Diffusion and ComfyUI is that, for fully local scenario on normal person's hardware (i.e. not someone's totally normal PC that happens to have eight 30xx GPUs in a cluster), the capabilities and speed are light years ahead of LLM space.
Just for comparison, yesterday I - like half the techies on the planet - got to run me some local DeepSeek-R1. The 1.58 bit dynamic quant topped at 0.16 tokens per second. It's about the same as it takes a SD1.5 derivative to generate me a decent-looking HD image. I could probably get them running parallel in lock-step (SD on GPU, compute-bound; DeepSeek on CPU, RAM-bandwidth bound) and get one image per LLM token.
I only use it on Windows with Nvidia GPU, but it should work both on Windows and Linux with CPU only and with Intel GPUs, as well as on Linux only with AMD. Though skimming the README some more, I also see Apple Silicon section, and one called "DirectML (AMD Cards on Windows)", so maybe AMD+Win works too.
As for use: you install ComfyUI from the link above, and then this:
https://github.com/ltdrdata/ComfyUI-Manager
to have UI for searching and downloading custom nodes (instead of having to install them by hand), and you're good to go.
Thank you from people holding NVDA.
- I asked grok to generate a list of racey prompts. - Has replicate generate them via script. About 10-20% are very poor, I filtered those out manually. - It also has NSFW guardrails, but a simple retry or word juggle gives you a chance to get around it.
I think I spent $10
Often the innovations from that world are ahead of mainstream AI research by years. You should see what coomers did for LLM sampling in order to get over issues with "slop" responses just for their own pervy interests. This is a full several years before the mainstream crowd ever cared.
Look mom, I can make some cool astrology images for you! Whoops, that's boobs. That too. And this one. Ehh, hold up, I need to add a pile of negative prompts first...
Even if we assumed equal amounts of effort, it wouldn't be surprising if a large corpus of nude images in the training data improved model results.
But maybe we should have better negative prompt presets for different levels of decency
There wouldn't be depth data so it would be inferred from shadows
Also, Kling 1.6 Elements works pretty okay if you use the same person/face for each element.
Kling also has lip sync.
Or this lip sync with replicate: https://replicate.com/bytedance/latentsync
Or there is HeyGen or D-ID or Synthesia, or tavus.io for full interactive digital twins.
Edit: when I say "nefarious" I mean you can use that tech to impersonate someone (eg. political reason) but for my case it's more the creeper type cloning someone for personal use eg. Replika
Tangent, the holo vtubers industry is interesting since they build up these characters with some unique persona/theme and then people follow that specific model, they could make themselves into an AI easily since it's a rigged 3D asset but of course it would be boring compared to the real thing
The most popular vtuber on Twitch is an AI tho
I'm not sure if that's truly AI since the Turtle drives her
Edit: if the source was open I'd believe it
It has always been a LLM. There is no human typing at insane speed to the TTS.
edit: but yeah the fact that so many people interact with her shows generated content can keep people occupied
Thanos/NFTs: where did that take you? right back to me
Thinking hardware with built in chain interface for proof
Oh man dating apps too
That's true love though, two people meet up IRL they're both like wtf who are you