We're fay away from it now, but I've seen less sketchy solutions being implemented.
We're fay away from it now, but I've seen less sketchy solutions being implemented.
That being said these shitty video models I believe are just an arms race between Meta and Google after the release of stable diffusion. Microsoft has a video version of CLIP that I believe will really change the game, but unless you have trained a model with video embeddings it's all going to look devoid of any narrative. Right now the models just look like a sequence of images with the same promt and some sort of continuity to make it look more video like.
Also, there's some interesting work with ML taking diffused light from around a corner and recovering the original pre-diffused silhouette.
In many ways, this is how we've learned the visual cortex is working.
The amount of actual neutral data you are seeing is way less than you'd think given your perceived visual fidelity.
The only practical issue is that distribution of AI hardware in consumer devices is going to noticeably lag behind POC on compounding cutting edge hardware in research environments, and no one wants to invest into obsolescence.
Maybe it will happen in the cellphone market though given the hardware refresh rates from carrier subsidies.