You're ultimately right, though, and that a true HD is only going to come from the raw film content. What the neural network gives us are essentially plausible higher-res hallucinations.
Edit: as per the other comment, if the original exists only on video and not film, perhaps this is the best we're going to get.
There are entire catalogue of overlay comparisons of different releases, encodings etc. [0].
Example: http://compare.bakashots.me/compare.php?setId=3896&compariso...
The main difference here is that the interpolation algorithm on your TV is online. It's handling 30 frames per second, over 9 million pixels per second. Doing the interpolation offline (ahead of time), you can take as long as you want, look at multiple frames to try to make better guesses, try multiple things and use some fitness measure to pick a winner, even a frame or a pixel at a time.
It's still interpolation.
In this case they're using machine learning to add additional information about textures that isn't in the footage broadcast. They can add frames by interpolation, but the ML texturising and detailing is not interpolation.
Starting with a blob, if you interpolate you get a smoother blob, with this process you get a more structured figure.
It can still look nicer than naive upscaling though.
> ... interpolation is a method of constructing new data points within the range of a discrete set of known data points
I don't think that's quite right, at least it doesn't jibe with what the DS9Doc people have been doing (which consists partly of remastering pieces of DS9 scenes):
https://www.indiegogo.com/projects/what-we-left-behind-star-...
I think the footage really was on film, but the issue was that it was composited with low-quality CGI effects, or something like that. So you can rescan the film, but you have to redo all the compositing (and probably with your own models because I'm guessing the original CGI didn't look that good). That's why a DS9 remaster is so expensive.
In the fully general case of arbitrary video this is true, but in practice it isn't.
You can gather information over time to do superresolution, and if you want to get super fancy you can build a world model (e.g. get more information about what an actor's face looks like from a close up shot, and apply that knowledge to less detailed shots).
I expect ML based upscaling to eventually produce some truly stellar results.
i imagine upscaling the Phantom Menace podrace to 4k, and giving the model a bunch of NEW rock texture info to use to create new detail.
Kinda of an automated way to combine these two ideas.
http://www.framecompare.com/image-compare/screenshotcomparis...
https://techcrunch.com/2019/03/18/nvidia-ai-turns-sketches-i...
What you dont want is every upscale to start looking homogeneous, so it would be best for a design team to specifically map old to new texture sets, giving each upscale a unique look.
In DS9 / Voyager as I understand it, they transitioned from using models to directly generating the whole shot with CGI at NTSC resolution. See https://memory-alpha.fandom.com/wiki/CGI#Acceptance
This means that in TNG they just had to recomposite the film with a newly created high-res phaser shot; whereas for DS9/Voyager they would have to recreate the whole shot in high-res CGI.
"They’re using the original Lightwave scene files for camera and model movement, lights, etc. It’s also the original 3D models and textures used on the show – and nothing has been updated in any way other than being rendered out at 1920x1080. It’s the raw CGI without any post work."
Well, that is the difference between regular upscaling and upscaling with neural networks. With a neural network, the additional information is being stored within the network during training and added to the video during the upscaling process.
Ultimately, you could argue, that this is just interpolating too, but the quality of the interpolation depends on the training material. If you would train it on an original and use such a neural network to upscale a lower resolution version you could end up with the original (a perfect interpolation).
So it all comes down to the quality of the training material and while AI Gigapixel seems to have quite good material, I wonder if the result could be improved by transforming the video as a whole and not just frame by frame, as that would give the NN even more information to interpolate on.
I've seen a few people say this in the thread. It doesn't seem accurate to me. Information is being created/hallucinated/interpolated. Re-scanning filmstock gets new information. If not, it doesn't matter how sophisticated your algorithm is (naive upscaling or deep learning), you're still interpolating values.
"While the popular Original Series and The Next Generation were mostly shot on film, the mid 90s DS9 had its visual effects shots (space battles and such) shot on video.
While you can rescan analog film at a higher resolution, video is digital and can't be rescanned."
Edit: why am I getting downvoted for asking a question? Sorry :-(
I assume your downvoting is because people (uncharitably) believed you were being an annoying pedant instead of asking a genuine question.
But you can infer visual detail from the information that is already there. Especially because ML uses information from the training set to help make sense of the information that is already there.
e.g. if I show you a picture of a key then you can figure out what the lock looks like because the key contains that information, even if it’s not visible.