Maybe I'm ignorant but what's a 1D image ?
Maybe I'm ignorant but what's a 1D image ?
There is no such thing ("2D image" is still useful, to distinguish from 3D.)
No it isn't? You don't say "here's that 2D image you wanted" when you send someone a jpg.
(+ the meaning in math as what comes out of a function....)
and time-based volumetric recording of 3D video-games are 4D?
(Assuming you do work with ML+videos) - it's surprising to hear you say you work with RGB instead of YUV - can you briefly explain how that's the case? I'd have thought that using luma/chroma separation would be much easier to work with (not just with traditional video tooling, but ML/NNs/etc themselves would have an easier time consuming it.
I take the impression that much of the time, color doesn't provide much signal and gives your model things to overfit on, so you collapse it down to grayscale. (Which is to say, most of the time you care about shape, but you don't care about color.) But I bet there are problem spaces where your intuition holds, I'm sure that there's performance the be wrung out of a model by experimenting with different color spaces who's geometry might separate samples nicely.
I did something similarish a few months ago where I used LDA[1] to create a boutique grayscale model where the intensity was correlated to the classification problem at hand, rather than the luminosity of the subject. It worked better than I'd have guessed, just on it's own (though I suspect it wouldn't work very well for most problems). But the idea was to preprocess the frames of the video this way and then feed it into a CNN [2]. (Why not a transformer? Because I was still wrapping my mind around simpler architectures.)
[1] https://en.wikipedia.org/wiki/Linear_discriminant_analysis
[2] https://en.wikipedia.org/wiki/Convolutional_neural_network
Video files are a linked list of 2D images.
> and time-based volumetric recording of 3D video-games are 4D?
Kind of, but it would be an oversimplification. Typically when we refer to the dimensionality of objects, we're referring to physical dimensions. Time is a temporal dimension. I think it would be more specific to say this is a linked list of 3-dimensional images, right?
Still, it was a fun project.
I wonder if your idea would work for lightfield captures, or time sequences of a lightfield.
But the more you understand the scene, the more you can potentially outright reconstruct, and in some contexts more loss would be entirely fine if the artifacts are plausible.
Note that I did this in '98 or so, when there was less of a computational budget, maybe what I couldn't hack back then is feasible today.
I like the sibling's suggestion about audio; if we were to adopt it, it would make a 1D element a "sample".
[1] People are often confused on this point, because in the course of everyday conversation we don't distinguish between the number of pixels on the side of a rectangle (which is a 1D quantity) and the number of pixels inside that rectangle (a 2D quantity). So if I say I have a 10 pixel by 10 pixel image, what I mean is that I have a grid with an area of 100 pixels, with sides measuring 10 pixel-widths by 10 pixel-heights (each a 1D quantity of length). If that looks awkward and tiresomely pedantic to you, well, that's why we just say pixels and let the details be implied.
If you're still skeptical, consider for instance that voxels are more clearly a unit of volume (think Minecraft blocks), and that pixels are obtained by subdividing a rectangle. Another useful way to think about it might be by replacing "pixels" with "dominos" and imagining making grids out of dominos, pixels can be tricky since you can't see their area yourself.
But in all seriousness, call it what you want, I happen to enjoy this minutia but understand many people see it as an impediment to clear communication. If you're working in the unusual contexts where the difference matters you probably know.
But I would still argue that an array of pixels doesn't represent a 1D image. If a 2D image associates areas with color or intensity values, a 1D image would associate intervals with color or intensity values (since intervals are the measure of 1D space which is analogous to areas in 2D space). In my mind, those are different data structures, but the difference is pretty nuanced and I would understand if people felt I was splitting hairs.
Given the question, "what's a 1D image?" I'd argue this is the more complete answer. But if we were to ask that question in the context of a real world problem, yours is likely to be the more useful answer.
Some computer graphics experts don't agree with this: http://alvyray.com/Memos/CG/Microsoft/6_pixel.pdf