3D Video Capture with Three Kinects
doc-ok.org
doc-ok.org
http://www.precisionmicrodrives.com/tech-blog/2012/08/28/usi...
As far as i know the new kinects use Time of Flight technology though (which is the reason they bought PrimeSense), which sends out short pulses of light, and times when they arrive back. Since there is no dot pattern to blur, the shaking technique won't work to my knowledge.
Nope. MS licensed first gen Kinect technology from PrimeSense. Apple is the company acquired it.
The ToF Kinect was developed fully in house. Here's the published paper on the depth sensor in ISSCC2013:
[A 512×424 CMOS 3D Time-of-Flight image sensor with multi-frequency photo-demodulation up to 130MHz and 2GS/s ADC]
http://ieeexplore.ieee.org/xpl/login.jsp?tp=&arnumber=675737...
Sadly we never got the funding to go further to do multi-cameras and I had to move on to other urgent things, so I'm glad to see others might get to solve it: imho there are many applications, even simple things like making better video conferencing using 3d capture viewed in the oculus :)
One suggestion however: the "fat points" pointcloud rendering of potree[2] might improve the appearance of the generated model instead of using meshes, could be worth a try.
-----------
[1] http://ivn.net/demo.html (you can skip the cheesy first minute of the video)
Toward the last quarter of the film, at one point he clips through the table, and it was shocking to see, but even after seeing it, when he moved back out, it still felt like a "person" more than a "CGI Ghost".
[0] https://en.wikipedia.org/wiki/Kinect#Open_source_drivers
There is enough 3D data in kinect stream, but 640x480 video is just pathetic.
Once you made something like this, then you'd start writing apps for it -- I would imagine you'd start off with virtual "pictures" for the walls that could have a web browser, spreadsheet, etc. built in. Then you could work up to truly interactive 3-D tools, but I'm not sure users could easily grasp moving to holographic toolsets right off the bat. It's an interesting marketing question.
Voxels are nice because they are well understood by most 3d developers, and have the same spatial resolution characteristics as we are used to on 2D formats.
Light fields on the other hand have easier capture going for them, don't change transmission formats (a light field can be transmitted in a 2d video or image), and don't suffer from interference problems.
I'm excited to see what happens.
Does someone know a DIY Lidar project or a cheap Lidar?
(Lidar are usually very expensive, e.g. the Lidar that Google uses for its autonomous cars cost 78k dollar.)
http://www.youtube.com/watch?v=B4g9J-aSF-c
Lots of interesting UX ideas in there.
There's no reason the color and depth data produced by the Kinects can't be encoded and compressed in a "regular" video stream.
A very naive approach (where you simply stitch the images side by side in a larger video stream) would require you to transmit 6 separate 640x480 video streams.
640 x 480 = 307,200 pixels
A 1080p video stream, that can now be easily broadcasted to any decent household internet connection, has 1920 x 1080 pixels.
1920 x 1080 = 2,073,600
2,073,600 / 307,200 = 6.75
So the 6 separate video streams would fit just right into it!
I'm not taking into account some factors, like:
* The effect of existing video compression algorithms on depth maps (might cause some severe artifacts, since they're tuned for color vision perception)
* The fact that the depth map has a single color channel, and can probably be represented more efficiently than a full color RGB 640x480 image.
* Framerate (60 fps is probably needed for a more immersive feel)
But I do think it's very feasible to stream this type of 3D video in real time with current Internet speeds.