Light-field videos: Part I
roblesnotes.com
roblesnotes.com
However I've been surprised at how there's hardly been any commercial applications of light-field technology. Lytro unfortunately folded despite having large amounts of capital, and the only company left is Raytrix, which uses light fields largely only for single-camera depth map generation. Google has released very nice tech demos, but that's about it.
*See here for a few test videos: https://www.youtube.com/watch?v=6Buj8WWhGrA&list=PLzhX-LcIzx...
Or: yes, this is a recent blog post.
Maybe a wifi enabled sd card is what's needed?
Edit: cool, there's even a hack that adds auto ftp uploading of new files to a specific brand of wifi enabled sd cards
https://bitbucket.org/harbortronics/flashair-ftp-upload/src/...
As painful as all of that was, there's no way I would have risked losing data by using wifi enabled cards. If you lose the image from one card, the entire take is lost. Trying to push that much data at the required throughput for 16-22 cameras is just not something that's going to work out in the field.
What were the hardware specs on the PC like? How many cameras could you record from simultaneously? Which OS were you using, and did you have to do anything special to make it work?
The main considerations are having enough USB bandwidth (hence the need for PCI-E expansion cards), dumping the USB camera data straight to the SSD (no re-encoding, use a proper Gstreamer pipeline or ffmpeg command), and ensuring you can sustain the write speeds (through RAID array). There was some bug in the USB drivers that caused the cameras to reserve more bandwidth than they actually used, limiting their numbers, but that was easy to fix.
My guess is that you probably re-encoded the stream, which would indeed drastically limit the number of cameras if their resolution was high enough.
I am looking at the footage from the A77 on Youtube. The resolution doesn't seem to be 4K. Apeman seems to do a lot of software upscaling. Their highest-end model, the A100, claims 20MP resolution, 4K video, and an Panasonic MN34120 sensor.
Panasonic claims 16MP resolution -- so 4MP less than Apeman claims, and a maximum of 22fps at 4k. This isn't a linear slide, and I think the best Apeman could do is grab 1080p at 30fps and upscale.
https://industrial.panasonic.com/content/data/SC/ds/ds4/MN34...
The A77 doesn't say what chip it uses, but it's a model down. And looking at two Youtube videos, I'd say it's grabbing at best 720p and upscaling, likely less, but most of the video has enough action that I could be wrong (compression relics do things as well when scenes change quickly).
OP: Can you post some video frames and see if this setup works as claimed? If it does, it'd be really neat to play with. I'd even be happy with true 1080.
I'd consider this setup more for photos. Lightfield video will be hard without frame-synchronized videos. But once you get photos, one can think about how to invest in videos next.
Resolution at 30fps is 3840x2160 pixels for the A77, but yes it's probably upscaled. Again at this point I am not too concerned about resolution at this point.
I recently got a spherical camera and I am trying to use it for photogrammetry. I also have an array of four 4k cameras with hardware synchronized shutters hooked up to an NVIDIA Jetson Xavier[1]. That system can record four 4k streams at once to the SSD. I wonder how many 4k streams a Jetson Nano could record, because then you could use 16 of these [2] and four to eight Jetson Nanos to make a camera system with all hardware synchronized shutters that could easily record the data and export it via the network. It would cost around $2500 though. These projects get expensive and I keep thinking I want a sponsor, but the slow pace using what hardware I can buy is probably fine for now.
I'm trying to do a complete photorealistic photogrammetry capture of hiking trails, so I can run my robot in simulation on virtual hiking trails and train real computer vision networks. Lately I've been wondering if there is a GAN in my future...
The frustrating thing about my project is the sheer amount of computation required. I really don't need a direct photogrammetry capture of a trail, an approximation would be fine to some degree. But I take like ten gigabytes of video data and then process each frame to find keypoints, run correlation on all these points, and all this (using COLMAP). This stuff can take days to process on my desktop.
Meanwhile there are neural networks that can compute depth from video in real time, and I wonder what it would take to stitch sequential depth estimations in to one 3D model with RGB textures in one continuous calculation. There's so much research to do!
By the way I found the work in this paper pretty fascinating. [3] Facebook is working on 6dof video recording and playback, which is quite the challenge on many levels!
[1] https://reboot.love/t/new-cameras-on-rover/ [2] https://www.e-consystems.com/4k-usb-camera.asp [3] https://research.fb.com/wp-content/uploads/2019/09/An-Integr...
https://www.e-consystems.com/nvidia-cameras/jetson-agx-xavie...
The encoding quality is good, but I have not tried lossless. I am encoding at 20,000 kb/s, which I am realizing is not super high. I may bump them all to 50,000 kb/s and see if it can still record. I am writing to an SSD over ESATA so I assume it can sustain four x 50mbps recording streams. I don't have much concern about the hardware encoder as those are pretty powerful.
I don't have an easy link to a raw image sample but I do have a stitched image from two cameras. This picture is extracted from Rover's two front video streams: https://reboot.love/uploads/default/original/1X/2e90f9e9e308...
If the hiking trails are accessible enough, you should have a look at SLAM technique. SLAM allows you to create smooth and rough approximate map of the environment through which you navigate your camera. Colorization of this map could be done by a GAN(might be an interesting side project).
I am adding some pointers below :
1. https://www.doc.ic.ac.uk/~ajd/ - Prof. Davison and his group's work is impressive in this area. 2. https://vision.in.tum.de/research/vslam - Prof. Cremers group have some SOTA algorithms in this area.
P.S: You don't need a heavy setup for this. A single or a stereo camera should do the job.
https://github.com/ivalab/gf_orb_slam2
I feel like vslam could be the the first step in a post-processing pipeline that would reduce a lot of the computational complexity of solving large maps. Once I can easily make large maps I can build simulated environments and use those for training an agent.
Thanks again for the tips!