HNHacker News
TopNewBestAskShowJobs

samuelm2

19 karma · joined August 6, 2024

submissionscomments
samuelm2··on OpenQuestCapture – An Open Source Meta Quest 3D Gaussian Splat Capture Pipeline
Hey all!

A few months ago, I launched vid2scene.com, a free platform for creating 3D Gaussian Splat scenes from phone videos. Since then, it's grown to thousands of scenes being generated by thousands of people. I've absolutely loved getting to talk to so many users and learn about the incredible diversity of use cases: from earthquake damage documentation, to people selling commercial equipment, to creating entire 3D worlds from text prompts using AI-generated video (a project using the vid2scene API to do this won a major Supercell games hackathon just recently!)

When Meta's Horizon Hyperscape come out, I was impressed by the quality. But I didn't like the fact that users don't control their data. It all stays locked in Meta's ecosystem. So I built a UX for scanning called OpenQuestCapture. It is an open source, MIT licensed Quest 3 reconstruction app.

Here's the GitHub repo: https://github.com/samuelm2/OpenQuestCapture

It captures Quest 3 images, depth maps, and pose data from the Quest 3 headset to generate a point cloud. While you're capturing, it shows you a live 3D point cloud visualization so you can see which areas (and from which angles) you've covered. In the repo submodules is a Python script that converts the raw Quest sensor data into COLMAP format for processing via Gaussian Splatting (or whatever pipeline you prefer). You can also zip the raw Quest data and upload it directly to vid2scene.com/upload/quest/ to generate a 3D Gaussian Splat scene if you don't want to run the processing yourself.

It's still pretty new and barebones, and the raw capture files are quite large. The quality isn't quite as good as HyperScape yet, but I'm hoping this might push them to be more open with Hyperscape data. At minimum, it's something the community can build on and improve.

There's still a lot to improve upon for the app. Here are some of the things that are top of mind for me:

- An intermediary step of the reconstruction post-process is a high quality, Matterport-like triangulated colored 3D mesh. That itself could be very valuable as an artifact for users. So maybe there could be more pipeline development around extracting and exporting that.

- Also, the visualization UX could be improved. I haven't found a UX that does an amazing job at showing you exactly what (and from what angles) you've captured. So if anyone has any ideas or wants to contribute, please feel free to submit a PR!

- The raw quest sensor data files are massive right now. So, I'm considering doing some more advanced Quest-side compression of the raw data. I'm probably going to add QOI compression to the raw RGB data at capture time, which should be able to losslessly compress the raw data by 50% or so.

If anyone wants to take on one of these (or any other cool idea!), would love to collaborate. And, if you decide to try it out, let me know if you have any questions or run into issues. Or file a Github issue. Always happy to hear feedback!

Tl;dr, try out OpenQuestCapture at the github link above

samuelm2··on Gaussian Splatting Alternative: WebGL Implementation of Nvidia's SVRaster
I was also surprised that my iPhone didn't do too well with it. Regular Gaussian Splatting runs pretty well on my iPhone, almost as good as my laptop (of course, my phone screen is a lot smaller so its rendering a lot fewer pixels). The fragment shader for this is very heavy compared to gaussian splatting, so maybe it has to do with that.

And I agree re: constantly drawing frames vs only drawing on change, that's a good optimization to make.

samuelm2··on Gaussian Splatting Alternative: WebGL Implementation of Nvidia's SVRaster
thanks! You can go to this link for a 1.2m voxel version of the pumpkin scene:

https://vid2scene.com/voxel/?url=https://huggingface.co/samu...

I'd try to go higher but I start running out of VRAM when training a scene with too many voxels. Maybe I'll spin up a runpod instance to train higher-voxel scene on a beefier GPU.

samuelm2··on Gaussian Splatting Alternative: WebGL Implementation of Nvidia's SVRaster
Right now, its pretty restricted to relatively static geometry. You can do simple transforms like scaling, rotating, etc to groups of voxels, so maybe down the line you'd be able to animated them!
samuelm2··on Gaussian Splatting Alternative: WebGL Implementation of Nvidia's SVRaster
Hi all! Several weeks ago, Nvidia released a voxel-based radiance field rendering technique called SVRaster. I thought it was an interesting alternative to Gaussian Splatting, so I wanted to experiment with it and learn more about it.

I've been working on a WebGL viewer to render the SVRaster Voxel scenes from the web, since the paper only comes with a CUDA-based renderer. I decided to publish the code under the MIT license. Here's the repository: https://github.com/samuelm2/svraster-webgl/

I think SVRaster Voxel rendering has an interesting set of benefits and drawbacks compared to Gaussian Splatting, and I think it is worth more people exploring.

I'm also hosting it on vid2scene.com/voxel so you can try it out without having to clone the repository. (Note: the voxel PLY file it downloads is about 50MB so you'll probably have to be on good WiFi).

Right now, there's still a lot more optimizations that would make it faster. I only made the lowest-hanging fruit optimizations. I get about 60FPS on my Laptop 3080 GPU at 2k resolution, and about 10-15 FPS on my iPhone 13 Pro Max.

On the github readme, there's more details about how to create your own voxel scenes that are compatible with this viewer. Since the original SVRaster code doesn't export ply, theres an extra step to convert those voxel scenes to the ply format that's readable by the WebGL viewer.

If there's enough interest, I'm also considering doing a BabylonJS version of this

Also, this project was made with heavy use of AI assistance ("vibe coded"). I wanted to see how it would go for something graphics related. My brief thoughts: it is super good for the boilerplate (defining/binding buffers, uniforms, etc). I was able to get simple voxel rendering within minutes / hours. But when it comes to solving the harder graphics bugs, the benefits are a lot lower. There were multiple times where it would go in the complete wrong direction and I would have to rewrite portions manually. But overall, I think it is definitely a net positive for smaller projects like this one. In a more complex graphics engine / production environment, the benefits might be less clear for now. I'm interested in what others think.