689 karma · joined December 31, 2015
E-mail: alex %{dot} floresescar %{@gmail}
Curious about Gaussian Splatting.
Most recently started work on splatinit.c It initializes 3D Gaussian Splats from an image and a depth map. It's fast and outputs to a typical 3D Gaussian Splat .ply file.
https://github.com/afloresescarcega/splatinit
I use the instant edit feature with directions on exact code changes.
Instant edits feature can surgically perform text edits fast without all the extra fluff or unsolicited enhancements.
I copied shadertoys, asked it to rename all variables to be more descriptive and pasted the result to see it still working. I'm impressed.
I copied a shader toy example, asked it to rename all the variables to be more descriptive and it edited just the variable names. I was able to compile and run in shader toy.
It has some challenges that it needs to solve to do great as a cross platform "general-purpose" programming language.
It's paradoxically high level with its syntax and ergonomics but is tied down to the same cross platform headaches like in low level languages (e.g. cpp). Linking across cross platforms requires lots of careful thought and testing. Unlike cpp, it's not super portable. It requires a hefty 30 MB runtime for some features of the language to work. Try static executable hello world.
That being said, it's possible. You can build cross platform applications with Swift, but you'd still have some of the same kinds of portability issues like in cpp but with nicer syntax and ergonomics.
If that's the case, I would rather give my money to a human driver than just Uber.
I believe that we should explore pretraining video completion models that explicitly have no text pairings. Why? We can train unsupervised like they did for GPT series on the text-internet but instead on YouTube lol. Labeling or augmenting the frames limits scaling the training data.
Imagine using the initial frames or audio to prompt the video completion model. For example, use the initial frames to write out a problem on a white board then watch in output generate the next frames the solution being worked out.
I fear text pairings with CLIP or OCR constrain a model too much and confuse
I've started a project to initialize each pixel in an image into a mesh of GS. I used a depth map to unproject them into space.
I've been very curious what a few training iterations would do to optimize my scenes. The original 3DGS implementation is not accessible on my hardware at the moment! I really look forward to your training implementation!
Wondering if I could then play PCVR with game streaming on the Apple Vision Pro
https://developer.apple.com/documentation/avfoundation/addit...
Currently a huge challenge is in real-time reconstruction. Approaches involve estimating point clouds from images then optimizing splats on those. Other approaches are using SLAM like LiDAR to have distance of points to the camera but then still optimizing on that.
Optimization is producing good results but takes iterations that are not suitable for real-time.
Pixel-wise splat estimation with iPhone LiDAR could produce good results but need help and expertise
I'm more excited about Samsung's micro OLED with no color filters that will push the brightness and/or reduce power consumption.
AFAIK, Apple's Sony micro OLED does have color filters. So there is definitely an opportunity for another display upgrade.
Thinking in 2D for a second, to get a nice crispy edge, you need a long and opaque splat to mark the boundary. Sometimes the long splat could wisp off leaving fuzzy artifacts.
Take this example: https://www.shadertoy.com/view/dtSfDD
Peyman Milanfar [1] suggested using bump functions instead. Bump functions would allow you to specify cut off intervals but still make the whole function smooth and continuous (good for my gradient optimization freaks)
Question for the authors, are there opportunities, where they exist, to not use optimization or tuning methods for reconstructing a model of a scene?
We are refining efficient ways of rendering a view of a scene from these models but the scenes remain static. The scenes also take a while to reconstruct too.
Can we still achieve the great look and details of RF and GS without paying for an expensive reconstruction per instance of the scene?
Are there ways of greedily reconstructing a scene with traditional CG methods into these new representations now that they are fast to render?
Please forgive any misconceptions that I may have in advanced! We really appreciate the work y'all are advancing!