59 karma · joined January 31, 2022
When we started building Pareto, we weren’t thinking about YouTube, but the data path looked familiar: upload, process, store, index, play.
Here's what the classic system design question gets right in production.
- Read: https://hebbianrobotics.com/blog/youtube-system-design-for-r...
- Try Pareto: https://pareto.hebbianrobotics.com
- Code open source: https://github.com/Hebbian-Robotics/pareto
Features: - Upload and view .ply file directly in the browser - Multiple camera trajectory animations (rotate, swipe, shake, forward) - Interactive orbit controls (drag to orbit, scroll to zoom, right-drag to pan) - No installation required – runs entirely client-side
The original ml-sharp can render the same video trajectories but requires a CUDA GPU. This viewer lets you explore the 3D output on any device with a browser.
I also added cloud GPU inference via Modal so you can generate splats without a local GPU (free tier available) [2].
Started off as an open source alternative to Wispr Flow for myself as I wanted to have more control over the formatting rules as well as model choice but after sharing with friends and presenting it at my local Claude Code meetup, I was encouraged to share it more widely.
The desktop app uses tauri so it is cross-platform compatible and I have tested it working on macOS and windows.
About ollama in pipecat: https://docs.pipecat.ai/server/services/llm/ollama
Also, check out any provider they support, and it can be easily onboarded in a few lines of code.
The integration to set up the WebRTC connection, get the voice dictation working seamlessly from anywhere, and input into any app took a long time to build out, and that's why I want to share this open source.
Unlike your average LLM benchmark, this benchmark focuses on location and time as variables since these are the biggest factors for networking systems (I was a developer for networking tools in a past life). The idea is to run benchmarks from multiple geographic locations over time to see how each platform performs under different conditions.
Basic setup: echo agent servers can create and connect to temporary rooms to echo back after receiving messages. Since Pipecat (Daily) and LiveKit Python SDKs can't coexist in the same process, I have to run separate agent processes on different ports. Benchmark runner clients send pings over WebRTC data channels and measure RTT for each message. Raw measurements get stored in InfluxDB, then the dashboard calculates aggregate stats (P50/P95/P99, jitter, packet loss) and visualizes everything with filters and side-by-side comparisons.
I struggled with creating a fair comparison since each platform has different APIs. Ended up using data channels (not audio) for consistency, though this only measures data message transport, not the full audio pipeline (codecs, jitter buffers, etc).
Latency is hard to measure precisely, so I'm estimating based on server processing time - admittedly not perfect. Only testing data channels, not full audio path. And it's just Pipecat (Daily) and LiveKit for now, would like to add Agora, etc.
The README screenshot shows synthetic data resembling early results. Not posting raw results yet since I'm still working out some measurement inaccuracies and need more data points across locations over time to draw solid conclusions.
This is functional but rough around the edges. Happy to keep building it out if people find it useful. Any ideas on better methodology for fair comparisons or improving measurements? What platforms would you want to see added?
Stack: Python, TypeScript (React), InfluxDB
Made this with Devvit at the LA Tech Week Reddit Hackathon last week. Currently, the questions are just based on metadata like upvotes, comments, and authors, but if you like this, I would be happy to work on more features like leaderboards and AI-generated answers.
You can also contribute to it directly here: https://github.com/kstonekuan/reddit-trivia-night
You can also join Chrome’s EPP to use it on webpages