HNHacker News
TopNewBestAskShowJobs

dr_blueberry

9 karma · joined May 12, 2026

submissionscomments
dr_blueberry··on Ask HN: What are you working on? (September 2026)
Working on real-world computer vision demos with a focus on fitness.

So far, they are based on ViTPose+ Large pose estimation in the context of different sports.

My demos include:

- Comparing dancers' sync performing the same choreography and computing a sync score

- Visualizing a runner's ankle path while running, as well as the average knee shape upon foot strike

- Counting chin-ups reps and measuring each rep duration

The code is open-source on GitHub:

https://github.com/jeremyipark/vision-demos

dr_blueberry··on I made a rock climbing tool using computer vision [video]
I prompted VLM Run’s visual agent Orion to segment all of the blue bouldering holds, and it did a good job! It is interesting that now we can prompt VLMs to segment all of the holds, rather than creating a new dataset from scratch to train a model.

With holds detection + pose estimation, I can show how each hold gets activated as a hand or foot uses it. Once we touch the final hold with both hands, the route is completed, and I show the overall path of my torso midpoint.

A tool like this could help climbers understand their movement better. I’m still very much a beginner at bouldering, so it would be great to get quantitative feedback.

There are definitely things to improve, but overall I’m encouraged by this first demo.

Models used:

- VLM Run’s Orion for segmentation (https://www.vlm.run/)

- ViTPose+ Huge for pose estimation (via Hugging Face)

- RT-DETR for person detection (via Hugging Face)

Shoutout to Daniel Reiff and his bouldering + computer vision project for the inspiration!

Link to Daniel Reiff's bouldering + computer vision blog: https://blog.roboflow.com/bouldering/

dr_blueberry··on Gemini Robotics 2 brings whole body intelligence to robots
Doing some initial testing of Gemini ER 2 within the Orion 2 visual agent harness. I'll share a few initial results and chat threads here. I'm impressed with how fast Gemini ER2 is.

1. Crop the segment when adding rice to the rice cooker + analyze a frame: https://chat.vlm.run/c/40b5edfb-6d15-47e8-bedd-c77b1ed92496

2. Extract 16 keyframes in a 4x4 grid + detect the water bottle: https://chat.vlm.run/c/a8542517-0021-4013-9663-a9f41f286e4c

3. Extract the frame at 0:03 + segment the lettuce + generate a 3D reconstruction: https://chat.vlm.run/c/494c7c65-4aa1-442f-a78a-bfd1409c784d

(I work at VLM Run).