HNHacker News
TopNewBestAskShowJobs

rocauc

1,027 karma · joined January 7, 2017

working on https://roboflow.com, developer tools for computer vision. rocauc@roboflow

prev built NLP products, farms

submissionscomments
rocauc··on Roboflow Playground: Try and Compare 30 Computer Vision Models
thank you for the model suggestions!
rocauc··on Meta Segment Anything Model 3
As someone that works on a platform users have used for labeling 1B images, I'm bullish SAM 3 can automate at least 90% of the work. Data prep is flipped to models being human-assisted instead of humans being model-assisted (see "autolabel" https://blog.roboflow.com/sam3/). I'm optimistic majority of users can now start deploying a model to then curate data instead of the inverse.
rocauc··on Meta Segment Anything Model 3
A brief history. SAM 1 - Visual prompt to create pixel-perfect masks in an image. No video. No class names. No open vocabulary. SAM 2 - Visual prompting for tracking on images and video. No open vocab. SAM 3 - Open vocab concept segmentation on images and video.

Roboflow has been long on zero / few shot concept segmentation. We've opened up a research preview exploring a SAM 3 native direction for creating your own model: https://rapid.roboflow.com/

rocauc··on Meta Segment Anything Model 3
The model supports batch inference, so all prompts are sent to the model, and we parse the results.
rocauc··on Meta Segment Anything Model 3
I tried it on transparent glass mugs, and it does pretty well. At least better than other available models: https://i.imgur.com/OBfx9JY.png

Curious if you find interesting results - https://playground.roboflow.com

rocauc··on Meta Segment Anything Model 3
Yes. But also note that redistribution of SAM 3 requires using the same SAM 3 license downstream. So libraries that attempt to, e.g., relicense the model as AGPL are non-compliant.
rocauc··on Meta Segment Anything Model 3
Yes. It's a custom license with an Acceptable Use Policy preventing military use and export restrictions. The custom license permits commercial use.
rocauc··on A down detector for down detector's down detector
yes, downdetectorsdowndetectorsdowndetectorsdowndetector is available.
rocauc··on Rivian's TM-B electric bike
The bike lane compliant vehicle category is exciting. Infinite Machine (infinitemachine.com) made me aware of this category with their Olto model, which is at a (surprisingly) superior price point.
rocauc··on Every vibe-coded website is the same page with different words. So I made that
Not nearly enough gradient for a vibe coded site :)
rocauc··on Edge AI for Beginners
One of the most common uses for edge AI not listed in this course is computer vision. You similarly want real-time inference for processing video. Another open source project that makes it easy to use SOTA vision models on the edge is inference: https://github.com/roboflow/inference
rocauc··on Search all text in New York City
Reminds me of NY Cerebro, semantic search across New York City's hundreds of public street cameras: https://nycerebro.vercel.app/ (e.g. search for "scaffolding")
rocauc··on Show HN: Ten years of running every day, visualized
both the endeavor and the site are super cool - congrats on 10 years. interaction on the graphics would be a nice touch to select into a specific run. went looking for the code on your GH! https://github.com/friggeri
rocauc··on Tesla drives into Wile E. Coyote fake road wall in camera vs. Lidar test
In the 2019 fatal Tesla Autopilot crash, the Tesla failed to identify a white tractor trailer crossing the highway: https://www.washingtonpost.com/technology/interactive/2023/t...
rocauc··on Tesla drives into Wile E. Coyote fake road wall in camera vs. Lidar test
I wonder how long until techniques like Depth Anything (https://depth-anything-v2.github.io/) provide parity with human depth perception. In Mark Rober's tests, I'm not sure even a human would have passed the fog scenario, however.
rocauc··on AI Demos
Meta deeply comprehends the impact of GPT-3 vs ChatGPT. The model is a starting point, and the UX of what you do with the model showcases intelligence. This is especially pronounced in visual models. Telling me SAM2 can "see anything" is neat. Clicking the soccer ball and watching the model track it seamlessly across the video even when occluded is incredible.
rocauc··on Video Surveillance with YOLO+llava
A suggestion: I'd swap llava for Florence-2 for your open set text description. Florence-2 seems uniformly more descriptive in its outputs.
rocauc··on Segment Anything Model and Friends
SAM 2's key contribution is adding time-based segmentation to apply to videos. Even on images alone, the authors note [0] the image-based segmentation benchmark does exceed SAM 1 performance. There have been some weaknesses exposed in areas of SAM 2 vs SAM 1, like potentially medical images [1]. Efficient SAM trades SAM 1 accuracy for ~40x speedup. I suspect we will soon see Efficient SAM 2.

[0] https://x.com/josephofiowa/status/1818087122517311864 [1] https://x.com/bowang87/status/1821021898928443520?s=46&t=9K-...

rocauc··on SAM 2: Segment Anything in Images and Videos
One thing its enabled is automated annotations for segmentation, even on out-of-distribution examples. e.g. in the first 7 months of SAM, users on Roboflow used SAM-powered labeling to label over 13 million images, saving over ~21 years[0] of labeling time. That doesn't include labeling from self hosting autodistill[1] for automated annotation either.

[0] based on comparing avg labeling session time on individual polygon creation vs SAM-powered polygon examples [1] https://github.com/autodistill/autodistill

rocauc··on Show HN: A football/soccer pass visualizer made with Three.js
"Load Example" was very helpful to get a sense of what this does. Awesome build. +1 to the other comment wanting a breakdown of what the colors mean.

Also, combining this with real-time in-game camera play could be really powerful, too[0]. Like illuminating details during the game.

[0] https://x.com/skalskip92/status/1816461263829889238

rocauc··on I am using AI to drop hats outside my window onto New Yorkers
i work on roboflow. seeing all the creative ways people use computer vision is motivating for us. let me know (email in bio) if there's things you'd like to be better.
rocauc··on Vision Transformers Are Overrated
Pulling out a key part of this post from a DeepMind 2023 paper[1]: “Although the success of ViTs in computer vision is extremely impressive, in our view there is no strong evidence to suggest that pre-trained ViTs outperform pre-trained ConvNets when evaluated fairly.”

Another common constraint in vision vs language is the long tails are very long in the visual world. There's a number of domains where you have very little examples to learn (defects are designed to happen infrequently; rare species for identification show up, well, rarely). And pulling from the blog: "But small models ... benefit greatly from the exact type experiment of outlined in this post: strong augmentation with limited data trained across many epochs."

[1] https://arxiv.org/pdf/2310.16764.pdf

rocauc··on Suno AI
suno has improved fast. I remember when they released Bark in April ‘23. it was good. but this new model is fun. props to the team.
rocauc··on Show HN: I scraped 25M Shopify products to build a search engine
Really neat. I tried your search for red shoes, and I found some, er, unexpected imagery on page 1.

One thing you could do is add semantic search so when a user searches "red shoes," the index returns images that look like red shoes even if the metadata doesn't say anything about color or item types. To do this, I'd use a model like CLIP. Here's an example of using CLIP and Supabase to do semantic image search: https://blog.roboflow.com/how-to-use-semantic-search-supabas...

rocauc··on Advent of Technical Writing: Navigation Structure (Day 1 of 24)
In your inference project example, what examples do you place in Getting Started vs common usage examples? In general, where is the best place for usage examples - alongside the methods they use, or in an independent section?
rocauc··on Crowd Funding the Release of OpenCV 5
I can't recall the true first time I used OpenCV because it's simply so embedded in all image processing work. I've backed the project, and my company has contributed as well.
rocauc··on Brev: Start fine-tuning and training models in < 10 minutes
brev's execution to make LLM finetuning approachable, educational, and fun has been nothing shy of inspiring
rocauc··on The CEO of Airbnb wants you to charge less for your house
The place I most like Airbnb is when traveling in a group. There’s a different dynamic to having my family or friends together in the ‘unscheduled’ times when traveling.
rocauc··on Ask HN: Who is hiring? (October 2023)
Roboflow | Multiple Roles | Full-time (SF, NYC, Distributed) | https://roboflow.com/careers?ref=whoishiring1023

Roboflow is the fastest way to use computer vision in production. We help developers give their software the sense of sight. Our hosted tools are used for image/video collection and annotation; dataset exploration with foundation models; model training; deployment to anywhere. Here's a simple webapp I made recently: https://josephofiowa.com/rps-embedded-vision23/

Over 250k engineers (including engineers from 1/2 Fortune 100 companies) build with Roboflow. We now host the largest collection of open source computer vision datasets and pre-trained models[2]. We are pushing forward the CV ecosystem with open source projects like Autodistill[3] and Supervision[4]. And we've built one of the most comprehensive resources for software engineers to learn to use computer vision with our popular blog[5] and YouTube channel[6].

We have several openings available, but are primarily looking for strong technical generalists. Our engineering culture is built on a foundation of autonomy & we don't consider an engineer fully ramped until they can "choose their own loss function." At Roboflow, engineers aren't just responsible for building things but also for helping figure out what we should build next. We're builders & problem solvers; not just coders. (For this reason we also especially love hiring founders.)

We're currently hiring full-stack engineers for our ML and web platform teams, a web developer to bridge our product and marketing teams, several technical roles on the sales & field engineering teams, and our first applied machine learning researcher to help push forward the state of the art in computer vision.

[1]: https://roboflow.com/?ref=whoishiring1023 [2]: https://roboflow.com/universe?ref=whoishiring1023 [3]: https://github.com/autodistill/autodistill [4]: https://github.com/roboflow/supervision [5]: https://blog.roboflow.com/?ref=whoishiring1023 [6]: https://www.youtube.com/@Roboflow

rocauc··on First Impressions with GPT-4V(ision)
+1
Page 1 of 3Next →