HNHacker News
TopNewBestAskShowJobs

rocauc

1,027 karma · joined January 7, 2017

working on https://roboflow.com, developer tools for computer vision. rocauc@roboflow

prev built NLP products, farms

submissionscomments
rocauc··on Show HN: ASL Classifier Built with CoreML and Roboflow
Did you find the model to perform better / worse when the background varied? I see it's all wood table examples in the gifs.
rocauc··on Show HN: Send Audio Messages in Gmail
Nice - how did you manage to store audio locally yet enable both parties to have access?
rocauc··on PP-YOLO Surpasses YOLOv4 – State-of-the-art object detection techniques
Thank you very much for the kind words.
rocauc··on PP-YOLO Surpasses YOLOv4 – State-of-the-art object detection techniques
Thanks for the copyedits - I've updated to "Joseph Redmon."
rocauc··on PP-YOLO Surpasses YOLOv4 – State-of-the-art object detection techniques
I am guilty of using your dumb idea with satisfactory performance.

I'd recommend EfficientNet (or one of its may variations) for an off-the-shelf SOTA classifier. https://paperswithcode.com/sota/image-classification-on-imag...

rocauc··on Benchmarking the Major Cloud Vision AutoML Tools
Really appreciate it; we had fun writing it. More of this to come.
rocauc··on Benchmarking the Major Cloud Vision AutoML Tools
The 'hard costs' were the training time for each model ($125) and conducting inference on the valid set ($15) for each.

Now, we had a few failed experiments like trying to run the COCO dataset through each tool. And none of this includes the implicit cost of the computer vision engineering time.

rocauc··on Yoloface-500k: ultra-light real-time face detection model, 500kb
Thanks, I see the table. Are the source datasets available for creating additional benchmarks?
rocauc··on Yoloface-500k: ultra-light real-time face detection model, 500kb
Purpose-built, small, and fast models appears to be the inevitable evolution for computer vision.

Where can the "Easy Set, "Medium Set, and "Hard Set" evaluations referenced in the "Wider Face Val" be found?

rocauc··on Responding to the Controversy about YOLOv5
Thanks for your followup thoughts. Pleased the conversation has evolved to focus on architecture and performance rather than naming alone.

re: GitHub comment deletion - We determined we should engage when we have easier to reproduce results, so we moved quickly to share them. That's what this post and diligence is, and I've re-engaged on the issues thread in question [1] with fuller remarks + reproducible Colab notebooks.

[1] https://github.com/AlexeyAB/darknet/issues/5920#issuecomment...

rocauc··on YOLOv5: State-of-the-art object detection at 140 FPS
Appreciate it!

Nice, I really respect research coming out of NIH. (Happen to know Travis Hoppe?) Coincidentally, our notebook demo for YOLOv5 is on the blood cell count and detection dataset: https://public.roboflow.ai/object-detection/bccd

We've seen 1000+ different use cases. Some of the most popular are in agriculture (weeds vs crops), industrials / production (quality assurance), and OCR.

Send me an email? joseph at roboflow.ai

rocauc··on YOLOv5: State-of-the-art object detection at 140 FPS
Agreed!

Crucially, we're tracking "out of the box" performance, e.g., if a developer grabbed X model and used it on a sample task, how could they expect it to perform? Further research and evaluation is recommended!

For size, we measured the sizes of our saved weights files for Darknet YOLOv4 versus the PyTorch YOLOv5 implementation.

For inference speed, we checked "out of the box" speed using a Colab Notebook equipped with a Tesla P100. We used the same task[1] for both - e.g. see the YOLOv5 Colab notebook[2]. For Darknet YOLOv4 inference speed, we translated the Darknet weights using the Ultralytics YOLOv3 repo (as we've seen many do for deployments)[3]. (To achieve top YOLOv4 inference speed, one should reconfigure Darknet carefully with OpenCV, CUDA, cuDNN, and carefully monitor batch size.)

For accuracy, we evaluated the task above with mAP after quick training (100 epochs) with the smallest YOLOv5s model against the full YOLOv4 model (using recommended 2000*n, n is classes). Our example is a small custom dataset, and should be investigated on e.g. COCO. 90-classes.

[1] https://public.roboflow.ai/object-detection/bccd [2] https://colab.research.google.com/drive/1gDZ2xcTOgR39tGGs-EZ... [3] https://github.com/ultralytics/yolov3

rocauc··on YOLOv5: State-of-the-art object detection at 140 FPS
Hey all - OP here. We're not affiliated with Ultralytics or the other researchers. We're a startup that enables developers to use computer vision without being machine learning experts, and we support a wide array of open source model architectures for teams to try on their data: https://models.roboflow.ai

Beyond that, we're just fans. We're amazed by how quickly the field is moving and we did some benchmarks that we thought other people might find as exciting as we did. I don't want to take a side in the naming controversy. Our core focus is helping developers get data into any model, regardless of its name!

rocauc··on YOLOv5: State-of-the-art object detection at 140 FPS
Great points, and hoping Glenn releases a paper to complement performance. We are also planning more rigorous benchmarking nonetheless.

re: PyTorch being a confounding factor for speed - we recompiled YOLOv4 to PyTorch to achieve 50 FPS. Darknet would likely top out around 10 FPS on the same hardware.

EDIT: Alexey, author of YOLOv4, provided benchmarks of YOLOv4 hitting much higher FPS here: https://github.com/AlexeyAB/darknet/issues/5920#issuecomment...

rocauc··on YOLOv5: State-of-the-art object detection at 140 FPS
I remember following this as it came out (and learning windshield wipers should be called "swipey bois")

Surprised and happy to hear you're seeing high labeling quality.

We'll re-host with credit on https://public.roboflow.ai What license is this?

rocauc··on YOLOv5: State-of-the-art object detection at 140 FPS
EfficientDet was open sourced March 18 [1], YOLOv4 came out April 23 [2], and now YOLOv5 is out only 48 days later.

In our initial look, YOLOv5 is 180% faster, 88% smaller, similarly accurate, and easier to use (native to PyTorch rather thank Darknet) than YOLOv4.

[1] https://venturebeat.com/2020/03/18/google-ai-open-sources-ef... [2] https://arxiv.org/abs/2004.10934

rocauc··on Training YOLOv4 on a Custom Dataset
YOLOv4 was published[1] on April 23 with COCO weights, but there haven't yet been resources on how to adapt its architecture to your own domain. This post walks through setting up Darknet and training on a custom dataset in Colab.

FWIW, in our tests, we saw the highest mAP (89.5) on a new task[2] compared to EfficientDet and YOLOv3.

[1] https://arxiv.org/abs/2004.10934 [2] https://public.roboflow.ai/object-detection/bccd

rocauc··on SQL for the Rest of Us
this is a great breakdown on the WHY of SQL.

giving clarity as to the what / why of SQL empowers those that don't yet know it how to make better asks for data in their organization.

rocauc··on Earnest Capital Trailhead
+1 to Pioneer's model
rocauc··on Should I apply to Pioneer app?
Yes, I'd recommend it. Pioneer gamifies your To-Do list, keeping you accountable and increasing your output. There's little downside to trying it out.

I think there's been a fair number of Pioneer projects that go onto YC, too

rocauc··on Inside of a Tractor Cab [video]
The first self-driving tractors were used in research environments in 1996(!)

[1] Automatic steering of farm vehicles using GPS. https://onlinelibrary.wiley.com/doi/abs/10.2134/1996.precisi...

[2] First results in vision-based crop line tracking https://ieeexplore.ieee.org/document/503895

rocauc··on Illustrated FixMatch for semi-supervised learning
As others have noted, this is data augmentation, and it's incredibly useful to increase variation in training data to help decrease overfitting.

It's not a silver bullet. It won't capture the natural variations that happen in the real world.

But new forms of augmentation (like OP) are helping us get closer.

For example, MixMatch creates "mosaic" images by combining images across the training set [1]. In object detection, bounding box only augmentations are improving models by introducing variation [2].

And an anecdote: I work on https://roboflow.ai , and we've seen customers make production-ready results from datasets <20 images based on techniques like these.

[1] https://arxiv.org/abs/1905.02249 [2] https://arxiv.org/pdf/1906.11172.pdf

rocauc··on A popular self-driving car dataset is missing labels for hundreds of pedestrians
The scary part isn't necessarily this dataset, but that unlabeled data causes a silent decrease in model performance -- which can be esp important for underrepresented classes.
rocauc··on Tutorial: Using the TensorFlow Object Detection API on a Custom Dataset
Colab Notebook: https://bit.ly/rf-mn

Associated Dataset: https://public.roboflow.ai/object-detection/bccd

rocauc··on Show HN: ChessBoss – enhancing physical chessboards with computer vision
Nice - I like how your project demonstrates a Bird’s Eye View angle works well.

We’ll aim to support the view from a seated player on each side of the game.

rocauc··on Show HN: ChessBoss – enhancing physical chessboards with computer vision
So, as we think about making this into a production app, a feature you’d request is the ability for a user to add input as to why a given move is being suggested?

At present, we’re simply outputting the recommended move from StockFish, which also does give a “Why.” Perhaps letting users add that commentary in the app sufficiently solves that need.

rocauc··on Show HN: ChessBoss – enhancing physical chessboards with computer vision
Nice. We’ll dive into your repo, too.

How did you handle a queen occluding a smaller pawn behind it? Simply more training data?

rocauc··on Show HN: ChessBoss – enhancing physical chessboards with computer vision
We think streaming and creating game logs will be great features to add.

We’ll likely share to the 147k /r/chess when we have an app others can demo. Good call.

rocauc··on Show HN: ChessBoss – enhancing physical chessboards with computer vision
Great minds! Love that you deployed to a Pi – I’ve thought about the same to complement or replace smartphones.

Can you shed some insight into your ML process? One thing we did to simplify the vision problem is capture images from the same perspective (hence our tripod). We labeled 2894 objects across 292 images. We had 12 objects to detect: each piece for black and white. We struggled with occlusion, especially if a pawn is behind a queen.

← PreviousPage 3 of 3