YOLOv5 on FPGA with Hailo-8 and 4 Pi Cameras
fpgadeveloper.com
fpgadeveloper.com
What are real life actual useful cases for this tech?
I can imagine in manufacturing: detecting defects or layout mismatch - that's one.
Is there any open source project that uses a image recognition library to achieve any useful task? All I've seen from board partners seem to at most provide very simple demos, where a box with label is drawn around an object. Who actually is using that information, how and for what?
I've also been a part of the Kinect craze and made 3 demos (games mostly) using their SDK and still have a very hard time defending this tech in eyes of coworkers that only see this as a surveillance tech
Drawing bounding boxes is a common end point for demos, but for businesses using computer vision there is an entire world after that: on device deployment. This can be on devices like an NVIDIA Jetson (a very common choice), to Raspberry Pis to central CUDA GPU servers for processing large volumes of data (maybe connected to cameras over RTSP).
Note: There are many models that are faster and perform better than YOLOv5 (i.e. YOLOv8, YOLOv10, PaliGemma). Roboflow Inference that our ML team maintains has various guides on deploying models to the edge: https://inference.roboflow.com/#inference-pipeline
The system works quite simply: I start with an existing object detector and train it with a small (<100) number of manually labelled images. Then during inference, I move the scope's field of view using motor commands to put the center of the tardigrade at the center of the field of view.
This technology is very useful for doing long-term observations of tardigrades (so, useful for science).
That makes me want to revisit my previous idea: boiling soup spillage detector. I once had a google meeting with a cooking soup to keep an eye on it and thought, heck, that seems like a nice exercise for a visual detector finetune
To make it truly ready for production science, I'd need to put more work into making the model robust. I'd also like better object tracking, so I could track multiple unique tardigrades.
If you want to see even better examples, take a look at DeepLabCut, https://www.mackenziemathislab.org/deeplabcut especially the video examples.
As for the rpi 5 combination, the power draw is relatively low. The whole thing clocks in at about 14 watts on an rpi 5 which allows us to run this platform off of a battery. With 26 tops, this setup can contend with the jetson xavier nx (21 tops, ~$500) and the jetson orin nano (40 tops, ~$500) for for a cost of around $170. Furthermore the cpu on the rpi 5 is generally more performant than the xavier nx.
Specifically, this is an excellent vision module for real-time object detection from multiple cameras if setup properly, while maintaining access to the prolific raspberry pi hardware ecosystem which is typically cheaper than the jetson ecosystem.
Do you have any suggestions on how to proceed further? So far, I've procured a Jetson, five cameras, a stand to fix and calibrate the modules, and a cam array hat to equip four cameras and the jetson. I was checking out VPU and NPUS and other hardware as well but struggling to identify compatible hardware. How can I get ahead and build such model to test and validate in 3 Months of time ?
At a hotel, we had a problem of luggage carts going missing. There’s a few ways to deal with that. A generic one that would support other, use cases would be to let the camera tell you the last room it went in. Likewise, outdoor cameras might tell you which vehicles had a customer walk in the hotel and which might be non-guests.
I ought a Kinect but never got it working correctly with a Mac. What is the state of the art here ? Is there active development anywhere ?
It’s a detailed breakdown of a technically impressive project and your main takeaway is the 5 seconds in the demo where a guy covers his face?
Kudos to the author for making something neat and sharing it.
Crazy, right?
Any scenario in which undo harm is brought upon someone because they were a passerby on the street in a video, is a such a reach that you have to question how grounded in reality they are. It's some deep "I'm the main character" level thinking.
Perhaps with a single camera you could port this to fit on a Zynq 7000 footprint with something like a Pynq Z1 or Numato Styx, which are in the $250 hobbies price point for example.
The Zynq Z7030+ chip would probably be able to get the job done but aren't as common - https://www.lcsc.com/product-detail/Programmable-Logic-Devic.... The Kintex 410T-480T are available new <$160 from reputable vendors, used <$50 from less-reputable vendors - they do have the performance (and overall IO) for this task.
Regarding the Hailo featured in the article, a few of us have been messing with it on a Raspberry Pi 5[2], and it offers more performance in a similar power envelope. The major downside is availability. I can buy the Coral on many electronics supplier sites, but Hailo seems to be selling through 'Product inquiry' right now, which is not easy to navigate as an individual!
[2] https://pipci.jeffgeerling.com/cards_m2/hailo-8-ai-module.ht...
BTW, is it possible to have 10 of these connected to a single board/cpu ?
https://www.lcsc.com/product-detail/Programmable-Logic-Devic...
(420T/480T are priced similarly new.)
Didn't have enough time to dive into it yet and still working on some other project, but this still tickles the back of my head and would be cool even if I could only run mnist on it.
To use it with an FPGA accelerator you also have to build all the "hardware" to run said openCL kernel efficiently, manage data transfer, talk to the host, etc for the FPGA. This is very foreign if you're only used to doing software design and still very nontrivial even if you've done FPGA work, though I think there are some open hardware projects around doing this.
Its hard to handle multiple data streams of video, and I think I maxed out around 20 camera feeds before the computer slowed to a crawl. NVME storage and better internet would defiantly push the limit on what is possible, but unfortunately it is impossible to get good internet where I live. Also, cameras are not cheap and neither is flash storage. But I do know this stuff sees use in both commercial and I bet they use some better version for the defense industry.
I'm thinking of an external camera (weather resistant) and hardware. The hardware could be a small computer that connects to the camera (maybe with wifi?), and runs the YOLO model.
Does anyone have a comparison between Hailo and, say, a mid or high-end GPU or a TPU?
$200 in prototype quantities [2] which is 12% the price of a 4090 - but perhaps the price drops when you order them in bulk?
They claim it compares favourably to an 'Nvidia Xavier NX' for image classification tasks, providing somewhat more FPS at significantly lower power consumption. 218 fps running YOLOv5m on 640×640 inputs.
They're completely silent about the amount of memory it has, but you can fit int8 YOLOv5m into about 20 MB so it'll certainly be an amount measured in megabytes rather than gigabytes.
Their target market is "CCTV camera that tracks cars and people" rather than "Run an LLM" or "train a network from scratch"
[1] https://hailo.ai/products/ai-accelerators/hailo-8-ai-acceler...
"One of the key reasons for the performance improvement is that RAM is self-contained without the need for external DRAM like other solutions. This decreases latency a lot and reduces power consumption."
Not sure how much RAM is included on the chip, but I'm also thinking in the tens of MB range, certainly not gigabytes.
[1] https://www.cnx-software.com/2020/10/07/learn-more-about-hai...
Process a stream of data - yes. Machine learning on large set of data - no.
So... maybe just about?
Sorry, but I can’t just ‘cool-project-bro’ this one. Does nobody else have the faintest misgivings about where we’re at right now: human surveillance as just scratching a technical itch?
Apologies if this comes over all grumpy, but wow. Seriously.