Jetson AGX Xavier
nvidia.com
nvidia.com
I've started doing machine learning experiments with it finally. (See [1] for details)
There's a few tricks to getting the best performance. You want to convert your neural network to run with NVIDIA's TensorRT library instead of just tensorflow or torch. TensorRT does all the optimized goodness that gets you the most out of the hardware. Not all possible network operations can run in TensorRT (though nvidia updates the framework regularly). This means some networks can't be easily converted to something fully optimized for this platform. Facebook's detectron2 for example uses some operations that don't readily convert. [2]
But then if you're new like me you've got to both find some code that will ultimately produce something you can convert to TensorRT, and you also need something that you can easily train. I've learned that training using your own dataset is often non-obvious. A lot of example code shows how to use an existing dataset but they totally gloss over the specific label format those datasets use. That means you've got to do some digging to figure out how to make your own dataset load properly in to the training code.
After trying a few different things, I've gotten some good results training using Bonnet(al) [3]. I was able to make enough sense of its training code to use my own dataset, and it looks like it will readily convert to TensorRT. Then you load the converted network using NVIDIA's Deepstream library for maximum pipeline efficiency [4].
The performance numbers for the AGX Xavier are very good, and I am hopeful I will get my application fully operation soon enough.
[1] https://reboot.love/t/new-cameras-on-rover/
[2] https://github.com/facebookresearch/detectron2/issues/192
For dealing with layers not supported by TensorRT, you might want to try to export to onnx instead and then use tvm[1] to compile your model for the hardware. I have not used it on nvidia boards yet, but I had good experience on other less powerful ARM boards, and as tvm docs show some examples running on jetson tx I imagine the AGX Xavier is likely fine too.
Second biggest complaint is deploying Jetsons in production environments. Dev kits aren't production stable, so you either need to build your own carrier board or find one pre-built, and frankly that's just a giant pain to do.
Third biggest complaint is having to flash Jetsons manually. Misery.
A production-ready x64 Jetson that you could order directly from NVIDIA would be my dream. Add up all of the shortcomings and overhead of ARM Jetsons and IMO you do not have a viable device for shipping AI solutions at scale.
It's not so bad, you just need a beefy ARM machine to build the containers in CI. It would be silly to build a Docker container on the Jetson itself. You would never use an embedded device for compiles and builds, why would you build Docker containers on one?
Comedy answer: iPad Pro
I mean, the iPad Pro does have a relatively beefy processor. If only you could run arbitrary code on it.
Gives you a working Alpine Linux installation which you can download and install packages for normally, all within the bounds of the normal Apple sandbox, with decent enough performance.
It doesn't have SSE or MMX yet, so eg Go and Node aren't usable at this point. But a shocking amount actually does work perfectly, so it's only a matter of time as more instruction sets are implemented.
https://blog.treasuredata.com/blog/2020/03/27/high-performan...
More expensive, more power consuming? Sure. More sanity? Massively better dev experience? Massively better production/ops experience? Absolutely.
Or just have Jenkins etc trigger the launch of one when your build job needs it.
I have a bin full of ARM single board computers, and while the hardware on all of them is pretty much up to the task, the software support from all the vendors has been terrible. I'm in the process of switching to Nvidia hoping it would be the exception.
If anyone from Nvidia is reading this, please do everything you can to convince the bosses to allocate the resources required to support a linux machine properly. It takes much more than it seems.
These devices are designed to be in production for a long time, so heavy investments now on the software support are going to give value for a long time. Rather than dragging the feet and slowly getting it right over time and devaluing the product in the process.
Fingers crossed!!
Also I can see something close to that being a competitor to the XSX and PS5. NVIDIA made the Shield and that led to a design win with the Nintendo switch -- wouldn't it be nice to play Nintendo games in VR?
Additionally it seemed like there's a decent ecosystem around it of board suppliers, which then translates into pretty good software support.
But again, not sure if that's still the case. I chose them for a project in 2015 and have not regretted it since - but have been able to use the exact same IC since then for other projects so I'm not sure if the experience would be different if starting from scratch today.
Although in a similar HN thread a few days ago, people pointed out that docker buildkit does do cross-architecture compilation, which has been slow, but works on anything!
I do this every day from a Linux x64 host using qemu-user. What problems are you hitting ?
> Second biggest complaint is deploying Jetsons in production environments. > > Third biggest complaint is having to flash Jetsons manually.
I don't think these two use cases are the goal of the Jetsons. Feel like more the goal of the EGX devices, which can be programmed, updated, etc. fully remotely.
They're made as dev kits for people building "autonomous machines like delivery and logistics robots, factory systems, and large industrial UAVs". Deploying with Docker and running devkits in production isn't what I'd call normal in such applications. Usually you need to deal with that "giant pain" of properly integrating with your hardware. Flashing would usually happen in the factory as part of the process, either by flashing the flash before soldering/inserting it or through some exposed contacts on the board.
We were able to run OpenPose[2] at 27FPS, which we found was even faster than running it on K80 on AWS p2.xlarge. It was a pain to install caffe and all the dependencies on an ARM processor, but it worked out eventually.
We were able to train and run Tensorflow 2 models quickly also. Felt like using an actual GPU at a fraction of the cost.
[1] https://www.youtube.com/watch?v=AF8zmTaa17s
[2] https://github.com/CMU-Perceptual-Computing-Lab/openpose
It's much stronger than an RPi and could fill the gap between RPi and Arm-based server platforms.
If you have thousands of remote sensors collecting Gbs and GBs of real time data, for ~1000$ you can add a "streaming" supercomputer to your sensor to analyze the data in place and save on network and storage costs.
Notice however that the announcement is for an Nvidia AGX product, which is for autonomous machines. The Nvidia "edge" products for processing data on the sensors are called Nvidia EGX.
For autonomous machines, you often need to analyze the data in the machine anyways, e.g., you don't want a drone falling if it looses network connectivity.
I've worked on fruit sorting machine and there was about 20ms to make decision if the object passed or not + there was continuous streams of 10000s of objects per second to classify. The computer vision/classifier had to be both fast and reliable about spitting the answers, which was actually more important than precission of classifier itself.
I am designing a four wheel drive robot using the NVIDIA AGX Xavier [1] that will follow trails on its own or follow the operator on trails. You don't want your robot to lose cellular coverage and become useless. Even if you had coverage, there would be significant data usage as Rover uses four 4k cameras, which is about 30 megapixels (actually they max out at 13mp each or 52mp total). Constantly streaming that to the cloud would be very expensive on a metered internet connection. Even on a direct line the machine would saturate many broadband connections. Of course you can selectively stream but this makes things more complicated.
Latency is an issue. Imagine a self driving car that required a cloud connection. It's approaching an intersection and someone on a bicycle falls over near its path. Better send that sensor data to the cloud fast to determine how to act!
On my Rover robot it streams the cameras directly in to the GPU memory where it can be processed using ML without ever being copied through the CPU. It's super low latency and allows for robots that respond rapidly to their environment. Imagine trying to make a ping-pong playing robot with a cloud connection.
I am also designing a farming robot. [2] We don't expect any internet connection on farms!
[1] https://reboot.love/t/new-cameras-on-rover/ [2] https://www.twistedfields.com/technology
Edit: Don’t forget security! Streaming high resolution sensors over the cloud is a security nightmare.
But 'edge,' as used in context of AI, is also a wink-and-a-nod that the device is inference-only (no learning, no training). The term "inference only" doesn't sound very marketing-friendly.
AFAIK the "Roomba learns the layout of your house" type of edge learning is generally done with SLAM rather than neural networks. There might be other applications for edge learning, of course.
I think the magic camera interconnect is CSI/CSI2 and it's not really flexible enough. You either have really short copper interconnects, or unavailable fiber interconnects.
What would be cool is if csi to ethernet were a thing. either low latency put-it-on-the-wire or compressed. I don't know, maybe it is. But make it a standard like rca jacks.
https://leopardimaging.com/product/nvidia-jetson-cameras/nvi...
I haven't tried them, but am considered them for project.
I got my cameras from e-consystems and they’ve got some USB3 cameras that could do it. At least I’m pretty sure. My USB3 cameras just showed up and I haven’t tried them yet.
The other thing to look at is using hardware transcoding in an NVIDIA GPU, a raw 4k feed from a 4k camera is huge, but transcoded to h.264 or better h.265 the footage is much more playable from disk in my experience. It may help with live footage. Here are some notes I made when setting up GPU transcoding.
Using xstack filter:
https://trac.ffmpeg.org/wiki/Create%20a%20mosaic%20out%20of%...
Using h265 (required to support the extra resolution of 4x4k) http://ntown.at/knowledgebase/cuda-gpu-accelerated-h264-h265...
https://devblogs.nvidia.com/nvidia-ffmpeg-transcoding-guide/
What is your monitor resolution? If you can't display the full resolution, I've found that using the nvidia encoder hardware to resize each 4k stream to 720p makes transcoding much faster. I've added my video conversion scripts to github so you can see how I've done that. https://github.com/tlalexander/rover_video_scripts
You can also contact e-consystems, as they seem eager to provide application support. Finally feel free to email me to the email in my profile, or better yet create an account and ask the question on my website http://reboot.love so other people can see our conversation and benefit from what we learn.
EDIT: I JUST saw that you meant one local display and one remote display. Sorry busy day. In that case the other poster mentioning gstreamer is spot on, and I believe econsystems has some gstreamer plug ins at least for some cameras, or maybe nvidia does...
You'll probably want to use the CSI-2 interfaces to connect the cameras, but that depends. CSI-2 was developed for cell phones and is hard to run over long distances. It's optimized for low-power and designed for very short interconnects. We had a ton of problems using it at the last company I worked for. I really wish there was a competing standard for embedded cameras.
Making the edges smarter allows them to react and adapt on smaller timescales.
If you chucked one random human on a desert island, they'd probably die. Chuck a dozen, they have a better chance of survival. Chuck a thousand, you might have a civilization.
Conversely if you say chucked 2 or 100 rabbits on an island - end result is probably going to be an island full of rabbits.
EDIT: The 8GB Module seems to be $679 here[1]. This makes the $699 or the 32 GB Developer Kit seem like a steal. Still, too expensive for play, I guess I'll stick with my Jetson Nanos for a while...
[1] https://www.arrow.com/en/products/900-82888-0060-000/nvidia
So another way to put it: its tensor cores do feed-forward calculations, but no backpropagation, and no weight updates.
Nite that the number of CUDA kernels and amount of memory available is smaller, if compared to descrete Volta GPUs.
[0]: https://developer.nvidia.com/embedded/jetson-nano-developer-...
The economics don't really make sense for TI/ADI DSPs imo. If you had an application where you needed a chip just to do DSP you'd probably use an ARM core instead - but the applications engineers at TI/ADI will gladly help you find a product in their catalog that has more features integrated into it (like ADC/DAC, even analog front ends for audio/RF, USB/Bluetooth stacks) for your product.
Basically there's no market to kill, from what I've seen.
Got it.
2. TPUv1 is a matrix multiply ASIC that requires a host CPU to do anything. This thing is a SoC that includes both a CPU and a GPU. The CPU is pretty fast for what it is - much faster than e.g. raspberry pi, see https://www.phoronix.com/scan.php?page=article&item=nvidia-j....
3. not sure how you know whether this is more expensive than a TPUv1, since the TPUv1 was never sold or available outside of google.
A much better comparison would be between this and the Edge TPU development board.