PRU tips: Understanding the BeagleBone's built-in microcontrollers
righto.com
righto.com
For example, in a recent project, I used one of the PRUs to generate a precise 40MHz square wave clock signal with 40% duty cycle, and the other to read the signal pin of a camera module into shared RAM. It worked extremely well, allowing me to obtain camera data at hundreds of FPS, and freed up the main CPU to do some fairly heavy image processing - all without involving an expensive camera capture rig or an external PC.
Bela is being used at the Augmented (Musical) Instruments Lab in London (and now the community) to make rich, responsive digital musical instruments. It has a Web IDE, supports C++, Pure Data, Faust, SuperCollider, etc., but again thanks to the BB's PRU's, it supports audio-rate sensor sampling!
Site: http://bela.io/
Code: https://github.com/BelaPlatform/Bela
Videos: https://www.youtube.com/channel/UCgWd1Q2dcWdqCGNl5BijFsA
Paper: "An environment for submillisecond-latency audio and sensor processing on BeagleBone Black" http://www.eecs.qmul.ac.uk/~andrewm/mcpherson_aes2015.pdf
while (1)
{
uint32_t sample = __R31;
uint32_t change = sample ^ pEcap->sample;
if (change)
{
// Calculate which bits have seen a desired edge (it changed, and we trigger on that change)
uint32_t edge = change & trigger;
if (edge)
{
uint32_t ts = ReadTimestamp();
int bit = 1;
int iBit = 0;
while (iBit < 30)
{
if (edge & bit)
{
// Store the ts (or the difference in ts) into the capture table for this bit
// and increment the slot index (aICap)
pEcap->ECAP[iBit][pEcap->aICap[iBit]++] = (pEcap->differential & bit)? (ts - tsLast[iBit]) :ts;
pEcap->aICap[iBit] %= 4;
// Next time we calculate the difference from this time
tsLast[iBit] = ts;
}
iBit++;
bit <<= 1;
}
pEcap->edgeDetected |= edge;
}
pEcap->sample = sample;
trigger = sample ^ pEcap->edgeUpDown;
}
}
It detects an edge within a few dozen nanoseconds (low jitter). A while loop like this would kill a main processor thread; and it would have terrible latency when other threads were scheduled during a trigger event.And it detects edges on 30 pins in parallel! I could work on the "which pin had a trigger" code to reduce the period calculation from X30 to log(30) but I have no need for that fine latency in my current application.
This is great for projects where you control the whole stack and don't plan to support anything else. I guess the only downside to going down this path is that you're locking in the Beaglebone Black as your sole hardware platform and losing some modularity.
For example it would be great to control a 3D printer by running something like OpenGB (http://opengb.readthedocs.io/) on the main CPU and something like Marlin (https://github.com/MarlinFirmware) on a PRU. But that would require a Beaglebone-centric approach which wouldn't work on other hardware combinations.
That's just an observation though - I'm really impressed that the PRUs are there!
I'd say that most embedded related projects will have to deal with newer revisions of their hardware, maybe because an older part is no longer available or because people came up with more intelligent or less buggy circuits over time. So that's already a few (prob. very minor) variations you'll have to support.
And besides very trivial projects, you should always try to have at least one "dummy" implementation of everything, to facilitate automated testing of your code.
So, in your case, yout 3D printer controller could support the internal PRU of the Sitara, an external servo controller, or some dummy library that just logs positions to a file, for testing.
OpenGB uses an abstract base class (called IPrinter) to describe a printer interface. At the moment there exists a Marlin implementation and a Dummy implementation (as you describe) of IPrinter.
Other comments mention BeagleG and MachineKit. It should be pretty trivial to add IPrinter implementations of both of these.
Thanks for the inspiration! :)
MachineKit is using the PRUs with LinuxCNC already and you can control 3D printers with it!
Have raised an issue for this: https://github.com/re-3D/opengb/issues/18
I've added an enhancement issue to OpenGB for adding BeagleG support: https://github.com/re-3D/opengb/issues/17
Kudos to Ken for lifting the curtain a bit on the Sitara's PRUs!
It's a shame there isn't more developer availability of these cores. I'm kind of shocked that the way to program these is still just through assembly, but really not that surprised.
This is not true anymore (was in the beginning though)- there is a C/C++ compiler available.
Regarding C vs asm, I started writing some 8mhz 4-bit capture code in C but found asm was easier to reason about timings. One line of code is 5ns, done. (apart from some memory access).
There is also a C compiler available, I haven't used it, but Tridge (of Samba fame) gave a demo at linuxconf a couple of years back in which he launched a plane (remotely) and ran the guidance software on a BB, with the PRUs doing all the servo/etc work coded in C ..... and compiled the linux kernel onboard at the same time ....
I'm not saying that using an FPGA would be easier necessarily--indeed, likely not just because the tools a so terrible--but there are many options. Last I checked the documentation for the PRUs was pretty bad. Mostly a smattering of wiki pages and a couple powerpoints. And even that is a huge improvement over even just a year or two ago.
1) Read the fine print when it comes to processor manufacturers telling how much time something takes. Although PRU subsystem is deterministic and __most__ instructions take 5ns, there are quite a few cases which take an order of magnitude more time[1]. Sure, the access times might be deterministic but that doesn't make it easy to know how much something will take.
2) Remote-controlled airborne vehicles need a lot less computation power than I expected. PX4 runs it's "main loop" at just 400Hz and I've seen PX4 or ArduPilot devs (might even be tridge) saying that 50Hz would be enough. Sure, you need accurate timing for PWM outputs (~1us resolution) and most important: low jitter.
The first point has bitten me personally- I found it non-trivial to get reliable < 100ns interrupt jitter on Cortex-M4. It really got down to what was happening on the bus between CPU/memory/peripherals at the point interrupt was supposed to fire (e.g. getting the documented latency of 10-odd cycles when CPU is idle but a lot more when there is a DMA transfer in progress).
1: http://processors.wiki.ti.com/index.php/AM335x_PRU_Read_Late...
I would turn to the Cypress PSoC stuff for that actually. http://www.cypress.com/products/microcontroller-mcu-and-prog...
They have just that little bit of programmable logic to catch whatever silly thing that just has to respond in 400nS, for example.
The PSoC 4 series also operates at 5V I/O levels. That is actually getting remarkably difficult to find, nowadays.
> If you want to perform real-time operations, the BeagleBone's ARM processor won't work well since Linux isn't a real-time operating system.
I get that Linux is not a real time system, but the article seems to be making the implication that the ARM processor cannot support real time operations. I was under the impression that the real-ness of time was determined solely by the operating system and not by the hardware. Is this not the case, or is the article just making the assumption that no one is going to port a real time operating system to the BeagleBone?
So yeah, it's more than just the OS.
So they are perfectly fine for some subset of real-time tasks, where you just need some sort of guarantee like "this thread will not be starved of CPU for 10 milliseconds". But that kind of latency and jitter isn't going to work out when you want to generate something like a fast PWM signal or send out audio over I2S in software or generally do any task that relies on reproducible (fast) timings. You could connect a servo to your real-time Linux generated PWM signal and it would tremble like it's drunk.
And honestly, modern beefy ARM cores like the BeagleBone uses have been tuned for throughput above everything else that you will find it extremely difficult to get anywhere near the timing that a PRU can do and still have anything resembling a useful system.
There's a nifty presentation from TI [1] where they have the ARM core do nothing but toggle one of it's GPIOs and it takes 200ns for that to go through all the various layers of caches, pin muxing, buses, interconnects, what have you. And that's a GHz processor!
1: http://processors.wiki.ti.com/images/3/34/Sitara_boot_camp_p...
The PRU ran a bytecode execution engine that was handcoded in assembly. The core of the engine was a loop:
- perform the routine for the current bytecode
- when the routine finished, spin on the PRU's embedded timer (the IEP) until 256 ticks (1.28us) had passed since the start of the cycle
- advanced to the next bytecode and repeated
External devices were connected either via I2C (those that didn't need super-tight timing) or to the pins controlled by the PRU. Some of the bytecodes controlled execution flow like loops while others were specific to the application, things like turn light on, move lens, trigger camera, etc.
There were a few more bells and whistles, like memory locations where flags were set by the engine to let the CPU know how progress was going.
It was a fun project. Unfortunately, the manufacturer of one of the core devices decided to stop making it, so that start-up had to go in a different direction.
Code: https://github.com/machinekit/ PRU links: http://blog.machinekit.io/2013/06/beagle-bone-pru-links.html