A closer look at Raspberry Pi RP2040 Programmable IOs
cnx-software.com
cnx-software.com
If you've ever played one of the Zachtronics games (in particular Shenzhen IO) it feels like the programming in one of those games. Tight constraints but major possibilities when you think about it for a while.
SATA is a 6Gbps port, while HDMI is a 10Gbps port.
The PIO ports discussed here are on the order of 100kHz, roughly 1-million times slower than HDMI, and 600,000 times slower than SATA.
Your discussion point of "here's a VHDL block" seems to understand the general issue. You need a non-trivial amount of FPGA-magic (LUTs) to implement logic and routing at those speeds. SATA has some kind of error-correction code if I remember correctly... so its not exactly easy to parse those messages.
I am just surprised that this is not more widespread among major players as a way to reduce costs and increase flexibility. Though I'm pretty sure I'm overlooking the core of the issue here haha
I think I’m many cases the flexibility is great during the prototype phase. But those are used at lower volume. When you move to production, you’d want to have the cheapest BOM as possible.
That's not what the datasheet[1] says:
When outputting DPI, PIO can sustain 360 Mb/s during the active scanline period when running from a 48 MHz system clock. In this example, one state machine is handling frame/scanline timing and generating the pixel clock, while another is handling the pixel data, and unpacking run-length-encoded scanlines.
Still not SATA speeds though.
[1]: https://datasheets.raspberrypi.org/rp2040/rp2040-datasheet.p...
https://hackaday.com/2021/02/12/bitbanged-dvi-on-a-raspberry...
Even though I did not intend to say we should have SATA and HDMI on the RP2040 itself (I didn't know if it was possible), it still proves that having realtime control on the IOs opens the door for way more functionalities than SPI/I2C/UART specific ports. All of it using the same SoC and potentially less pins.
Having the same level of control on any device would be beneficial in my opinion.
> The PIO ports discussed here are on the order of 100kHz
PIO runs at the system clock, 125 MHz by default, overclocks of over 400 MHz have been reported stable. A single PIO can clock out 32 bits every cycle (with a DMA and a memory system that can feed this), giving you a total of 4 Gbps.
Running at full throttle like that, especially for a decent length of time is tricky if you want to actually do anything other than blast out bits but 16 or 8 bits per cycle is a lot more straight forward, so 2 or 1 Gbps.
DVI output has already been demonstrated, running two displays at 480p: https://github.com/Wren6991/picodvi
The DVI is maybe more of a party trick than something you'd do in production hardware but it does demonstrate how capable the PIO can be. You could happily implement the same concept in a more performant device and reach 10 Gbps or more in a reasonable way.
The differentiate their price according to features and so they extract more money. And by writing your code to specific peripherials, it's harder to switch to another mcu.
And they have large libraries of proven hardware peripherials and code which make it harder for competitors to enter. Why would they want to compete with open-source pio libraries?
The raspberry pi foundation doesn't care about all that. So they created this chip.
https://twitter.com/ZodiusInfuser/status/1357067388928335873
0 -> 1 -> 2 -> 3 -> 0 -> 1 -> 2 -> 3... is counting up. Formally: newstate - oldstate == 1.
0 -> 3 -> 2 -> 1 -> 0 -> ... is counting down. Formally: newstate - oldstate == 3.
0 -> 2 is ambiguous. Carry on the "same direction as the last time". Formally: newstate - oldstate == 2.
0 -> 0 is standing still, no movement. newstate - oldstate == 0
-------
Except of course: 0, 1, 2, 3 are gray coded, not normal binary coded.
* 0 == 00
* 1 == 01
* 2 == 11
* 3 == 10
Anyway, convert the 00 / 01 / 11 / 10 raw hardware bit-stream into the above numbers (0, 1, 2, 3). Then it becomes VERY easy.
Any I/O is basically compression. So lets think about what we're actually outputting. We're converting a raw bitstream from two GPIO pins paired (Ex: 00, 01, 01, 01, 11, 11, 10, 00, 01, 11), into a singular message: such as +6.
There are two applications of a quadrature encoder that I can think of.
1. Virtual Pot -- Converting the messages into a kind of "-100 to +100 slider", such as a volume control knob. In this case, you want a regular update schedule at human-interface speeds (~1000 updates / second or slower).
2. Rotational Velocity sensor -- Converting the messages into a speed. You're willing to batch the bitstream up as slowly as possible, maybe waiting for a rollover event (+128 or -128). And you just interrupt on those overflows.
-------
Since the input format is already set, we now think about the output format: how to represent +6 and or -12?
Based on the PIO instruction set, it seems like +6 and -12 probably would be easiest as two separate registers in two different state-machines.
Every time an "increment" is detected, the FIFO-associated with +1 gets a +1 bit added to it. Every time a "decrement" is detected, the 2nd FIFO associated with -1 gets a +1 bit added to it.
----
In pseudocode:
IncrementSide(){
currentState = 00; // default
while(1){
nextState = in(GPIO0) | in(GPIO1);
if(currentState == 00 and nextState == 01){
push 1;
}
if(currentState == 01 and nextState == 11){
push 1;
}
if(currentState == 11 and nextState == 10){
push 1;
}
if(currentState == 10 and nextState == 00){
push 1;
}
currentState = nextState;
}
}
"nextState" and "currentState" probably is just the X and Y registers of the PIO state machines. The above is "very pseudocode" as I don't really understand PIO yet.I'm saying "push 1", but "push 1 into the OSR". It looks like the OSR has some kind of auto-push mechanism, but if that doesn't work then a manual-push might be needed in the code proper. Hard to say from the docs alone.
DecrementSide would be just the inverted if-statements:
if(currentState == 00 and nextState == 11){
push 1; // Decrement side checking for a decrement
}
--------This above methodology would only work for #2 ("velocity"), and fails for #1 ("positional"), because the increment-side and decrement-side would be updating at varying rates.
I probably can make a positional-decoder instead of a velocity one using similar principles. Or maybe by somehow ensuring that the +1 and -1 messages go into the same FIFO to be picked up by the host CPU.
The idea is route "0" messages to /dev/null, compressing the input stream. The host still sees the important +1 and -1 messages. The exact mechanism for that is up for debate, but the simple if-else state machine above seems to accomplish that to a limited extent.
------
EDIT: Hmmm, maybe a singular FIFO can be done if we have push1 and push0 for the two kinds of messages: 1 representing +1 and 0 representing -1.
Anyway, my point is that there's lots of solutions here. It doesn't seem very difficult to me conceptually. Just a lot of experimentation needed to know exactly how those 9 instructions work and the exact mechanisms of the ISR / OSR / X reg / Y reg.
Ex: 00 -> 10 -> 00 is interpreted as 0 -> 3 -> 0, or -1 transition then a +1 transition.
01 -> 11 -> 01 is interpreted as 1 -> 2 -> 1, or +1 followed by a -1 transition.
All possible combinations of bounces (00 -> 01 -> 00, 00 -> 10 -> 00, 01 -> 11 -> 01, 01 -> 00 -> 01, etc. etc.) have this +1 / -1 or -1/+1 property.
Example for the ESP32: https://github.com/espressif/esp-idf/tree/73db142/examples/p...
I would not be surprised if you don't need PIO for this on this chip.