Raspberry Pi Pico does line rate 100M Ethernet
github.com
github.com
RP2040: https://datasheets.raspberrypi.com/rp2040/rp2040-datasheet.p...
RP2350: https://datasheets.raspberrypi.com/rp2350/rp2350-datasheet.p...
- The original Pico was not built around the RP2040 as its central part ("includes" sounds to me like it was an addition)
- The Pico 2 includes a RP2040 (in addition to the RP2350) which runs PIO
Neither of which are true. I'm guessing some other people had a similar reaction.
This has made much faster systems not being able to process packets at line speed. A classic was that standard Gigabit network cards and contemporary CPUs were not able to process VoIP packets (which are tiny) at line speed, while they could easily download files (which are basically MTU-sized packets) at line speed.
context switching between processors will reduce cache coherence and hence hits, but yea, it might be worth the tradeoff on busy systems
It's a Cortex M33, so there's no meaningful cache to speak off. Access to all memory takes essentially the same amount of time. If you're really worried about access time you could probably use SRAM banks 8&9 (each 4k, with their own connection to the AHB crossbar) and flip-flop between the two - but I highly doubt it's going to have a measurable impact.
It's a physical layer, so yeah, of course it will.
At first I thought it was the new Pico 2 (RP2350), but no, it’s the old Pi Pico with RP2040.
Is there enough room to have it control the ethernet port for another weaker or perhaps more powerful microcontroller?
Can you combine multiple picos with one being the ethernet stack and another that modifies certain packets?
Are there any other interesting things that can be done?
Well there is a whole unused core and plenty of built in SRAM. Seems like a good way to have an open-source version of Wiznet chips [1]. It could support full protocol offloading like Wiznet's or a lower-level raw packet sender/receiver like the ENC424J600.
I don't have the time to give it a shot myself, but I could try to help if needed.
Back when I was doing a dumb-server/smart-client desktop environment. Something like this would have been pretty cool. It needed a tiny API to save files, but the bulk of the environment worked as a static server.
It would be interesting to see a short writeup of what kind of magic was required to achieve this, as there have been multiple failed attempts before this.
I'm also curious about the performance boost from 2.81Mbit/link failure at 150MHz to 65.4Mbit/31.4Mbit at 200MHz. That doesn't sound like basic processor bottlenecks, but rather some kind of catastrophic breakdown at a lower level? Does it just occasionally completely fail to lock onto an incoming clock signal or something?
If so, might it be possible to use two RX PIOs, automatically starting the next one via inter-PIO IRQ when a packet is finished? That'd give you an entire packet receive time to reset the original PIO, which should be plenty.
.wrap_target
irq set 0 ; Signal end of active packet
start:
wait 1 pin 2 ; Wait for CR_DV assertion
wait 1 pin 0 ; Wait for RX<0> to assert, signalling preamble start
wait 1 pin 1 [2] ; Wait for Start of Frame Delimiter, align to sample clk
sample:
in pins, 2 ; accumulate di-bits
jmp PIN, sample ; as long as CRS_DV is asserted
.wrap
It's run at a fixed 100 MHz, regardless of system clock speed, via controlling
the PIO execution rate a fraction of the system clock speed. So, for a 300 MHz
system clock, the PIO is clocked once every three system clocks. I'm speculating that the extra two clocks (at 300 MHz) allows more setup time to
the PIO inputs. The [2] above enables an extra two PIO clock delays before
executing the next instruction. I tried changing this from zero to three at 100 MHz system clock (i.e. a PIO system clock divisor of one), and wasn't able
to fix the problem. Though it should be noted that the LAN8742 isn't a very
forgiving chip - I've seen RX Data Valid (DV) go metastable when the TX clock
is interrupted/changed, so another pass through might be worthwhile.BTW, Sandeep's original code clocked the RX PIO SM at 50 MHz, pushing all the samples to the output FIFO, and relied on the processor getting interrupted at the falling edge of DV to figure out what samples constituted a packet.
Is this an effective rate, or just the reflection of a hardware limit?
7 byte preamble
1 byte SFD
6 byte dst MAC
6 byte src MAC
2 byte ethertype or length
46-1500 bytes of payload (ignoring “Jumbo” frames and 802.1q tags)
4 byte CRC
12 byte IFG (which is silence, but still counts for time on the wire)
Add it up and you have 1538 bytes “on the wire”.
TCP overhead for IPv4 is 20 bytes for IP(v4) (no options) and 20 bytes for TCP (again, no options).
So 1460 bytes of data for 1538 bytes on the wire. 1460/1538 = 0.949284
So for 100M Ethernet, 94.9284Mbps is “perfect”.
As far as I can tell, someone has figured out how to send Ethernet packets at a relatively high rate using hardware with very limited CPU. Cool, but what can you _do with that_? If the RPi Pico has the juice to run interesting network _application-level traffic_ at line rate it's more intriguing, but I doubt that anyone's going to claim that can serve web traffic at line rate on this device, for example.
What am I missing?
For example, the Oric-1/Atmos computers recently got a project called "LOCI" which adds USB support to the 40-year old computer[1], by using an RP2040's PIO capabilities to interface the 8-bit DATA bus with a microcontroller capable of acting as the 'gateway' to all of the devices on the USB peripheral bus.
This is amazing, frankly.
And now, being able to do Ethernet in such a simple way means that hundreds of retro-computing platforms could be put on the Internet with relative ease ..
[1] - https://forum.defence-force.org/viewtopic.php?t=2593&sid=2d3...
This "very limited" microcontroller has two cores. Both of them can execute about 25 instructions per byte for generating "application-level traffic". You could definitely saturate a 100 Mbps connection with just one core.
Should be cheap, right? Though 1Gbit version might still be expensive..
This was AFAIR based on empirical knowledge, nothing scientific.
So a Pi Pico running at 300MHz pushing 100Mbit is something that is not totally unexpected, if you consider the low-power, low-cost CPU design in a Pi Pico (and the fact that you have to push the bits manually on the wire).
It's still a nice feat that they pulled this off!
For 100M Ethernet this requires 148,809 packets per second.
Edit: for 1538 octet frames, one need only process 8,127 packets per second.
What could go wrong?
As the article calls it, the gold standard. If a device is capable of forwarding/switching packets at the smallest packet size line rate on all interfaces at the same time you don't have to think too much about its performance when designing your network. Haven't worked much with hardware for a few years but it was common that Cisco switches were not capable of this.
IMIX makes sense on devices that are not capable of small packet line rate like firewalls where bandwidth is much more costly and need to be sized appropriately.
> Both DUTs can achieve line rate performance on all ports with an NDR of 170 Bytes for the 88-LC0-36FH-M line card and 215 Bytes for 8201-32FH router. Same values were observed for both IPv4 and IPv6 traffic. This exceeds all real-life deployments requirements regardless of position in the network.
The 9000 series analysis reports something like 400B packets to hit line rate.
Fundamentally, everyone has to scale their internal bus width and clock rate to hit the headline numbers, always at the cost of small frame performance.