Memory Mapping an FPGA from an STM32
serd.es
serd.es
* non-4-byte-sized writes randomly lost about 1/million writes if QSPI is writeable and not cached
* non-4-byte-sized writes randomly rounded up in size to 2 or 4 bytes with garbage, overwriting nearby data about 1/million writes if QSPI is writeable and cached
* when PC, SP, and VTOR all point to QSPI memory, any interrupt has about a 1/million chance of reading garbage instead of the proper vector from the vector table if it interrupts a LDM/STM instruction targeting the QSPI memory and it is cached and misses the cache
Some of these have workarounds that I found (contact me). I am refusing to disclose them to STM until they acknowledge the bugs publicly.
I recommend NOT using STM32H7 chips in any product where you want QSPI memory to work properly.
After burning enough company time chasing bugs through ST's crappy silicon, I've had to just swear them off entirely. We're an Atmel house now. Significantly fewer (zero) problems, and some pretty nifty features like UPDI.
I suspect ST only ever tested it with their single PSRAM they intend this mode for. My intent is to use indirect mode and manually poke the peripheral, though DMA will have to happen still.
Back on the PIC32MX platform there was a similar type of bug that doesn't exist anywhere else but to me: If any interrupt fires while the PMP peripheral is doing a DMA, there is a 1 in a million chance that it will silently drop 1 byte. Noticed this because all my accesses were 32bit (4 bytes) and broke horribly at the misalignment. The solution is to disable all interupts while doing DMA.
So far I have it working quite reliably (my test firmware does a loopback test with 100K reads/writes of a 32-bit register at the start that I had written with intent of using it for link training of the PLLs to optimize read/write capture timing but never ended up using as such) and my iperf test can push tens of thousands of packets per second without issue.
The IMXRT1064 is around $7 and is also an M7 core with an HS USB PHY, programmable PLL-connected LVDS clock output, 2 EMACs, excellent hardened IP generally.
The big thing holding me back was that their crypto accelerators were all locked behind NDAs (a dealbreaker for F/OSS work) while the ST ones are documented in the freely downloadable datasheet you can just google up.
But I did find some third party wrapper libraries that seemed to be able to use the crypto registers so it might be possible to figure things out from that. I haven't tried yet.
The other issue I had with the RT is that they lacked internal flash so PCB complexity is slightly higher than with a STM32.
Keep in mind the dual-core 11xx chips are a bit harder to boot than the rest of the line - but you probably need the power domain flexibility for most FPGA projects (1064 has way fewer practically-usable 1v8 banks.)
> crypto accelerators were all locked behind NDAs
I've been able to use every bit of hard IP and high-assurance boot from registers using no vendor code whatsoever.
Here's what you are looking for:
https://github.com/JayHeng/imxrt-level2-boot/blob/master/dev...
> The other issue I had with the RT is that they lacked internal flash
The IMXRT1064 has a 4MB Winbond QSPI chip in-package, by the way!
> PCB complexity is slightly higher than with a STM32.
The Xilinx FPGA that is sitting next to your MCU incurs multiple orders of magnitude more PCB-complexity than a little QSPI flash, haha.
You'll hit almost no bugs if you keep accessing the same address in a loop. Lucky you :)
Have you hit issues with the FMC? From what other people are telling me, the OCTOSPI is full of land mines and the FMC is pretty decent. The worst errata I've encountered so far is two dummy clocks with CS# asserted at the end of a read burst.
MicroKVS expects to be able to memory map data fetches (uncached), but is fine with using indirect access for writes.
This saves a chip on the board, reduces the amount of PCB routing required, and eliminates use of the sketchy OCTOSPI peripheral entirely. Testing that out is on my list of things to do on this board eventually.
AXI (and all memory-mapped bus protocol schemes) becomes very very pleasant. SV interfaces get you 5% of the way there, though!
Also - I was under the impression that S1000-2M is a higher-end material, not cost-optimized? (But not Rogers, of course.)
For higher end digital work I typically reach for Taiwan Union TU872SLK (Df 0.009) which also has a better range of prepregs and glass styles available to help minimize fiber weave effect. Still quite a bit lossier than e.g. RO4350B but far less expensive and if you have decent equalizers on your SERDES the difference is typically not significant unless you're making some kind of humongous backplane. I get wide open eyes with just a tiny bit of post-cursor emphasis on the TX FFE at 10.3125 Gbps on TU872SLK for my typical shortish high speed tracks (FPGA to SFP+ cage).
I haven't seen these as directly-advertised options at any of my usual suspects.
I have some 25/100G stuff in the pipe for probably some time next year that I plan to make with them too.
Their website undersells, I get the impression most of the actual sales contacts are word of mouth. I talk to my sales rep by skype mostly (the alternatives are expensive international phone calls or wechat).
The really cool thing is that you get a 10+ page QA report with every order including measured copper/dielectric/soldermask thicknesses, hole sizes, ionic contamination measurements, and a ton of other metrics. And they send the TDR strips and polished cross section with every order as their way of saying "look, we actually did the QA, double check our measurements if you don't trust us". (I actually have repeated some of the measurements to spot-check and got results within a few percent of their QA department, no surprises there).
And they don't make silent gerber changes or anything. They do a full CAM review and send you working gerbers and a list of suggested DFM tweaks for you to sign off before beginning manufacture. If something doesn't look right you have a chance to say "wait there's a problem".
For example, one time they wanted to make a really large width adjustment for impedance on some RF traces that I had carefully modeled in an EM solver. But they didn't make a bad board without telling me, they flagged it on the CAM review and we went back and forth before realizing the mistake was on their end (they had calculated impedance assuming solder mask over the traces, while they were actually exposed copper). They re-ran the numbers which then closely matched my simulations, I signed off on the modified design, and the board was manufactured without issue.
For some development hardware, we had Elixir running on the ARM of a Zynq Ultrascale, running in tandem with some digital logic. It required one C code "port" that integrated with UIO to expose the registers to our application and then we had a great programming environment.
Elixir for embedded doesn't get talked about that much, but that is actually the origin story of Erlang (software component of telephony hardware). Basic language features like binary pattern matching work very well, and the concurrency approach makes it very easy to write clean performant real-time software. We had a lot of functionality that did used digital logic and then had the stateful stuff in software and it worked very well.
Plus, I could then do stuff like trivially spin up a Web UI with a graphical display of all the register state, that updated live (Phoenix LiveView). And be happy that that running wasn't going to interfere with the realtime stuff.
We did this using Nerves which is a Linux platform set up to boot the BEAM and nothing else (e.g. no init system, just a special pid 0 binary that boots the BEAM and lets that handle all other processes). It had some plus points like making firmware upgrade trivial and simplifying the system, but not being a "normal" linux platform was a bit irritating sometimes. You could equally well just run Elixir as an application normally.
You get to own every aspect of your toolchain and with that will come a lot of power.
Are you familiar with:
https://github.com/corundum/corundum
Perhaps you can build a support package for your platform.
Would not surprise me if the M4 was there and fused off (i.e. same die as multicore H7 offerings), but it's not active.
Do you know if it's fabbed in house, TSMC, or Samsung? I've seen ST silicon from all 3 foundries but the only thing I've seen stated publicly is 40nm. When I get it opened up it should be easy to tell, TSMC and Samsung processes have distinctive features on them that I recognize by sight.
Now this would be a cool blog post!
The MCU is for control plane only. Several hundred Mbps between the control and data plane is more than enough for a SSH management CLI and poking registers on the FPGA to move a port to a different VLAN in response to a CLI command or add an ACL rule or something.
But a STM32 is more than sufficient for the management interface on both.
(Like I like retro stuff and during COVID I bought an old DSP56k dev board with a book about the assembly language but oh boy, oh dear)
PIO has extraordinarily sloppy timing (skew in all categories) compared to the cheapest and smallest FPGAs.
Checking sketchier places Win-Source has the CLG400 package for $22.20 and even the cheapest aliexpress seller wants $4.84 for something marked as a 7Z010 that may or may not be legit.
Also "fight the chip" is pretty much the definition of what I did last time I did a zynq project. Just give me a plain FPGA and MCU with no wizards or GUIs or automatic code generation.
I've ordered trays (and they send the OEM tray) - unique barcodes, legit.
> Just give me a plain FPGA and MCU with no wizards or GUIs or automatic code generation.
You can pretty much cut out all of their tools and get a pure Yocto/Vivado TCL build for the bitstream for the 7 series Zynqs. Very low touch.
Their IO planner (in the Vivado IP integrator) is somewhat necessary for complex peripheral scenarios and is one of the few things I ever use Xilinx GUI applications for anymore.
On the chance they're half reasonable, thanks for the link.
I've previously struggled to roll the dice for higher end parts as the cost difference isn't as extreme and had some obvious reballed parts a few years ago. If they're OK then their $20 XC7K325T will be at the top of my list...
In the way that that Aliexpress vendor lists the 7010 parts at 1/10th the price of LCSC, some of their $20-60 listings are also shockingly cheap in comparison.
Thanks for your write-ups btw, been following glScope development for a while.
Nowadays, I only often wish I had their ARMv8 chips instead of the old ARMv7 32bit architecture because that’s just showing its age, but that’s par for the course of using ARMv7, and doesn’t affect the PL side (much, except for interfacing sometimes).
I've always wanted to do an FPGA project but haven't looked seriously into where to start. Can the Zynq 7010 handle something like data transfer from a 4K image sensor to a USB 3 transceiver?
> PIO has extraordinarily sloppy timing
Do you have any data on this? I was under the impression the PIO has fairly precise timing if you set up your clocks right, but maybe I've been misled here.
The Zynq 7010 is less-than-ideal for this because you'd have to use some kind of USB3 interface PHY - which would increase cost and be pretty limited functionality-wise.
If you use an FPGA with transceivers (some 7 series Artix chips - https://www.lcsc.com/product-detail/Programmable-Logic-Devic..., most 7 series Virtex/Kintex chips, all US/US+ chips), you can implement USB3 without an external PHY: https://github.com/enjoy-digital/usb3_pipe
The Z7010 probably has the area to do this type of translation but not the transceivers. There are other chips in the 7 series Zynq family with capable transceivers, but they are much more expensive ($15-$35 from CN).
> I was under the impression the PIO has fairly precise timing if you set up your clocks right, but maybe I've been misled here.
No measured data, but when I was once implementing JTAG and SPI at 50MHz+ with an extremely overclocked chip, the edges were very inconsistent in relation to each other and in pulse width - 5-15ns (estimating from memory, they were sloppy.)
PIO is very precise within its specified capabilities, this range is just very low compared to cheap FPGAs.
> PIO is very precise within its specified capabilities, this range is just very low compared to cheap FPGAs.
Good to know, I don't plan on pushing RP2040s to their limit any time soon. They're still excellent for lower speed projects.
Embedded is about solving problems more physical in nature, as you are physically closer to reality in nearly all aspects.
--------
An MCU + FPGA project could implement... say... the VFIR IrDA (Infrared) protocol at 16Mbit.
Traditional IrDA is widely supported at SIR and MIR levels (upto 1.152MBit or so). Anything faster and the equipment has basically been lost to the 1990s (and never was very popular anyway).
IrDA I'd explain as a remote-controller on steroids. Its infrared based (like TV Remote Controllers), so you need to line up both devices and have them looking at each other. Infrared can reliably travel about 3 meters over the open air in a variety of conditions. IrDA allows for bidirectional communications. Its a truly wireless protocol, albeit one that requires significant alignment to function correctly. But ~3 meters is good range and practical for many applications.
Nominally, you could use an entire MCU to handle the encoding / decoding of these light-pulses. However, that's a bit redundant. Its far more cost efficient to dedicate a few LUTs in an FPGA to the task.
Yes, the MCU is needed for the final application-level / OSI layer 4/5/6/7 aspects of IrDA protocol. But the lowest PHY and MAC levels of the protocol can and (probably) should be a small section of FPGA.
Upgrading from standard MCU 1MBit to 16MBit would be a 1600% improvement to communications compared to what's readily available with commercial-off-the-shelf solutions. If you've determined that IR Communications is good for whatever purpose you're using, maybe the 1600% improvement is going to be useful.
------------
EDIT: The "physicality" of this is because photodiodes react very quickly to light pulses. And an expensive enough transistor can amplify that at the ~100MHz speeds needed to run VFIR (at least in theory. I've never done this).
The FPGA (or MCU if you go that route...) just needs to clock at 100MHz or so, and interpret the start-of-frame and end-of-frame signals, while also interpreting a few other low-level details. Overall, this turns the sequence of light pulses into bits-and-bytes for higher-level processing (which code can and should handle).
That's a bit of an overly broad statement.
I have several times at work done projects that use FPGAs to do things faster than a computer can do, e.g. some specific processing of a full 10Gbps UDP stream, which is much easier to get 100% reliable without packet drop in digital logic than it is in software land.
The FPGA+CPU combo allows you to do a LOT of things. I tend to use just a straight Zynq so I already have my memory buses wired up from processing system to digital logic, but this is an interesting architecture.