In llvm binaries, the code is abstracted/re-ordered, optimized/deduplicated, and flow may operate completely differently with minimal changes. One may gain 11% to 15% raw computational throughput performance, but have no latency guarantees on what the pipeline will do or when. This will often eventually cause intermittent CDC mismatch under load, as the compiler pseudo-randomly decides to do something silly every time something is updated.
https://en.wikipedia.org/wiki/Clock_domain_crossing
The current best solution is to side-load special fpga memory access modules into the multi-tasking kernel space for handling DSP data stream filtering on Zynq.
https://en.wikipedia.org/wiki/Finite_impulse_response
https://www.analog.com/en/resources/evaluation-hardware-and-...
Best of luck, =3