There are boards with 4 sfp+, ones with 2 sfp+ and 2 QSFP+, and even one with 4 QSFP28 (and UltraScale+ XCVU9P)...
https://www.aliexpress.com/store/group/FPGA-DEV/620372_25030...
they sound like great targets for your work...
Straddling is an attempt to mitigate this issue. Instead of only staring packets in lane 0, the interface is adjusted to support starting packets in several places. Say, byte lanes 0 and 32. Or 0, 16, 32, and 48. Now, when you have a packet end in byte lane 0, you can start the next packet in the same clock cycle, but in byte lane 16 or 32. This increases the interface utilization. The trade-off is now the logic has to deal with parts of two packets in the same clock cycle, and it has to deal with multiple possible packet offsets.
The specific annoyance with PCIe packets is that the max payload size is usually 256 bytes, but every packet has a 12 or 16 byte TLP header attached, which really screws things up when combined with the small max payload size.
Right now, 40GB is the sweet spot in lower cost surplus hardware: E.g. you can get Arista DCS-7050QX-32 for about $500 shipped on ebay all day long.
100GB/25GB switches are still really expensive.
Funny you mention that switch, we bought one of those off of eBay for our testbed as it supports PTP.
Also, for optical switching applications, one of the most important factors is how long it takes to bring up the link after switching. Because of this, we have no interest in spending time on 40G and 100G interfaces because interlace deskew takes hundreds of microseconds, and 100G also requires FEC which takes hundreds of microseconds to lock. So we're focused on 10G and 25G and running multiple links in parallel, which also provides more architectural flexibility. I added 100G support for three main reasons: the CMAC license is free, so why not?; supporting 100G makes the project a whole lot more interesting than only 10G or 25G, and it provides a simple way of testing the core NIC datapath.
Got any pointers to the sort of optical switching components you're using?
[I've been out of the networking business professionally for almost a decade now, so I'm a bit out of touch with the state of the art in optical stuff--- I was somewhat surprised recently to learn of the existence and low cost of LR4 40gb optics. :P]
Take a look at: https://circuit-switching.sysnet.ucsd.edu/
And: https://arpa-e.energy.gov/sites/default/files/UCSD_Papen_ENL...
The current generation of switches that we're working on uses diffraction gratings patterned onto glass hard drive platters, installed in a modified hard drive, spun by a custom motor controller that's synchronized to the NICs via PTP.
The cost of switch ports and interconnects could all be dumped into making host interfaces faster, allowing for the switching time to be reduced.