ExaLink Fusion – Ultra low latency switch
exablaze.com
exablaze.com
"Consistent latency as low as 350ns for all packet sizes" (http://www.arista.com/en/products/7150-series)
So this is about 3x faster than current devices.
Calling this the fastest switch in the world and then printing 5ns is misleading. It is 5ns as a Layer 1 patch panel. Not really what you want for market data which needs multicast, PIM, IGMP snooping, BPG, ACL, etc.. For the 110s is this multicast with all features enabled? Let me know if I missed a link.
BTW, if you look back when the 7124S was released there were others that built a switch based on the Bali chip like BNT. An important point to note is that the chip is line rate multicast but that is not enough. Processing the joins/leaves and programing the chip is a function of the software and that depends on the quality of the code and the CPU system in the switch. Arista won because of this. Here is a link to a bake off from 2010.
http://www.networkworld.com/article/2241525/virtualization/a...
The technology involved in enabling 64 octet frames to be shoved around in 300ns is fascinating. The software control of these systems was just as fascinating. These chips had a high level of programmability in the frame handler. How to use that programmability was an open question when I left Fulcrum.
One very interesting bug, but one that did not really matter, was handing 1G at line rate. For those that do not know a bit of a background.
So you have a 10G chip that connects to a PHY chip.
http://en.wikipedia.org/wiki/SerDes
There are 4 x 3.25G lanes to the PHY. Now lets say you want to put a 1G SFP in that port. Well the path uses only one of the 3.25G lanes, not all 4. All the of logic and timing is really built around the packets going over the 4 links.
So anyway, if you ran a test at 100% line rate with a 1G optic there would be drops if your packet size was not divisible by 4. The stand packets sizes that are commonly used for RFC2544 testing are 64,128,256,512,1024,1280,1518. So all the sizes work expect for 1518. You could do 10G at line rate @ 1518 packet size but not 1G! It is very important for everyone to understand that this only matters in lab testing and has ZERO impact in a real world environment. If you changed the IFG on the link it would run at line rate.
http://en.wikipedia.org/wiki/Interpacket_gap
The Fulcurm chips were amazing. I very much enjoyed using them.
The claim in the (original) story title is 110ns layer 2+ switching. I'm sitting in the summit right now, no details on what 2+ means, but, the target market is market data distribution (and this was discussed in the slides) so I would assume that they've thought about this.
5ns is additional functionality available in previous products (http://exablaze.com/exalink-50) which is also in this product.
From the sides, the idea is that they can set up multiple broadcast groups with the 5ns device to distribute market data, and then aggregate responses together with with 110ns (or in some cases 100ns) device, all in one box.
So the comment "can setup multiple broadcast groups" sounds like the Cisco Nexus 3548 Warp Mode. The market did not take to that so not sure how this is going to change that. Like I said, I have not reviewed all the data yet.
An important step for verification for all this is to put it in the hands of David Newman for a test, either paid via his private company Network Test, or publicly via Network World. A testing house like EANTC would also be great.
As you can might have guessed from my post I know a bit about this world (I am not involved anymore). At this point it is more of a historical curiosity on my part more then anything (I still have a number of friends that run HFT networks at the banks). Given what I saw in testing and claims during the early HFT rush I always want to see 3rd party testing from a reputable testing house showing a real market data feature set being used.
Here is a starting list:
1. BGP running 2. Mulitcast w/PIM 3. IGMP snooping 4. ACL configured 5. 1 port to max port of switch fanout. 6. 1% to 100% line rate on the feed. 7. Packet size ranges, both fixed and mixed. 8. Join/leave time. 9. Max groups joined.
Can I ask why you'd want BGP running in a private LAN? This seems like an odd thing to do, especially in Colo environments.
Since everyone has the same latency from the feed, the goal is to take the feed in and fanout to host as fast as possible. So ideally one port on this switch is to the exchange and the others direct to servers or if you need more scale to another lay of low latency switches that the server farms connect to. For all the money spent on this it is really pretty simple from a networking point of view.
Am I missing something here? This thing is fast when used as a patch panel, but that is cheating and isn't really switching.
Which means simple apps wont even have their traffic leave the switch. I can think of a ton of uses for that, especially from a security perspective.
The Solarflare device had (has?) the FPGA in a strange place, behind the NIC controller. AFAIK, the Exablaze NIC is pure FPGA which again saves on latency.
exablaze might not even be able to keep their ip as it remains subject to dispute: http://meanderful.blogspot.com/p/exablaze-and-zomojo.html
The bio you submitted to STAC says you worked at Exablaze: "He is an experienced software and hardware developer and has worked for a collection of start-ups including Exablaze,Zomojo..."
https://stacresearch.com/system/files/summit/files/stac_summ...
The Exablaze announcement was made today at the London STAC summit at which I was invited to speak on unrelated work.
I am quite surprised that no one mentioned the Cisco nexus 3548. Switching at L2 with 110ns latency is not that impressive considering that the Csico nexus 3548 switches packet at (L2/L3) with 50ns (with warp span turned on) http://www.cisco.com/c/en/us/products/switches/nexus-3548-sw...
110ns is the device in its capacity as a full layer 2(+) switch. The "warp span" equivalent would be using the layer broadcast groups functionality which runs at 5ns, 10x faster.
Has that advantage of having quicker rdma than 10gig ethernet. so while the switching may be slower, the processing is faster because data is piped directly into memory.
and you can see from the product brief (https://www.vitesse.com/products/download.php?fid=4548&numbe...) many people use it in telco oriented gear for back-planes and the like.
Here is my take on why you shouldn't buy an ExaNIC from Exablaze (http://meanderful.blogspot.com/2014/11/dont-buy-exanic-from-...)
Just remember that I'm very biased...
--Matt.