A Peek Inside a 400G Cisco Network Chip
nextplatform.com
nextplatform.com
Reading between the lines a bit it might have been used in one of the Nexus 5K's - which would put it at around 8 years old, depending on how we're counting.
The point is that 12.5Gb/s SerDes pretty much means they'd be practically to 10/40G. At least in the DC networking world this puts it several generations back.
The original 5k (which was a Nuova Systems product before cisco spun them in), had a 1040G crossbar (nicknamed Altos) and a bunch of 80G line chips (nicknamed Gatos).
https://www.intel.com/content/dam/www/public/us/en/documents...
I think that a comparison to the FM4000, which is the programmable series of parts in the "Monaco" family, would be more fair. Here's their datasheet: https://www.intel.com/content/dam/www/public/us/en/documents...
The FocalPoint ASICs were, as far as I know, some of the first to support (a demo-quality implementation of) OpenFlow in hardware. When Intel bought them, they released the datasheets, which is neat.
As a real-world example, these ASICs were used in Arista's 7100 series (c. 2008) switches. They published a two-part "Technical Evaluation Guide" for those switches which are, among other things, an interesting glance at how switches are constructed out of ASICs. Part 1 (https://local.com.ua/forum/index.php?app=core&module=attach&...) shows the topology of each switch (starting on page 13).
The 7124 is a single 24-port FM4224 with all 24 ports connected to front-panel ports. The 7148S has three FM4224 ASICs; each is connected to 16 front-panel ports and uses 4 ports (40 Gb/s) to connect to each of the other two ASICs in a ring. Intuitively, this means that it's possible that those inter-ASIC connections could cause bottlenecks (if e.g. all 16 ports connected to the first ASIC try to send 160 Gbit/s of traffic to the 16 connected to the second ASIC, they'll saturate the 40 Gbit/s of connectivity between the ASICs). Therefore, Arista also offered the 7148SX, which is non-blocking but needs six (!) FM4224s to make it happen!
There are field trials and pre-prod switch/route chips and optical/DWDM gear that support 400G Ethernet using 56Gbps SerDes. This chip uses older 12.5Gbps SerDes, current-gen 100Gbps stuff uses 28Gbps SerDes.
For anything over 100G (currently) you're making some sacrifices on distance and may even need some guard bands.
(Short-haul/intra-DC may well be a different story but I doubt that it's very different.)
Unless you are pushing google/facebook/Comcast level traffic - there are very few use cases. Apparently, Google/Facebook uses their own network hardware.
As a result it's actually fairly common to find new DC fabrics (read: inter-switch connections, not end hosts) being built with 100G because there's no significant economic disadvantage to doing so. That said, the pricing for inter-site 100G is still high enough that it hasn't commonly made its way to smaller organizations.
I think you're having the same conceptual issue that I had when first reading it.
The 400G in the article title is referring to the bandwidth of the chip, not the bandwidth of any particular port.
The ports, as mentioned elsewhere here, were probably either 10G or perhaps 40G.
Very cool.
L2 instruction cache that also has an on-chip interconnect that links the clusters and caches to each other as well as packet storage, accelerators, on-chip memories, and DRAM controllers together. This interconnect runs at 1 GHz and has more than 9 Tb/sec of aggregate bandwidth
I keep seeing elephant guns and experimental German artillery when I read that!
https://news.ycombinator.com/item?id=15244655
I wish I could set up a notification of some kind to know when someone gets Erlang/Elixir running on this chip. It would be a great platform for stress-testing Go concurrency as well.
Beyond that, I would really like to see Octave running on it because it's the only approachable vector programming language that I know of. The holy grail for me is to be able to use something like the MATLAB libraries at 1000 times their current speed to simulate the interesting stuff.
I spent my teens writing blitters for shareware games and found that even then, the cache mostly got in the way. Processors like the PowerPC 603e had a pretty substantial cache miss penalty that was on the order of 5-20% for me depending on the situation. It was difficult to come up with appropriate cache hints for even relatively minor random access. I tried disabling the cache, but that made it even slower than a 601. So that's where my head is at, and the Epiphany sounds perfect. Here's a quick link for anyone curious:
https://www.parallella.org/2016/10/05/epiphany-v-a-1024-core...