Arm’s Neoverse V2
chipsandcheese.com
chipsandcheese.com
The C-model and RTL model outputs are often also compared with each other as a correctness validation step, as they should ideally never diverge. (ie, implement twice, by two teams, and cross-check the results).
Those simulations are terrifically slow for larger chips, so there is a surprisingly small number of workloads that can be run through them in reasonable time. So there tend to be even more simulator implementations that sacrifice perfect performance emulation for 'good enough' performance correlation (when surprises can happen). Being able to come up with a non-exact simulator that perf-correlates with real hardware is an art in itself.
First, they create detailed software models (usually in C++) of their chips to estimate performance as closely as they can before laying out a single transitory. These models can run code just like a real hardware device, albeit slowly.
Once the chip is designed, verilog simulators are programs used to generate the exact logical output of a circuit, which can be used to measure performance on a workload. However, this method is even slower than the first!
For larger workloads and higher speed, they use extraordinarily expensive FPGA-based platforms called Emulators. This allows circuits to be run at speeds in the MHz range before ever being sent to a fab. Booting an OS, running a complex multicore workload with shared memory, they can measure almost any workload. But this method is not available until late in the design phase and the boxes themselves are prohibitively expensive from being deployed very widely.
The software models are the most useful for estimating performance, as long as they are written early and well :)
How does this work? Do they model at the transistor level, or at the level of logical functions, or..? I'm particularly curious how this can estimate performance if it's anything higher-level than a direct transistor-for-transistor, layout-aware, emulation.
I'd be really interested in learning more if there's anything you could share, please. I can find info about chip design software and languages like Verilog (as you mention) but not this sort of modeling.
[1] https://en.wikipedia.org/wiki/Hardware_description_language [2] https://freecomputerbooks.com/langVHDLBooks.html
Let's take a simple example: Instead of modeling a 64-bit adder in all its gory transistor level detail, you can just have the model return the correct data after 1 "cycle" or whatever your ALU latency is. As long as that cycle latency is the same as the real hardware, you'll get an accurate performance number.
What's particularly useful about these models is they enable much easier and faster state space exploration to see how a circuit would perform, well before going ahead with the Verilog implementation, which relatively speaking can take circuit designers ages. "How much faster would my CPU be if it had a 20% larger register file" can be answered in a day or two before getting a circuit designer to go try and implement such a thing.
If you want an open source example, take a look at the gem5 project (https://www.gem5.org). It's not quite as sophisticated as the proprietary models used in industry, but it's a used widely in academia and open source hardware design and is a great place to start.
That's the behavioral model part and IRL they do basically the same thing to decide what behavior they actually want the hardware to do.
The next step is the circuit-level model done in verilog which actually simulates the logic-gates and does involve viewing a signal at every clock cycle.
"A Survey of Computer Architecture Simulation Techniques and Tools" - IEEE Access 2019 - Ayaz Akram, Lina Sawalha - https://ieeexplore.ieee.org/document/8718630
For more see also: https://github.com/MattPD/cpplinks/blob/master/comparch.md#e...
Usually System Verilog instead of C++ but it has C++ interfaces
In other words, they are a slightly more accurate version of something like QEMU, although I guess I should point out they can generate traces that can be fed into tools to model HW perf, ex gem5.
And idea about the cost and performance?
TBH this makes sense, as pretty much all the ARM code in the wild will be using NEON.
How can we be sure about this since Apple does not disclose any details about their chips?
> Apple's cores are actually completely custom, with no input from ARM the company.
Apple is one of the founders of ARM, with other two being Acorn Computers and VLSI Technology.
The ARM Cortex-X and Neoverse V cores are intended to compete with the Intel P-cores (like in Sapphire Rapids or the big cores of Raptor Lake) and with the AMD normal cores. These cores are optimized for high single-thread performance and for workloads where low latency is important.
The ARM Cortex-A5xx cores are much smaller and slower than any Intel or AMD cores.
I am hoping post ARM IPO we will have Cortex X5 and N3, V3 announcement. Also waiting if Apple A17 will gain any more double digit percentage IPC improvement. Personally I dont think that will happen.