You've got the right general idea already :)
> I understand that ARM sells core designs, and that each company assembles them / designs them in the way suits their performance needs, and then gets them manufactured?
Most consumers of ARM cores are simply interested in integrating them into a larger design; usually an SoC. Not so much in tweaking performance, though certainly they'll choose whether to prioritize performance or power in their application.
> But what does that really mean? What is the "user" doing? Are they arranging them like kids' Lego blocks on a die surface? Dragging and dropping blocks of code?
Pretty much like Lego blocks. Most shops just want an SoC that does X, Y, and Z with Foo requirements. So they grab an ARM core, an HDMI 2.0 RX core, and H.265 core, and glue them together.
Depending on what tools they're using this "gluing" is specified in different ways. You can design it somewhat abstract in a block diagram, where you specify all the cores you want, specify what (virtual) pins from each core connect to what others, and maybe tweak a few parameters on some of the cores. It looks kinda like this: https://www.altera.com/content/dam/altera-www/global/en_US/i...
Or there are even higher level tools that let you specify and hook these things together in a GUI specifically designed for building designs like this. They look like this: https://i.stack.imgur.com/yud23.png
But those and other higher level tools all basically just have a compiler that converts specifications into code (Verilog or VHDL).
That is then handed to another compiler which creates a netlist. Think of this like assembly code but for hardware. It specifies the whole design at the level of logical operations. 2-bit AND here, 4-bit Full Adder there, etc. Finally that netlist is thrown through _another_ compiler which does final place and route. Place and route is where the netlist is converted into the transistors and wires for the actual die, and then rendered out into all the layer masks that will be sent to the fab. (I'm glossing over a few details here. E.g. place and route is actually working from a library of transistor designs for each possible logical gate, pre-designed by the silicon fab that your going to send your masks to.). It sounds simple, but until place and route your design is just an abstract spaghetti of logical operations and connections between. Place and route has to solve an NP-Complete optimization to figure out where, on the physical die, all the transistors are going to go, given a set of constraints (transistors need to be close enough to their neighbors to achieve the performance requirements).
Anyway, stepping back, the design phase is where you specify various parameters for the cores. These parameters typically change things like enabling/disabling features of the core. For example on an ARM core you might disable the math co-processor to save space/power if you don't need it; make the pipelines smaller, etc. These are configuration options; you aren't editing the ARM core's code. The core itself specifies how its code changes depending on configuration options.
The place and router phase is where you have some control over power/performance. You can tell the tools to focus on power if your design needs to be low power. It'll then design the transistors in such a way that it'll use less power but also lose some performance. Or the opposite.
Of course how the cores themselves are designed is important to power versus performance, and ARM cores probably have some configuration options that affect their behavior in that regard.
And then the silicon you target is of the utmost importance for design trade offs. Targeting 7nm fabs versus something larger will usually mean a more power efficient and performant design, but will have _much_ higher up-front costs (the physical masks you deliver to the fab cost tens of millions of dollars or more).
To be clear, this is the level at which _most_ consumers of ARM cores are working at. They're just putting Lego blocks together. The ARM core is just a black box in that regard. ARM delivers the core's "code" to you as a pre-processed netlist. You can't see its real code (you just see a chaotic spaghetti of logical operations). But some companies have more special needs and want to tweak the core in specific ways (e.g. Apple). They'll have special deals with ARM that give them access to the core's source code where they can make custom tweaks. But this is rare.
Some companies have their own IP which they might integrate into a design. Their own cores, coded from scratch. These are coded in Verilog or VHDL, for the most part.
> Like, why doesn't ARM just create the chips that the end OEMs want?
ARM cores get used in TONS of custom silicon. Wifi routers, drones, cell phone chips, cell phone coprocessor chips, etc. So part of it is that ARM just couldn't possibly design and build all these different kinds of chips.
It's also that ARM does what ARM does best: design ARM cores. That's a job that takes an entire company all to itself to accomplish. Anything else is just beyond the scope of their company (for now).