OSs will need to be careful not to move a process from one core to another that supports a subset of the instruction set it was expecting.
It was a completely different situation, the Cell's SPE were vector coprocessors, they were a completely different architecture than the PPE but they were also interacted with explicitly from the PPE.
Also while they have different instructions sets they don't have a different arch, i.e. they have a "shared base" of instructions which likely are good enough for many background applications. Like updaters, downloaders, background mail fetching programs etc.
In the end it fits in with their "always connected" approach. Through I'm still somewhat skeptical about the "always connected" approach. Like even in many first world countries always connected is just not a think for many users. Even on phones it not fully given through much more then on laptops.
The program machine code can be scanned when loaded to look for specific assembly opcodes to determine the required capability of the CPU to execute it on. The code with instructions not fitted for Atom will be sent to the main CPU only.
Edit: Just a thought. The OS can install an illegal-opcode exception handler. When a process is first run on the small CPU, the unsupported opcode will raise the exception. The exception handler can simply set the processor affinity of the process to the main CPU and put it to sleep. The OS will handle it like it normally would - putting the process in the run queue of the affined processor.
That works as a heuristic, but it's not perfect, since JITs and self-modifying code are a thing.
I expect the chip will raise a fault and the OS will move the process.
[0] https://www.neotextus.net/hpca12.pdf
(I think the "3.2 QuickIA Software Support" section is interesting, if nearly a decade old by now)
Runtime detection would be the only actual concern here, but you can easily just advertise the common baseline. As in, just pretend the sunny cove core doesn't support AVX512. The only problem then becomes the big core is potentially slower than it could be, but given how rare things like AVX512 is in typical desktop applications will anyone actually care?
EDIT: Oh, and this appears to be what they're basically doing. Even though Sunny Cove itself supports AVX-512, it's being disabled in this application: "One thing we can confirm in advance – the Sunny Cove does not appear to be AVX-512 enabled."
Trap the undefined instructions, and dynamically replace them with a jump to emulation code if running on a little core.
Then let the OS and software folks do a systemwide profile to find out which bits of code most frequently run on which cores. Then configure the compiler not to output any instructions not available on the little cores for code which usually runs on the little cores.
This was more or less what VMware pioneered to deal with privileged instructions. I expect there are thickets of patents involved, even if the earliest have expired. But it's likely there is or would be some licensing agreement.
Disclosure: I work for VMware, though not in VMs per se. Speaking for myself only.
GPUs get away with it because they are really a completely different kind of processor, but even there GPU processing is under-utilized for this reason. It's really a pain.
I’d bet money that the different ISAs are full x86 and a subset of x86 which is a great idea. X86 has a lot of old instructions that almost No one uses. Perhaps the small cores will trap and move to process to the big core if it uses an old instruction.
I'd bet the instruction set difference is mainly level of SIMD support. Sunny Cove will have AVX2, Tremont won't have enough execution units to go wider than SSE2.
You don't even need to bet, it's in the article:
"Both SKUs will feature one big ‘Sunny Cove’ CPU core, along with four little ‘Tremont’ Atom CPU cores"
It's a different uArch, but the same ISA. The extension support would be the only concern here, which is super minor.
To clarify my concerns (I haven't dug in to the specific instruction set differences):
Assume the main core supports AVX2 and the smaller cores don't. Which core do you execute the code on? Which one will get you the best performance per watt? How do you account for that in the OS scheduler? What do you want to optimise for?
If your code is compiled for AVX2, it'll fail on the small cores unless it does continuous runtime checking (which is expensive, but given processes can migrate between cores, presumably necessary).
https://medium.com/@jaddr2line/a-big-little-problem-a-tale-o...
I don't see that as required? If you catch the illegal instruction signal that CPUs throw you can just run it on the other CPU since the instruction counter would not have incremented. There's a delay in the catch and retry on the other CPU but i don't see the big deal here?
The advantage is you avoid storing fp registers unless you are going to use them.
That flag could easily determine what you can run where.
Unless you think about embedded co-processors that are generally here to control a complex subsystem (like video processing or something like that, arguably GPUs fit in that description) but in general those aren't handled like real CPU cores at the OS level, they have dedicated drivers or userland libraries dedicated to a specific purpose.
Something tells me that if Intel wants this architecture to be popular they'll have to work on a tighter and more transparent integration, otherwise this is going to end up like the Cell.
It's a pretty interesting approach though, I'm genuinely curious to see how that's going to end up working.
If these Intel CPUs are meant to be general purpose I wonder how that's going to work out.