Intel Discloses Lakefield CPUs Specifications
anandtech.com
anandtech.com
Are the little cores’ instructions at least a complete subset of the big cores’? Are we going to have some ridiculous situations where the little cores are completely pegged but the OS can’t migrate their processes off to the big core?
Or do kernel programmers need to start chasing Intel and trap/software implement every single future AVX8192AESNISSSE instruction Intel jams into future instruction sets to provide Xeon market differentiation?
The only complication here would be if they have differing extensions like AVX512, but that's easily solved by the OS by just advertising the common baseline. Nothing about this looks difficult to support?
What does that mean? Advertise to whom? The process/process loader? Does it mean that I can’t compile with -mavx2 anymore? What if I do?
The extensions are the whole problem.
Runtime detection is the process querying what extensions are available, and then selectively using those. You adjust what the query returns to only return the common set.
Runtime detection has been a pretty standard thing for well over a decade now - it's how we all manage to run the same compiled binaries over the years despite variability in SSE & AVX support. You don't download different versions of Chrome/Photoshop/Gimp/Premiere/Blender/Whatever compiled for different CPU micro-architectures, do you? You might if you run Gentoo I suppose, but that'd be about it.
> Does it mean that I can’t compile with -mavx2 anymore?
You already can't if you're shipping binaries to users unless you only support Skylake & newer? There's a lot of CPUs currently in use that don't support AVX2. So... you either already have this problem and you're familiar with it, or you're not doing this and it's moot.
Edit: Also: AVX2 is a lot older than Skylake. You're probably thinking of AVX512.
But it looks like Intel is doing this anyway, as Sunny Cove in this application has had its AVX-512 removed anyway: "One thing we can confirm in advance – the Sunny Cove does not appear to be AVX-512 enabled."
But this shoots the big core in the foot. You run at the lowest common denominator and the big core doesn't have the advantage of higher clocks. And this might just lead to a lot of software that only runs on the big core.
Imagine you wanted to fly around between multiple points but if you want the option to ever switch to a bus then your plane will circle around each airport until it's as slow as the bus. Or you can opt for "plane only".
A better mix would have been simply using lower clocked, lower powered cores of the same type or really close derivatives of the big core where the manufacturing process and clocks are what keep power low. Not a mix of Ice Lake and Atom. But right now Intel would throw everything at the wall to see what sticks.
And it seems like a good way for developers to make sure their software stays on the big core.
AVX2/AVX-512 is great in HPC workloads, for example. But nobody is running an HPC workload on a 7w netbook, now are they?
What is useful on Xeon and what is useful on Atom are different. This is an Atom-class SoC used in Atom-class applications, not a Xeon-class one.
> You already can't if you're shipping binaries to users unless you only support Skylake & newer?
Huh? My 6+ years old gaming PC supports AVX2, and it definitely doesn't have a "Skylake & newer" CPU!
AVX2 support started at Haswell, or Intel core 3rd generation. We're at gen 10 now.
While there certainly are still a lot of systems without AVX2 support, new games (and other performance hungry software) requiring it would not be completely unreasonable.
(I recall a recent news story about how unawareness of this among people writing or documenting compilers has started to cause problems.)
> (I recall a recent news story about how unawareness of this among people writing or documenting compilers has started to cause problems.)
You can't blame them!
Except that they almost always do this runtime detection once, on startup, and then choose/thunk codepaths accordingly. If the OS just happens to start my avx2 process on a little core (and how is it going to know better?), that's going to turn off all of my optimizations, regardless of where the process subsequently gets migrated to.
You already can't if you're shipping binaries to users unless you only support Skylake & newer? There's a lot of CPUs currently in use that don't support AVX2. So... you either already have this problem and you're familiar with it, or you're not doing this and it's moot.
Except nobody in 30 years of x86 dev expects to get a different answer from CPUID during runtime.
That's partly what they did. From the article: "One thing we can confirm in advance – the Sunny Cove does not appear to be AVX-512 enabled."
Maybe they also fused off AVX & AVX2 support in the Sunny Cove core as well, we'll see.
And disabling AVX in cores that otherwise support it is already a common thing - see the Pentium & Celeron lineups that Intel currently sells. They don't have AVX/AVX2, even though the cores inside them definitely could offer it.
Which they do, Atom tops out at SSE4.2.
All that can be done with SSE and probably with x87 instructions, but, still, software will try to pick the best option.
If a binary does that, the OS can just not migrate that binary across cores.
The only mechanism I can come up with is to detect illegal instruction traps on small cores, then flag the thread as big-core-only and re-start the execution at the bad instruction. That's not ideal but maybe workable.
(If it traps on the big core, too, it's just a bad instruction and SIGILL is raised to userspace like usual.)
SSE itself tops out at 4.2, but Tremont does support the newer SHA extensions.
AVX appears to be the only missing thing, but AVX isn't even standard across Intel's other lines, either. The Pentium line doesn't support AVX either, for example, even though they are using Skylake & newer micro-architectures.
So you currently can't assume AVX support, and you still won't be able to assume AVX support. Why does this matter?
Operating systems move threads between cores. If those cores support different features, threads that are migrated to low-feature cores might experience illegal instruction traps despite correctly checking for instruction features.
For this to work, there would need to be a small set of undemanding tasks that account for a lot of machine time. Feeding video to the GPU? All sorts of GUI compositing and housekeeping? Handling network connections in the browser?
I don't think this is a good explanation, but it's fun to think about.
Even just looking at a browser, I might have half a dozen generic tab processes open, each using a small amount of CPU. But then I navigate one to a game, and I want that particular tab to get near-exclusive access to the big core while the others use only the little cores.
But what happens when the user opens up a photo viewer app and suddenly wants those photos to be synced right now?
If your code is running on a recent iPhone -- the heterogeneous-core platform I'm most familiar with -- then the answer is that the kernel will immediately detect the priority inversion when a foreground process does an IPC syscall, bump up the priority of the no-longer-background process, probably migrate it to the fastest core available, and make it run ASAP. Then, once the process no longer has foreground work to do, it can go back to more power-efficient scheduling.
This kind of pattern is super common, and it would be way more annoying and perilous to try to split tasks into always-foreground and always-background.
1. https://www.mono-project.com/news/2016/09/12/arm64-icache/
> The OS can install an illegal-opcode exception handler. When a process is first run on the small CPU, the unsupported opcode will raise the exception. The exception handler can simply set the processor affinity of the process to the main CPU and put it to sleep. The OS will handle it like it normally would - putting the process in the run queue of the affined processor.
And generally speaking, when an application first starts issuing SIMD instructions, that's probably not the a great time to be interrupting it even if it only needs to happen once.
the issue is that once you have an active process using a feature only available on the larger cores, you can't shut off the larger cores to save power without paying a large latency to wake up that process.
Yes that's the idea.
> Which would be slow I would assume.
How expensive do you think a trap is? It takes about the order of 10 billionths of a second.
> If you never switched back, wouldn't any process using advanced vectorized instructions (like anything using a decent libc) be permanently pinned to the large core?
I think you can switch back next time you schedule.
Ok, yes, then we're on the same page. I would still think that would be slow? You'd need a full transition-to-kernel and context switch before you could execute again, which AFAIK would take at least microseconds…unless you think there would be a faster path to resume execution?
No that's around 30 ns on modern hardware I believe.
are you sure about that? I would expect at least a couple of orders of magnitude more just for the userspace->kernel transition.
edit: for what is worth, a syscall it takes 250ns on my (admittedly vintage) machine. That's using the lowlatency sysenter path. An interrupt is probably going to cost more.
Anyway the cost of scheduling on another core is going to dwarf that.
edit2: for reference, this was a Sandy Bridge turboing at 3.5 Ghz during the test. With spectre mitigations on (which is going to be a good chunk of that overhead).
At least this only has to be done once, and frankly these AVX instructions are slow initially anyway.
FWIW I believe only some of them, and those are just some (or all?) AVX-512 instructions. I think AVX2 is implemented on the main die and doesn't have to be powered up first.
AVX (256 bit) instructions also suffer a penalty. Both the 256 and 512 bit instructions resulted in a slowdown for 9 microseconds featuring a quarter the instructions per clock, but the 512 bit instructions resulted in an additional penalty of 11 microseconds without executing instructions.
The first penalty was associated with voltage, and the second with frequency. Heavier 256 bit instructions would probably have resulted in the frequency transaction as well.
[1] well, I hope the larger cores support a superset of the smaller cores, otherwise it would really be insane.
Linux already tracks how long ago a task used AVX-512. I assume the same mechanism could be used to track AVX as well.
https://git.kernel.org/pub/scm/linux/kernel/git/tip/tip.git/...
The only way to make big.LITTLE work and allow task migration between them is to have them support the same instruction set exactly. So, I suspect that that is why AVX512 was removed from the big core here, and I suspect the atom cores aren't exactly garden variety either. They probably support most other extensions of modern x86 CPUs, at a slow speed. For example, AVX2 can be microcoded by running multiple pieces of the data through the ALU, one at a time.
I'm not a programmer, but doesn't compiled code often offer multiple codepaths, checking flags at runtime? So if your code runs on the big core, it finds avx256 is supported, and uses the avx256 codepath. If it runs on the little core, it finds avx256 is not supported and it takes a legacy floating point codepath instead.
It depends a lot on which instructions differ and how the feature flags work, of course. They might have set it up in a way that's fine.
[1]: `target_clones` -> https://gcc.gnu.org/onlinedocs/gcc/Common-Function-Attribute...
I don’t think users would be happy if it were a lottery whether their code runs slow, using microcode instructions, or fast, on the better CPU, especially given that, I guess, the difference in instruction set is for vector instructions (because adding those is what’s makes a CPU big nowadays)
So, performance oriented code will pin its performance-critical threads to the fast CPU. Now, the question is: what code won’t try to claim the faster CPU, given that benchmarks typically run programs while there is no contention from other programs?
Do apps have CCX pinning logic for Ryzen? No? Why would they add it for an Intel processor that probably won't even be sold in volume.
If the microcoded vector ops cause your code to run slow, then you will be consuming lots of CPU time, and should be migrated to the big core anyway, right?
https://ark.intel.com/content/www/us/en/ark/products/202777/...
There's no mixed instructions as Anandtech speculated. The Sunny Cove core was cut down to match what the Tremont cores support. So no AVX at all, no weird extension mismatch, no OS headaches beyond the expected big.LITTLE headaches.
Anecdotically, I have a "netbook" that has a Goldmont+ Celeron N4000 that works very respectable for everyday business (web browsing, office suites, watching videos, etc) but crawls when trying to run scientific code. 100x slowdowns at worst, 10x slowdowns on the parts that take "nice and easy" (wrt a 7200u notebook that I usually carry around).
You mean not much. The question was apart from scientific computing how much does it hurt. Per your own comment it's only worth maybe 20%, and that everyday needs work fine.
It'll show up occasionally in some things that would be relevant to a device with this SoC, like noise cancellation, but very little else. And even then you can do noise cancellation even better on a GPU (RTX Voice says hi), and this does still have a GPU and a decent one at that (~500 gflops), so it's not even that simple.
To clarify my concerns (I haven't dug in to the specific instruction set differences):
Assume the main core supports AVX2 and the smaller cores don't. Which core do you execute the code on? Which one will get you the best performance per watt? How do you account for that in the OS scheduler? What do you want to optimise for?
If your code is compiled for AVX2, it'll fail on the small cores unless it does continuous runtime checking (which is expensive, but given processes can migrate between cores, presumably necessary).
https://medium.com/@jaddr2line/a-big-little-problem-a-tale-o...
I don't see that as required? If you catch the illegal instruction signal that CPUs throw you can just run it on the other CPU since the instruction counter would not have incremented. There's a delay in the catch and retry on the other CPU but i don't see the big deal here?
The advantage is you avoid storing fp registers unless you are going to use them.
That flag could easily determine what you can run where.
Unless you think about embedded co-processors that are generally here to control a complex subsystem (like video processing or something like that, arguably GPUs fit in that description) but in general those aren't handled like real CPU cores at the OS level, they have dedicated drivers or userland libraries dedicated to a specific purpose.
Something tells me that if Intel wants this architecture to be popular they'll have to work on a tighter and more transparent integration, otherwise this is going to end up like the Cell.
It's a pretty interesting approach though, I'm genuinely curious to see how that's going to end up working.
If these Intel CPUs are meant to be general purpose I wonder how that's going to work out.
Trap the undefined instructions, and dynamically replace them with a jump to emulation code if running on a little core.
Then let the OS and software folks do a systemwide profile to find out which bits of code most frequently run on which cores. Then configure the compiler not to output any instructions not available on the little cores for code which usually runs on the little cores.
This was more or less what VMware pioneered to deal with privileged instructions. I expect there are thickets of patents involved, even if the earliest have expired. But it's likely there is or would be some licensing agreement.
Disclosure: I work for VMware, though not in VMs per se. Speaking for myself only.
I’d bet money that the different ISAs are full x86 and a subset of x86 which is a great idea. X86 has a lot of old instructions that almost No one uses. Perhaps the small cores will trap and move to process to the big core if it uses an old instruction.
I'd bet the instruction set difference is mainly level of SIMD support. Sunny Cove will have AVX2, Tremont won't have enough execution units to go wider than SSE2.
You don't even need to bet, it's in the article:
"Both SKUs will feature one big ‘Sunny Cove’ CPU core, along with four little ‘Tremont’ Atom CPU cores"
It's a different uArch, but the same ISA. The extension support would be the only concern here, which is super minor.
OSs will need to be careful not to move a process from one core to another that supports a subset of the instruction set it was expecting.
It was a completely different situation, the Cell's SPE were vector coprocessors, they were a completely different architecture than the PPE but they were also interacted with explicitly from the PPE.
The program machine code can be scanned when loaded to look for specific assembly opcodes to determine the required capability of the CPU to execute it on. The code with instructions not fitted for Atom will be sent to the main CPU only.
Edit: Just a thought. The OS can install an illegal-opcode exception handler. When a process is first run on the small CPU, the unsupported opcode will raise the exception. The exception handler can simply set the processor affinity of the process to the main CPU and put it to sleep. The OS will handle it like it normally would - putting the process in the run queue of the affined processor.
That works as a heuristic, but it's not perfect, since JITs and self-modifying code are a thing.
I expect the chip will raise a fault and the OS will move the process.
[0] https://www.neotextus.net/hpca12.pdf
(I think the "3.2 QuickIA Software Support" section is interesting, if nearly a decade old by now)
Runtime detection would be the only actual concern here, but you can easily just advertise the common baseline. As in, just pretend the sunny cove core doesn't support AVX512. The only problem then becomes the big core is potentially slower than it could be, but given how rare things like AVX512 is in typical desktop applications will anyone actually care?
EDIT: Oh, and this appears to be what they're basically doing. Even though Sunny Cove itself supports AVX-512, it's being disabled in this application: "One thing we can confirm in advance – the Sunny Cove does not appear to be AVX-512 enabled."
GPUs get away with it because they are really a completely different kind of processor, but even there GPU processing is under-utilized for this reason. It's really a pain.
Also while they have different instructions sets they don't have a different arch, i.e. they have a "shared base" of instructions which likely are good enough for many background applications. Like updaters, downloaders, background mail fetching programs etc.
In the end it fits in with their "always connected" approach. Through I'm still somewhat skeptical about the "always connected" approach. Like even in many first world countries always connected is just not a think for many users. Even on phones it not fully given through much more then on laptops.
Intel can’t meet deadlines which forces Apple to restructure their plans constantly.
We are on 14nm+++++ or something?
In 2008 Apple bought PA Semi to design processors for iOS, make processors for Macs is likely something they've been working on since way before Intel was late on 10nm.
It always seems like a good idea to do your own thing when your supplier is stumbling. The problem then is you're on your own. If they find their footing you're in trouble -- either you've blown a huge pile of money developing something which you then don't use because the competition has something better, or you use it anyway and get to relive the final days of Sun Microsystems.
And the same thing happens if anybody can beat you. If Intel can't but AMD does, you lose. If AMD can't but Qualcomm does, you lose.
Worse, success precipitates failure. If you make an in-house processor which is only a couple of percent faster, that's boring. You spend a lot for a little. If you make one which is more than a couple of percent faster, that's war. Intel can't have that. Google can't have that. Samsung can't have that. Microsoft can't have that. Even if you're bigger than any of them, you're not bigger than all of them. So they double their R&D, combine their resources, whatever it takes, and soon you're Sun Microsystems again.
https://www.gsmarena.com/benchmark-test.php3
I see a lot of Android devices at the top of those lists.
Apple likes to cherry pick. For example, they require everything on iOS to use their browser engine, then they spend a lot of time optimizing their browser engine for their CPUs. But that's not superior hardware performance, it's software optimization, which they could do with whichever commodity CPU they chose as well.
Intel does the same thing. They spend resources optimizing popular software for their CPUs. Qualcomm not so much.
it's not apple simply being control freaks (although they are and can afford to be), it's intel shooting it's own foot.
Yeah, Intel's chips are not the biggest fish in the Apple fish fry.
I'm not even an Apple customer, I'm just forced to use them by my employer. (And it feels like most employers nowadays.)
[1]: https://www.theverge.com/2019/3/5/18251264/macbook-pro-2018-...
[2]: https://support.apple.com/15-inch-macbook-pro-battery-recall
[3]: https://support.apple.com/13inch-macbookpro-battery-replacem...
[4]: https://thenextweb.com/apple/2016/10/27/cant-connect-new-mac...
[5]: https://techcrunch.com/2018/09/01/an-ode-to-apples-awful-mac...
[6]: https://www.vox.com/the-goods/2019/7/3/18761691/right-to-rep...
¹I now limit my use of Apple keyboards to an absolute minimum.
While Intel showed this, Apple was currently selling iPhone 11s with processors that would clearly outpace this. The A13 has two large cores and four small cores, whereas either one of the large cores outperform the comparable 'Core' core in this, and I'm betting the smaller cores do as well, possibly at less power. Supposedly their ARM< Mac CPU is a 12 core architecture, I imagine with four large cores and 8 small cores. Thats a massive hardware difference...did I mention that the 12 core CPU is supposed to be build on TSMC's 5nm? Intel is trying to get current and honestly, maybe Apple will ship a Macbook Air with it.. but probably not.
The evidence of this is flimsy at best. It's possible the A13's big cores are faster than Intel's Sunny Cove, but also not very likely, and certainly not established fact. The only real evidence here is Geekbench, which is a highly questionable benchmark that also has extreme OS dependencies (compare results for the same CPU on Windows & Linux, for example).
When/if an Apple releases something running MacOS with their custom ARM cores then you'd finally get a good comparison to isolate the CPU's performance by itself.
We already knew what to expect from Intel Willow Cove, we will have to wait and see what A14 has in store for us.
So not as fast as the latest processors on a per core basis, but still pretty damn fast.
Can't really make assumptions, but you would think with the higher thermal envelope Apple can boost clock speeds making them even more competitive.
OS and compiler dependencies, really. On the other hand, comparing macOS and iOS results may well be more meaningful (they use the same compiler and much of the OS is shared, especially at layers likely to significantly affect performance).
Geekbench (v4 especially) is definitely flawed, but the results in this case don't seem out-of-line with what we see on specINT for example.
e.g., https://browser.geekbench.com/v4/cpu/compare/15541405?baseli... comparing shows an A12X roughly matching a Sunny Cove 1060NG7 on single-thread performance (and it mostly follows where I'd expect Intel's strengths to be, with a wider vector unit and specialized instructions); in multi-thread performance the A12X unsurprisingly wins, but it also has double the number cores (okay, four of them are little, but that's still providing extra capacity to get work done).
Lakefield, Foveros, 10nm were suppose to arrive in 2018. And Apple would likely have known about this roadmap before that time.
Looks like there will be lots to unpack in WWDC, unfortunately it wont be a staged keynote like it used to.
*Edit: I wanted to add that I left about 2 years ago, and thought the project had been cancelled, along with so many other 10nm products.
Why left? If you dont mind sharing :)
Yes, BK was fired right around the time I gave my notice.
> Why left?
A lot of reasons that would sound like griping. I was recruited by Google for a similar job, but with a chance to build something from the ground up. That sounded like the adventure I was looking for, so I took the offer. I must say, it's been an adventure!
Thanks. I guess that does sums it up and align with lots of other similar sentiments I heard over the years.
Off Topic: I think we live in a world where we put too much emphasis on the positives and neglect any negativities. So we end up having any valid complains to be viewed as moaning or griping. I wish as a society there are better ways to handle these rather than simply ignoring it, which would leads to all sort of bad things happening as we seen in the world today.
I'm going to switch from a Pixel to an iPhone this year. The CPU is just clearly better, and seems to actually matter for taking photos and web browsing. (The Pixel still takes like 5 seconds to post-process a photo. iPhones do it instantly.)
Those "tests" almost invariably suck at measuring anything useful or being accurate.
The main area where normal people notice this is in web usage, where Mobile Safari has handily outpaced Android browsing for many years — see e.g. https://discuss.emberjs.com/t/why-was-ember-3x-5x-slower-on-... from 2014. How much that matters depends on how much a particular website is limited by single-core JavaScript performance — well-engineered sites probably don't have a huge impact but it's quite noticeable on anything which has a bloated SPA and the web has been moving in the latter direction for years.
Gaming is the other area where this is fairly noticeable but that varies both in where the bottlenecks are (CPU vs. GPU), how prominent the effect is, and the relative quality of the ports so it's harder to do a fair comparison.
I appreciate your effort with giving more background, although I think 2014 is a bit dated with Firefox Quantumn becoming more of a thing on Android. Can't remember if the Preview has it yet or not.
It also includes :AI:
The Pixel 4 has renamed it the "Pixel Neural Core" and has given it even more machine learning tasks and offloading such as Google Assistant and face unlock.
I just wanted a touch ID, notcheless, fast and current phone, and Apple did not offer one until recently.
I am not going back, and I am also now seeing the same thing with their laptops. Crappy keyboards, non optional and useless touch bars, low performance per dollar and software that gets in your way. No thanks. Windows 10, WSL, Surface Book/Go/Alienware, wow. Just good, open, usable stuff.
Literally phoning home every 2 mins:
Amd can do lower tdp to with underclocking undervolting. But should offer something out of the box.
Happy cpu makers are lowering power consumption we have global warming. People often talk about co2 with cars but less so with computers.
Cheat sheet to reading Intel CPU codes: if the number starts with 10 and contains a G it's an Ice Lake. Otherwise, if there is an Y somewhere, it's ultra low power. The last U letter means 15W, rarely 28W. Last letter T means a 35W desktop chip. E means embedded. Letter H means 45W mobile. All of these were Core chips, first letter N, J, Z means Atom. First digit (or two digits for 10) is generation which became an absolute mess past 7th gen, most important change is that 8-U is quad core where 7-U was dual core.
The other questions are perf/gram and perf/W (the latter matters more for temperature management if you’ve rolled batteries into the other two metrics)
They do say they’ve improved idle wattage by an order of magnitude. That’s good. (Imagine a laptop that stays on with the screen off for a week. Some idle cell phones can do that.)
But there are things that need lower power more than multi-thread performance.
It's also not obvious what that use case would be. You can put most laptops into standby and they'll run on battery like that for many weeks. What's the thing that needs more than that?
It would be interesting to compare idle power consumption, but for that we'd have to know what it actually is.
Then it doesn't matter what the standby power consumption is - it simply matters how quickly it can get back into a working state from OFF.
Wiki: https://en.wikipedia.org/wiki/List_of_AMD_accelerated_proces...
CPU-World: http://www.cpu-world.com/CPUs/Zen/AMD-Ryzen%20Embedded%20R10...
However those are Zen+ fabricated on 14nm, so not really the best that AMD could do in this area. They could fight either with a slowed down Renoir (that could meet the TDP requirements while not requiring any development) or a specific processor based on Zen2.
Also the "stacking" approach is really interesting, I wonder how far that approach may go though-- heat dissipation seems like it would quickly become a problem with multiple layers.
For power efficiency for dynamic loads, an asynchronous circuit design would win over current solutions. However designing asynchronous circuits is a magnitude more complex than a synchronous one.
But for parts like AVX, extensions that will only be available upon the main core, that would yield dividends and with that. However, may find they offer such extensions as a separate chiplet/stack and negate the issue of does this core support it as they all could tap into that.
What i'd find interesting would be the actual design and instruction set under the hood, the x86 gets translated via microcode and do wonder how much of the underlying way things work has diverged from that original instruction set.
And there's still benefits to lower speed sections as there's physical differences to the transistors to emphasize power consumption of switching speed that'd still continue into async designs.
On top of that, part of what we're seeing is dark silicon and the specifics of Dennard scaling. You can't light up the whole chip and not melt the chip, so you're going to see mobile TDP chips where you turn half the chip on or off at a time either way.
But biggest issue with any adoption of growth in asyn design is the tools and skills to do such work. Though it does prove hard to compare and most of the development in the async area been driven by the EM advantages for space based usage.
...wonder where they got that one from.
I mean if it can do minimalist 3D, with opengl ES, with good performance or the same performance of a playstation 2, it would still be interesting, since high end, bleeding edge GPU graphics are not always interesting for everybody (and developing on bleeding edge GPU seems like it requires a lot of work).
A dedicated GPU always made sense for high performance, but at some point, having a console or a gaming PC that can run games with a single big chip might not be a bad idea. I have been able to play wow classic on a laptop's i5 with integrated graphics, and it was just fine.
Although I'm not sure if that new CPU is just that kind of "hybrid" design.