Intel’s Thread Director: Assisting the OS to make task placement decisions
anandtech.com
anandtech.com
Professional me knows I'll be stuck figuring out that some bursty process that really needs p1 access all the time, in spite of usually being idle is going to be super pissed off when he gets paged at 2am. and he'll be stuck spending a lot of time figuring out how to pin that process to that core.
Worker drone me is going to be sad thinking that slack and chrome are snarfing the good cores while my compile times suffer.
Bit of a hot take, but it's a tragedy of the commons situation. Programmers are smart, they'll find tricks to grab the fast cores. There are a maybe zero organizations that can get alignment to keep important processes on fast cores, it's way better to just be fast and point the finger at other teams for being slow. What a time to be alive.
It's super cool. it has potential to be amazing. But even forcing every task to run on an E-core, and a heavy bias to fast cores being idle will be gamed. I guess it's better to know than be surprised.
Hey, maybe we could even have kept Flash around if it were easy to prevent flash ads from fighting each other to take 100% CPU. It might have taken some iteration to figure out how to map a simple GUI-appropriate slider to a scheduling policy (10% -> 10ms every 100ms?) but I bet it could have been done. It would have had a nice side effect of appropriately directing user outrage at offending applications, too.
I wouldn't mind the client side scripts so much if they behaved. Even with uBlock and such, random pages will peg my CPU, drain my battery. IIRC, allrecipes.com, imdb.com, and of course youtube.com.
Forcing Reader View often helps, but it's hit & miss.
The Great Suspender worked ok for Chrome. But I'm now 90% Safari and 10% FireFox.
Adjacent Idea: Browser history should also include other meta. Like page size. (weight), resources consumed, time spent on page, when page was closed, permissions requested & granted, etc.
In a way that can't wait tens of microseconds?
> and he'll be stuck spending a lot of time figuring out how to pin that process to that core.
Pinning to cores is really easy.
> Worker drone me is going to be sad thinking that slack and chrome are snarfing the good cores while my compile times suffer.
I'd say something about adjusting priorities but there's a much more important thing to realize here.
If your CPU didn't have this, you'd have 6 fewer cores. Or probably 12 in the next version. Even if your compiler can't use two of the fast cores, it's going to run much faster overall.
> tragedy of the commons situation
If you want to improve your own responsiveness at the expense of other programs you could already busy loop everything. And nobody does that.
I don't think it's something to worry about.
I was trying to reduce power consumption on my phone, and disabled all but one low perf core (the phone's SoC has 2 high perf and 4 low perf cores), and things were very noticeably laggy.
But, then, I tried with only two low perf cores enabled. Everything that I used my phone for (calls, messaging, notes, web browsing with an ad blocker) worked fine. So, the extra four disabled cores really didn't affect my ?average? use case.
I was concerned about consumption during active usage, not idle, but disabling the cores didn't affect power consumption for either (even with only one enabled).
It can advise on shifting cores around as fast as the OS is ready to, so the delay doesn't have to be worse than a normal sleep.
If you have a normal "active game" timer frequency, then Windows is probably checking in every millisecond, and it will have extremely accurate data to use. And that's even without any special mechanisms to react faster.
The 2x big hardware threads are from the big-core (SMT / Hyperthreading), while the 1x small hardware thread is the small-core.
Since this chip is 8-big cores + 8 little cores, the math works out.
-----
A future chip is rumored to be 8-big cores + 16 little cores, which can be implemented instead as 2x big-threads + 2x little-thread cluster (1x big core + 2 little cores per cluster).
Though of course, that depends heavily upon the implementation details of this "Thread Director".
An important difference though is that guests do not have to be aware of most of this scheduling behavior, the exception being NUMA. My guess at this point is we might see something similar to virtual NUMA topologies, but for big-little.
In this new model, some cores are hyperthreaded, and some aren't. Some cores are "full spec", some aren't.
Similarly, even if a hyperthread has low utilisation, if it's twin is busy you will see lower performance.
It's not like "logical CPU 3" is slower than "logical CPU 7"!
This is like a highway with equal width lanes. Sure, there might be more traffic in some lanes, but the lanes themselves are equal.
The new Intel CPU is like 8 wide lanes that can be used by up to 16 motorcycles or 8 cars (or combinations thereof), alongside 8 medium-width lanes only usable by small cars. It's bizarre for a desktop CPU.
PS: Looking at the die shots, it boggles the mind that they didn't include 16 efficiency cores! They're so tiny that it would have been a negligible area increase, but given the relative performance it seems like it would have been worthwhile. I'm guessing memory bandwidth limits are holding them back somewhere...
Let's say the thread starts on a high performance core and enables the code paths for features the efficiency core doesn't have. How will the software know to not move the thread to the efficiency core? If it did do so wouldn't it throw errors? I frankly don't see how they can solve this problem with software and not have to recompile everything to be hybrid arch aware. In the android space everyone throws their phones away every few years so this hasn't been an issue. To me it seems like Intel is creating another Itanic situation where no one is going to compile their software to target the hybrid paradigm.
It also drops AVX-512 support entirely, presumably because the efficiency cores don’t support it and the problem you mention isn’t easily solvable.
Does this mean we will get RPi style Gracemont SBC's or did Intel drop that after Edison?
I wonder if Zen 4 will have AVX-512 support. If so, given that it will use a 5nm process instead of this processor's 7nm process, it will absolutely blow it out of the water.
It would be somewhat ironic to see Intel being trounced by an instruction set they invented and then stopped using!
AVX-512 was a mistake, causing Intel lots of issues for the benefit of a new benchmark. It heats up and clocks down the chip for other work. And it takes up too much space.
Again, please, no.
Intel developed the instruction set and expected to rapidly shrink the die to 10nm and then 7nm, which would have fixed the power draw issues. The shrink never happened, and this then made AVX-512 look bad.
It's not AVX-512 that's at fault, it's the manufacturing process. Fix the process, and the instruction set can shine.
Many people writing vector code by hand say that they much prefer AVX-512 over its predecessors because it is complete, flexible, and powerful.
The only reason Intel didn't lose every benchmark against AMD recently is because for some workloads AVX-512 doubles throughput despite being hamstrung by the power draw and overheating problems.
A 5nm chip using AVX-512 might only need to clock down 10-20%, or not at all. Or just use turbo for a shorter period.
Besides the Broadwell instructions, all the instructions supported by Tremont (the previous core in the Atom series of cores), but not by Broadwell, are also supported, plus VNNI from Cascade Lake and some of the non-AVX-512 instructions introduced by Ice Lake.
I’m excited for this part because Tremont was never a thing normal people got to buy. I think it’s going to be a real pleasure to program these E-cores.
There could be a dirty state flag that indicates whether vector registers or other advanced instructions were used.
> If it did do so wouldn't it throw errors?
The kernel could catch those errors and move the thread back to the performance core and resume execution at the trapped instruction and mark the thread as non-schedulable on the efficiency cores.
However, there was a real problem a few years ago with some Samsung Exynos processors, where the cache line size was different between the big and the little cores.
A few programs made assumptions about the cache line size and had weird behavior after starting on one kind of core, then being migrated on the other kind of core.
they couldn't change the thread scheduler without a new major release?
pull the other one
I'd argue that thread schedulers are a core component of any operating system.
That being said: Windows8 had support for heterogenous processors (aka: ARM's big.LITTLE), so I'm not sure if a fundamental change was really needed here. Microsoft already put a lot of hard work for those Windows8 phones (which used the same kernel as Windows8 desktop).
The system and apps on it make use of multiple cores, that's all that matters for the scheduling point.
And it shouldn't be hard to get into Linux, I wouldn't really say this partnership notably hurts that.
Or - probably more likely - Microsoft, Intel, etc. want to keep that ~90% worldwide market share of theirs from shrinking.
Any numbers on JVM performance on them?
That seems unrelated to however well Java runs on current Graviton.
I still don't understand why is it necessary in an OS which can juggle threads hundreds times per second using priority system. At least ignoring power efficiency aspect and concentrating only on performance or perceived performance aspects.
This is the same motivation behind why high-performance software architecture uses a "thread per core" model. There are large throughput gains to be had by minimizing context switching and CPU cache sharing across threads.
Stuff like OpenMP is great for doing quick things like a parallel for cycle
Volkswagen doesn't need to worry about Polo sales just because Ferrari outperforms them.
These aren't going to the server yet.
Just like whatever Khronos does, AMD, Intel and NVidia collaborate with Microsoft in doig native DirectX hardware support and only afterwards those features show up on OpenGL and Vulkan extensions soup.
Hardware T&L, GPU shaders, Mesh Shaders, Compute Shaders, Ray Tracing, certainly not.
[0] https://www.khronos.org/registry/OpenGL/extensions/NV/NV_mes... [1] https://devblogs.microsoft.com/directx/coming-to-directx-12-... [2] https://www.youtube.com/watch?v=Ge427_2VORo&t=1607s
It's "just" hyperthreading on servers right now that could benefit from this.
Does anyone know what is meant by the scheduler analyses the program?