Intel Core i7-4770K Review: Haswell Is Faster; Enthusiasts Yawn
tomshardware.com
tomshardware.com
If you can stand the sunk cost it's a LOT cheaper than AWS. Horses for courses YMMV etc.
I buy not-for-resale chips on a dodgy basis for fuzzing (MOAR VMS!) but on the whole, bucks per MIP, or MIPs per watt - who cares?
Mainly, whether MIPs per watt matters depends on where you can get your stuff racked ;)
(x as many as I could afford). That does look like a good chip, but dodgy engineering sample trays in 2013 you can prolly get twice as many cores. That means NO WARRANTY and like, not for work.
That should be a way to do something wrong and interesting.
http://anandtech.com/show/7003/the-haswell-review-intel-core...
Unfortunately, it takes time to optimize code for each new architecture, and most benchmarks don't even bother with a recompile.
Not really griping that much; eventually people do do those kinds of benchmarks, but they do take time, and they're not really the target audience for someplace like Tom's. Understandable they would focus on gaming & general-use benchmarks.
I don't use MSVC, but /arch:AVX will get you the AVX instructions found on Sandy Bridge and Ivy Bridge. AFAIK, there is no generic "tune generated code for host" flag for MSVC, and AVX2 is not supported at all yet.
You'll probably need to pass optimization-level and vectorization flags to see more of an effect, but for the most part, the biggest gains will come when optimized libraries come out with new ASM- or intrinsic-based variants for the new ISA extension.
If you only want to compare intel to intel and want to look at linear algebra or FFT heavy workloads, that's a great place to start.
Also, I seem to remember reading that the x264 team got early access to Haswell and so were able to start optimizing their encoding routines for it. So if there's a benchmark based on that, that would be another good place to look.
The main exciting part of avx2 is that I can have more of a matrix in registers at a time. Depending on the details of the matrix multiply algs etc, this reduces the volume of loads and stores by some neat factor.
Likewise the support for gathers is neat, and makes for better performance with some of the standard matrix layouts that fundamentally have bad strides access.
What I'm really excited about is the parallel bit deposits / extract operations. They're meant for making it easy to write arbitrary bit permutation functions in a small number of rounds, but potentially could be used to support fast indexing into some nonstandard linear algebra layouts that have substantially better memory locality. I'm talking layouts that easily compete with blas / atlas etc with a lot less engineering work.
I'm really excited to get my hands on some haswell hardware in the coming months to experiment thusly.
Context for me being that I'm writing tools in Haskell to generate and run numerical codes and I think some of the things Haswell enabled are quite swell.
There's also that htm concurrency machinery that'd be neat to mess around with too.
What is the market for desktop CPUs where the desktop buyers don't also have dedicated video cards which can do a almost a magnitude better job than the APU?
Small form factors, like the iMac or Mac mini. Apple is pretty influential in Intel's roadmap nowadays. See also how the next-gen mobile chipsets are getting more powerful onboard GPUs to drive Retina displays.
EDIT: it does go up to nearly 20 watts, depending, esp if you have the HDMI plugged in. Technically I think you can get more out of a NUC than a mini in terms of raw MIPs. They're PCs, not toys like RPI.
In practice, with the advent of compute shaders and opencl, there is very little intense serial work that can't be done in parallel. My workflow nowadays is (using python as an example)
Slow? (assuming we know why it is slow and it isn't just maligned algorithmic complexity making something a runtime exponential where you can use a quadratic) -> Put it in C++. Still slow? -> Parallelize into tasks and put in a work threadpool (assuming there is a lot of this kind of work happening, else just numcores / X threads do stuff). Still slow? -> Port over to opencl (or if memory copies are unnecessary, opengl compute shaders) and keep the old implementation for backwards compatibility with systems lacking them.
The way the market is going these days, consumer devices will soon almost all be based around low power/small footprint solutions(e.g. atom, mobile, arm). Heck, it probably won't be too long before we start seeing some serious SOC solutions(what could a raspberry pi type device do in 2 years with a 4 or 6 core arm at it's core). So, that kills margins on the consumer market more or less. Or, well, at least it becomes a race to the bottom. What's left is enterprise and "the cloud", where efficiency and iops still rule. For these companies, CPU speed isn't a bottleneck, it's mostly core and io density. They simply want to run more jobs in a smaller space, not necessarily run the jobs faster. So engineering teams are starting to look very closely at GPUs as a means to get a huge improvement in core density for certain situations at the cost of having to rethink some of the software stack. Putting the GPU on the same die as the CPU makes a lot of sense from this perspective.
It'll be interesting to watch. If nvidia does have an arm platform in the works, they may win this race, though intel does have a huge advantage with being able to run multiple foundries at one. Part of me wants to believe that intel is capable of doing much more with their APU solutions, but are simply waiting until the future of the market is clearer. Once that happens, they may be able to beat everyone else simply by timing the market on newer designs(like how intel beat AMD after screwing up with itanium and P4).
Note: Having a low power SOC type solution is another side effect of all this, but I don't think that's the primary motivator for Intel's efforts here. Maybe for the atoms, but not this chip.
The power savings might get me to upgrade my laptop though.
Performance per watt is another story through ...