Google Cloud Platform is the first cloud provider to offer Intel Skylake
cloudplatform.googleblog.com
cloudplatform.googleblog.com
Since some GCP engineers are watching: Presumably we'll see some new zones to provide these processors, or will it be a limited release within existing zones? And if so, will you be moving away from homogenous zones in the future?
Disclosure: I work on Google Cloud.
Disclosure: I work on Google Cloud (and historically focused solely on Compute Engine).
The cache is also a whopping 56 MB.
Disclosure: I work on Google Cloud.
Doing a quick comparison between our Haswell servers and Skylake servers we do see a nice little speed bump of ~10-20% on matrix heavy code. The bump is negligible for most other things.
> You'll be able to use the new processor with Compute Engine’s standard, highmem, highcpu and custom machine types. We also plan to continue to introduce bigger and better VM instance types that offer more vCPUs and RAM for compute- and memory-intensive workloads.
[1] https://cloudplatform.googleblog.com/2016/11/power-up-your-G...
Even on the less complicated end, some years ago throwing a shitload of page cache at MongoDB was the only way to maintain high write loads, and eventually one reached a point where one had to keep an entire shard in cache. That bottom end threshold is lower than you think. I don't know if that's changed.
The key here is it's a personal development workstation on which he has these resources. If that's a requirement, I would argue there's something fundamentally wrong with his development processes requiring that much data locally without even a server.
I'd love to spin up a 32x512GB for a DB Server.
Disclosure: I work on Google Cloud (and helped a bit in our Skylake work).
Disclosure: I work on Google Cloud (but not compliance).
TL;DR:
- Google's global SDN is ultra-secure, and Google carries your packets on its network rather than dumping them onto public web. Google's undersea cables are shark bite-proof! [1]
- Google Data Centers are mostly homogeneous and rely on Google-build hardware rather than vendors. This greatly helps with securing infra - only one vendor (Google) to trust, only one set of best practices to follow, lower risk of misconfiguration, and exposure risk is minimal.
- > 600 security engineers.
- Encryption-at-rest and in-transit is ubiquitous.
[0] https://cloud.google.com/security/whitepaper
[1] https://www.youtube.com/watch?v=XMxkRh7sx84
(work on Google Cloud, not in marketing :) )
For security, I don't care whether my traffic as routed over a private network or public web as long as the data is encrypted, because a "private network" is only as private as thousands of miles of fiber and every single network facility it traverses can be. Which means 'not very'.
Does Google build it's own hard drives/SSD's and CPU's? If not, then that's at least 2 other vendors to trust since both CPU's and hard drives are related to hardware security.
On your first point. There is a very strong risk of inadvertent misconfiguration or your employees not following best practices.. Or even poor documentation. Google cloud by default gives you a global secure vpc that never traverses public internet, so there's less risk because the baseline of security is high. Sure, you can run VPN tunnels between data centers. Point is, with Google cloud you don't need to.
On the latter point, of course you cannot eradicate every vendor, especially within the obvious context of this specific Intel announcement... But, again, not having three dozen flavors of configs and four network router vendors does make a difference. I would encourage you to read the paper I linked to above (discussed this topic in great detail), and also perhaps [0].
Also security folks at Google, who are much more qualified to discuss this topic than I am, frequently post on hn. [1]
[0] https://cloudplatform.googleblog.com/2016/02/Google-seeks-ne...
It's also important to pay attention to which services you consume, as not all of them are certified to the same degree.
Technical management can sometimes be persuaded, given overwhelming evidence, that competing products/services are superior or more cost-effective but ultimately conclude "we'll never sell this solution to leadership".
The only thing I really miss from AWS is RDS's postgres.
I don't even want to talk about Azure.
Why do people use AWS? 1. Free tier. 2. Plethora of PaaS as you mentioned. 3. AWS was there first, has brand recognition and is a safe choice for management.
It's superior for running Windows at least, right?
Disclosure: I work on Google Cloud (so I want to sell you our services).
I'm working on a migration to it, so I'm not very experienced with it, but so far, it's been very painful compared to GCE or AWS, in which I've run production stacks. I'd rather not comment further, simply due to my relative new-ness to the service and the chance that it's just lack of experience. The customer service is, at best, run at a glacial pace.
> "It's superior for running Windows at least, right?"
I'm a Linux guy running a platform agnostic Linux stack, but I'd assume so. I get the feeling so far that it's really good if you want to run MS-SQL and .net, and garbage otherwise.
The only reason we're migrating from GCE to Azure is because our GCE credits are expiring, and Microsoft gave us the YC Credits offer for Azure. We'd rather stay on GCE if we could. Also, postgres RDS/CloudSQL is the one thing we miss from AWS
Disclosure: I work at Microsoft; opinions are mine.
I wasn't going to lay it out here, but, as I have your ear:
I'm not saying it's impossible to run a Linux stack on Azure - but man, trying to image the machine, for example, is a whole rigamarole.
Want to run your own image on a scale set? Oh, well, you need to craft a JSON template, by hand. There also appear to be limits on how many machines can run off an image.
ARM is a a mess (IMHO), and it's impossible to select a custom image when creating a new resource group. It also seems (correct me if I'm wrong) impossible to change the vnet of a VM/ARM after it's created. it also seems like ARMs can't share an existing vnet. Again, please, correct me if I'm wrong. I'm new to this service.
I may have to drop to running a bootstrap script to get my stuff working, but the idea of doing a curl | sh is pretty horrific to me, from a security perspective.
Non MS-SQL as a service? Nope.
The new managed disks are very nice. I like those a lot :)
Also, Azure times out my ssh sessions :(
When I last chatted with our Account Manager I mentioned Postgres, yeah. We're currently running our own on GCE, but it would be awesome to have it aaS, with replicas and automated backups that I don't have to keep an eye on all the time :).
For anyone else who is on AWS (maybe because of rds postgres, or because it's client work that needs to be on AWS), you can go outside of vanilla AWS to get a similarly great dev UX. I've used Convox + the Weave ECS AMI on a project, and it was pleasant.
Just getting into Kubernetes :)
I believe that whatever GCP does, does better. Faster network, faster spinning and turning off servers, simpler quotas, etc. GCP doesn't offer that many services as AWS and most of those are not very compatible with Apache tools (although they're doing their best, e.g. BigTable looks like Hbase from outside).
In our case, we're running ~28,000 servers on GCE and utilizing many services, especially PubSub, BigQuery, etc. The price we're paying for our setup is roughly 1/4 of a similar setup on AWS.
Additionally we're running our search engine at top of it with couple of other services.
You can check what we're up to here https://pex.com
Uhh, Google will terminate preemptible instances whenever they need the resources as well. Hence the name, "preemptible."
Also I personally find much easier to understand is the pricing. It's very consistent and simple. Preemptible servers are discounted by 80% from the regular ones. On AWS the price changes, also they only offer older generations while GCE offers any type at the same discount.
Disclosure: I work on Google Cloud.
Sadly, (because KNL doesn't have nearly the same volume) until Skylake became a real thing, I had no reason to update this. I'm planning on dusting it off now!
FWIW we're currently using GCP, generally love it, and I'm looking forward to trying out Skylake...
For the numerical, AVX stuff -- it should be mostly automatic for you, if you're already using optimized libraries at the core -- MLK, BLAS, that kind of stuff. They'll be transparently upgraded for you -- ideally -- to take care of these things. They normally check what your CPU is at runtime, and pick the fastest implementation among a few different choices it has.
You will need toolchains to support this all, but for the most part that likely won't be a burden unless you want to get your hands dirty and start it yourself -- inevitably, this should all mostly be "pre-canned". Your optimized linear algebra, vector, and math libraries are what will mostly concern themselves with this, not you necessarily. In fact, several of the things already available can probably use these new extensions! I bet if you're using Intel MLK for example, it will probably "magically" get faster on these Skylake machines by using AVX512 automatically.
If you want to understand more: you can always go grab an SSE/AVX reference, check your /proc/cpuinfo, and write a few simple things on your own to get a feel. Your toolchain will definitely support it :)
(Aside, I think you mean MKL.)
There is a 200MHz clock reduction when running AVX512 instructions. If your code makes heavy use of AVX512 there is of course still a big net win, but I'm curious of the impact with more heterogeneous workloads. We have an app that is a mixture of scalar and vector code. Some, but not all, of the vector code would benefit from 512 bit vectors. But how much does the clock slowdown when running this code bleed over into running the other non-AVX512 code? I guess I'm asking how quickly it clocks down, and how quickly the full clock speed is restored. Worst case it seems you could be running full time at a 200MHz slowdown due to blocks AVX512 instructions scattered throughout the application. Is that a valid concern?
I'd have to look up the specifics; but does AVX512 simply slow the clock, or does it actually have some kind of limited number of hardware ports? I wonder if some clock slowdown would be very much of an issue, since clock-for-clock, you should see better performance on Skylake anyway.
Just curious, what kind of workloads do you think you're looking at here?
In my case, yes, Skylake would still be a win over older hardware, but the question is whether to use AVX512 or not. The workload is a real time animation system with a bunch of nodes in a graph that get evaluated in sequence. Some nodes would benefit from AVX512, but others would not. So the question is, if we vectorize those nodes that would benefit and get a speedup there, will the other unvectorized nodes now run slower as a result of the lower clock speed, canceling out the benefit.
It sounds like your case is a much better fit for AVX512. Out of curiosity, have you tried running on Xeon Phi, which also supports AVX512?
SGX (secure enclaves)
TSX (Hardware Transactional Memory)
SGX can store SSL keys in a hardware protected enclave that can't be accessed by hypervisors/AMT so that can be useful for security sensitive stuff.
Cloudflare probably could have used MPX to prevent their recent leak with minimal performance overhead.
TSX is cool.
Yes, it might have stopped the CloudFlare case, if they were willing to pay (I believe) for L4 over-read protection overheads (2x I believe, with a LOT of variance between GCC and Intel's compiler), and give up multithreading[1] in their application, and probably deal with other false positives and unsupported things.
SGX is completely neutered in Skylake and totally useless without getting your enclave signed by Intel, unless Google has struck a deal or something. You can mostly ignore it. It's good for preparation of your application against future processors where you'll have control over this, though, I guess...
TSX would be nice, yes. It's deeply annoying it's taken them so long to get right and that they've stratified that feature amongst CPUs -- is there any _real_ reason my Kaby Lake XPS13 can't support TSX? I doubt it other than "market segmentation makes us more money". I guess now I'm just ranting, though.
(As an example, my Xeon D-1540, a Broadwell-family chip, advertises the TSX bits in CPU feature flags, so it's not errata'd off there.)
https://aws.amazon.com/about-aws/whats-new/2016/11/coming-so...
Great job GCP team!
It's something that I've wanted to play with for sometime. It's cool that GCE has them available as a service.
$ egrep '^(model name|microcode)' /proc/cpuinfo | head -n2
model name : Intel(R) Core(TM) i7-6820HQ CPU @ 2.70GHz
microcode : 0x9e
$ egrep -o ' (hle|rtm) ' /proc/cpuinfo | head -n2
hle
rtm
This is a mobile chip and an old stepping (and running the microcode that supposedly addresses the issue, presumably by disabling it on 'bad' chips?), so I'd be surprised if newer chips (especially server ones) didn't have it.Your calculator page is unusable on mobile due to fancy "material" form filling.
And the pre-release Skylake server procs seems to be gimped as it's missing a few features versus the actual official release Skylake server procs.
Disclosure: I work on Google Cloud.
Early work was done in AMD and we had an a1 series pre launch.
IIRC, g is for Google or Generic where there are no promises on architecture. And f was for fractional.
I've been out of Google for 2+ years now so reality may have drifted from that original scheme.
It hints there's Individual Accounts, but I see no way how to set it to that?
Disclosure: I work on Google Cloud.
Back to AWS it is.
disclaimer: personal opinion. I work at Google but not related to this area.
https://cloud.google.com/compute/docs/regions-zones/regions-...
For example, in us-west1-a, you're getting a 2.2 GHz base clock E5 Broadwell.
Disclosure: I work on Google Cloud.
It would also be mentioned in the article if it were.
My guess is that you have no idea how much effort goes into verification and testing of something as complex as a microprocessor. A significant chunk of NRE costs goes into verification and test AFAIK.
Though that was the FPUs in Intel processors, not AMD. So it's not very good precedent.
[1]: ISA => microcode is equivalent in some respects to C => LLVM IR
There was also an issue with transaction memory with both Haswel & Broadwell. The fix was to disable TSX support via micro code update. I wouldn't call disabling a fix personally. I doubt Intel even compensated the folks who bought it for the TSX support.
For example the TLB access and the fast path of the TLB miss are not microcoded. You only get to microcode if you have to set an accessed bit or a dirty bit, or if there is a page fault.
Regarding your criticisms of AMD's past processors, are you coming at this from a consumer standpoint, or are you just disappointed with their architecture and/or implementation?
I ask because the Phenom II seems quite loved by its users[1], and the FX series was also quite good for its price. I had a FX-8320 in my last PC and it was quite a capable CPU in my opinion, and many others seem to agree[2].
[1]: https://www.newegg.com/Product/Product.aspx?Item=N82E1681910...
[2]: https://www.newegg.com/Product/Product.aspx?Item=N82E1681911...
http://www.tomshardware.com/forum/246780-28-issues-stop-ship...
this post would have been interesting if they had included those tests.
Disclosure: I work on Google Cloud.
[1] http://www.intel.com/content/www/us/en/benchmarks/server/xeo...
Disclosure: I work on Google Cloud.
Skylake Xeon E3 is on the current older server chipsets..
Purley is supposed to be a significant improvement.
[1] https://cloudplatform.googleblog.com/2016/11/power-up-your-G...
Another commenter already brought that issue up, but thanks for pointing it out again. I still think that it's quite silly to claim that Ryzen Rev. A may end up being a paperweight based on a mistake that took place a decade ago. Whatever floats your boat, I guess.
And from what I read, it seems like it was an extreme edge case, so the TLB error was triggered only during specific workloads. Sucks to be AMD back then.
You have no idea what you're talking about if you think for a second that a large CPU vendor like AMD would delegate verification to "its customers". It's like saying that Boeing just builds planes and tells airlines to make sure they fly correctly before allowing passengers to board them.
This is uncivil and not ok on HN. Please take time to edit this kind of swipe out of your comments here.
Also, how was our argument a "flame war"? It only lasted for 3-4 replies, and was quite civil in my opinion.
I tend to always agree with your decisions, but this one is a bit too extreme.
When disagreeing, please reply to the argument instead of calling names. E.g. "That is idiotic; 1 + 1 is 2, not 3" can be shortened to "1 + 1 is 2, not 3."
Similarly, "You have no idea what you're talking about if you think for a second that a foo like bar would baz" can be shortened to, "A foo like bar wouldn't baz."
HN's rules apply regardless of how other commenters are behaving. If they didn't, we might as well have no rules, because it always feels like others are behaving worse.
We detached this subthread from https://news.ycombinator.com/item?id=13725460 and marked it off-topic.
Disclosure: I work on Google Cloud (and of course want you to pay us to use Skylake)
I already use a nice E3 Skylake workstation at work, thank you very much. Could be an E5 or a big i7 with more cores (I just rechecked on ark and I guess it's still not available publicly for Skylake, but I also guess this can not held for very long...), but we want it to be close to our product target, which is an E3, so E3 it is.
For tons of practical purposes, lots of reasonable E5 are comparable to what you can get with the biggest i7 -- and during the last gen for tons of practical purposes gen to next gen comparable CPU have only progressed slightly. Of course there are workloads where you want to use the most insane CPU, or some CPU with new features / better perf in some niche workload, and so over, but when you start to have such advanced needs I doubt a little that OTS platform solutions are better for the majority of people with highly advanced needs...
But yeah, on the mass of people, some will remain interested by "your" solution. Meaning mainly Intel solution, given how you advertise Skylake soo much...
How much technical (DevOps-y) do you have to be keep it secure & available?
Run your infrastructure however you want. This setup would never fly for a real organization.
I'm sure they would like some more customers :)