Compute Engine machine types with up to 96 vCPUs and 624GB of memory
cloudplatform.googleblog.com
cloudplatform.googleblog.com
Preemptible is a bit cheaper but are 600GB memory really worth it for short running applications? Until you loaded everything in memory your machine probably gets destroyed..
EDIT: Not sure about the exact CPU performance but should be quite close to what OVH offers here? With the same memory configuration and 2TB NVMe this still costs <$1,500/month (https://www.ovh.co.uk/dedicated_servers/hg/180bhg1.xml).
XGBoost will use all the cores you throw at it, and despite the recent work on GPU versions most of the time CPU cores are best.
I'll commonly run a model for 48 hours on an i7. I'd love to be able to try more models.
If (monthly) price is the main concern "the cloud" is probably not for you. Other people obviously massively value the benefits of it and pay the premium, and that is not suddenly going to stop for a new instance size.
[1] https://cloudplatform.googleblog.com/2017/09/extending-per-s...
Any provider letting you spin up bare-metal through an API almost certainly bills at much finer grain than monthly, although they may quote monthly prices to make it easier for customers to assess cost.
Add in flexible network storage options and integration with existing security infrastructure. You are thinking too small by comparing a single box to a single box.
It's never cheaper in the cloud.
Well, that's where cloud servers are great.
You'll need to break out Excel and calculate whether it's worth it with regards to usage.
Once rebuilt, using the database is fine on a normal server.
If all your data and other processing is in cloud A, moving some processing to B might not be feasible (moving lots of data takes time, security requirements may complicate setup)
I run semi-bandwidth intensive applications and DigitalOcean and LightSail are actually better deals than EC2 for the amount of bandwidth. $5/mo for 1 TB on DO/LS vs $90 for 1 TB on EC2.
We use a mix of dedicated hardware and DO/LS to meet our needs as bandwidth on the major cloud providers was just too expensive.
> notamy: I have an application on OVH (on the USD $3.50/month plan) that pushes/pulls >10TB/month
That entire discussion is recommended for anyone looking for a cheap VPS.
I seem to recall a friend having stability issues with their VPSes several years ago, so we stuck to the dedicated stuff from them, but it's been extremely good especially considering the price. Have you had good stability with their VPS services?
I am debating starting a Twitch -> YouTube video stream duplicator/archiver that would initially make money by auctioning available capacity with the long-term goal of being aquired by Twitch since their integration is so unreliable.
[1] OVH 10TB traffic throttling https://github.com/joedicastro/vps-comparison/pull/25
[2] Scaleway truly unlimited https://github.com/joedicastro/vps-comparison/issues/9#issue...
I actually tried the Hearthstone one and the very first clip in the current example 'Greatest Clips' (BJwDyxrplpo) appeared to miss the actual action (clicking "Disenchant" - which may actually have been the point since it could have been just a tease) but the rest of the clips seemed complete (and interesting).
I've thought about stream-jumping/recording based on simple indicators like increases in chat comments, viewers, followers, etc. How much of this could be built off Twitch's own 'clip' functionality (whether initiating them yourself or aggregating the manual curation of others -- neither of which AFAIK has an API right now) and collecting them later? Separate note I'm trying to hide in this pararagraph: don't overfit if you want to apply this tech to other streaming sites where real money is flowing (aka NSFW).
Personally I don't care so much about specific games on Twitch (except Street Fighter, which gets relatively little love but your videos are a real time-saver) instead of personalities. It might be worth offering this service to them, focused solely on collecting their highlights. had a tough time with the non-English streams but not sure what options you have there. I'm also interested to see how this will turn out for you using Twitch content if they notice that what you're doing is catching on. Twitch seems to be leaving a lot of low-hanging fruit behind for othes to capitalize on.
Feature wise: more playlists, maybe monthly and/or collecting the highlights of the highlights, with most comments/views/thumbs-up on previous YouTube videos. If there was a way to incorporate chat since most streamers don't in their videos, you should. https://github.com/PetterKraabol/Twitch-Chat-Downloader
Bug-wise, it seems like something is going wrong with the links as the end of this video, appx. 30 seconds of moving images but no links in Firefox with ad block [disabled as legacy]. (8ql3id1lJoM, ilkKuvuna10)
You're right about offering the service to the streamers - that's definitely the way to go to make a business out of it and it's something I've considered. However, I was mostly interested in doing the project for fun, and for some passive income, and making it a service would definitely not be passive.
The clips API returns language information about the clips, which you can use to filter them. Before that, I had to manually maintain a blacklist of non-English streamers.
I do monthly highlight videos, but they're solely based on clip views on Twitch - it doesn't use YouTube analytics, which I'm sure would improve the videos.
It is a cool idea to include chat - another thing I've considered but haven't implemented, though I've noticed some Twitch highlight channels (that do manually edited videos) do it. Thanks for the link to the downloader.
The links at the end of the videos are tricky - there's no api for that, so currently they're populated by a Firefox macro on a desktop that's supposed to be run every day - looks like there's an issue with it running! The better version would be to use a webscraper or headless browser to automate those clicks via the render server. That's what I'm supposed to be working on next, in fact...
In these cases the amount of money you spend on hardware virtual or otherwise is negligible. Depending on what you do it might just as well be a rounding error.
not so sure about that + a hoster who provides such a machine as bare metal wants a setup fee, needs time to setup and a minimum contract duration much longer than one month
guess there are not many hosters which have such a beast as bare metal in stock and available in few minutes (are they any hosters at all?); they will order sch machine themselves and you will wait at least a week
Minimum contract length: 1 month, total cost (including setup): $4,384 (on-going month to month cost thereafter: ~$2,191.20).
For that you'd get an aggregate total of 4TB RAM, 7.6TB HDD (SSD), 96 real Intel E5-1650 v3 cores (or 192 vCPUs) and 800TB of bandwidth.
Sprinkle with terraform/ansible/k8s/docker and you have a resilient, massively powerful compute cloud with no long term obligation that's about half the price of GCE if you keep it around beyond 30 days. Or another way to look at it: if you needed such a platform for two years, your second year would be free compared to GCE.
One major issue with this approach (versus GCE's "all in one" box) could be network performance bottlenecks depending on what task(s) you were using such a cluster for.
So we've established that aws is far more expensive now since the tco is taken into account in both cases.
If you can't tolerate the server being down for more than the lead time it takes to get a new one, you need to already have one on standby. The lead time is probably at least a couple weeks, but there's no guarantees since you're depending on vendor availability and hundreds of other things out of your control.
Depending on what you're running on it, you'll also probably want to test software upgrades and have fallback plans when you deploy.
The thing the "cloud" version gets you is zero lead time, along with the ability to spin up a second instance (or ten, if you want) while you deploy a new version or just want to do some testing.
Who monitors that PC, and does preventative maintenance, etc? If a part looks like its going to fail, where do you migrate your workload to, so you can take that server offline for repairs. (you need multiple servers)
Since you have multiple servers, how much does your 40GB networking cost (with 100GB uplinks) so you aren't constrained by the network? And what kind of storage network do you have, so that you can live-migrate these running machines around to different machines?
Lastly, if you co-locate the server somewhere, what does it cost for multiple redundant internet connections to the facility? And where is your failover facility, that is at least a few hundred miles away?
They're great boxes for cheap on-prem etl clusters.
I prefer sell R810s over those hps because they're 2U and have better power consumption with the same specs.
The total cost of ownership of a computing asset is several times greater than the cost of the actual asset.
Think of it this way: a dog can be obtained for a very nominal cost (or free) but the cost to house, feed, entertain, and provide healthcare for is non-trivial.
it's not unheard of for just the costs of deploying a new device into a large organization to be something like EIGHT TIMES what the cost is for the actual asset. That's just to get the hardware deployed and NOT the cost to keep it running.
Cutting down on TCO and streamlining the deployment of resources is a big part of the sell for cloud deployments. Particularly for computing assets that may otherwise spend a lot of their time idle.
How loud is it? How much energy does it consume while running? How hard is it to configure and keep running? What kind of firmware does it have and will it be a problem updating?
These are all the questions I would have before buying a beast like that...
Also, this machine you link to might cost $1,850 but:
- it is used, not new, so it can break any minute, and you don't have any warranty.
- on GCP just $2,100 per month buys you similary speced machine AND peace of mind.
- running this machine 24/7 poses some significant electricity cost.
EDIT: - $2,100 is probably lower than a salary you would pay an IT technician maintaining your machine(s).
well it's not just electricity. to run your server you probably at least need a network, ups and probably other things. this stuff especially gets ugly when you want to have a network with many servers. at least most dedicated server providers charge a ton of money for interconnection of servers. well at least ova actually provides a vRack for dedicated servers but its not always free.
Yeah, the main reason to go with Amazon in this case is if you only need the box for a few hours (ie: you're doing data science or similar). For long term high load use, bare metal, even managed bare metal like that is almost always cheaper.
Sure, it wasn't the highest performance per CPU, but I didn't have to buy the bare metal, I can scale up the number of cores if need be, and I can fire one up whenever I want one.
Frankly, I have a hard time understanding why the convenience of being able to call an API to get a VM (vs using an API to get bare metal) continues to be an advantage. I am reminded of the reddit articles about all the effort they went through to re-optimize their app (by batching queries) for the longer database latencies at AWS... Its like they never considered all that work might also apply to bare metal and save them even more money...
I would think lead time. To get the quote from your hardware vendor, get the PO approved from finance, get the box built and shipped, getting it racked and stacked in DC. This whole sequence can easily take longer than a month.
I totally believe that OVH is cheapest for European companies serving European customers, but that’s apples to oranges when we’re talking about American cloud providers.
those machines have a lot of value for specific workloads.
When GCP announces a feature that’s 1% better than AWS their employees flood the HN post, but when this story hits no one shows up.
I've been seeing that medium data keeps getting bigger (i.e. the features of traditional RDBMS are eating away at the need for specialized / distributed stores for data analysis). But so too does it appear that small data is getting a lot bigger too—just load that dataset into memory for analysis. 4TB of memory allows for pretty big "small data."
"I remember back when we used to do gradient descent to estimate linear models; back in the long ago when we didn't have 900 exabytes of memory attached to our NVidia Matrix Crusher 9000 linear algebra accelerator unit."
I mean before they'd train a model of 1,000 inputs and then test it against another 50 and call it a day. Now they want to train it against 1,000,000 inputs.
Am I completely off base? It's not my area, though I work with databases, my observation is that developers always want to use the most data possible even when it doesn't really provide any benefit.
The thing is, it didn't really matter; Postgres still had a ton of performance left over even after the product went live. If you can still fit it in RAM, why waste $$$ of dev time over the $$ cost of a bigger instance.
And whether or not more data is needed or collectible varies by discipline. Astrophysics collects way more data than they used to because 1. they need it. 2. instrumentation allows it.
Some kinds of data collection hasn't scaled up however. Surveying humans is expensive and labor intensive. And for many things that you might want to study about humans, you can't simply afix a sensor to them. So, what might have been only accomplished through big data, or medium data methods a few years ago can now be loaded into memory (i.e. small data strategies).
There are many applications where a single addressable memory space are still required/preferred.
Look into "tidalscale" if you want too see the logical extension for this.
https://www.intel.com/content/www/us/en/products/processors/...
Disclosure: I work on Google Cloud.
Interestingly, AWS have X1 and X1E-class EC2 offerings, two of which are significantly larger than this new Google offering.
Disclosure: I work on Google Cloud.
To me, legacy software would be something like using Novel PeopleManager 2002 or WordPerfect 7.
Postgres 10.0 is < a week old, and is a perfectly fine path to go for a brand new application. You are fooling yourself if you think you should just throw every new application in MongoDB or Cassandra because they are "web scale".
If you are using "Legacy Software" to mean "software that has existed for a long time", then I guess sure. But there are many pieces of software that are very new, and could benefit from a single instance with a lot of cores.
The most common use of "Legacy Software" is "Old crusty stuff which needs to be replaced", which is NOT at all the case for MySql or Postgres.
Nah, with scaling, I‘m more referring to approaches of eg Cassandra instead of „add another read replica“.
https://blogs.saphana.com/2014/12/10/sap-hana-scale-scale-ha...
HANA's roots lie in BWA (SAP BW accelerator), whose main game was distributing large amounts of data across commodity clusters.
HPC people are much technical and less conservative, in general, and can sort themselves out.
I.e. just like Azul Systems' system, it's for running bad enterprise applications that you cannot scale with regular means.
I think you can store every single uncompressed frame of a bluray movie in memory of an AWS X1... but even then, so what?
I never quite understood why this is necessary if they're cutting up larger machines. Why should the total network cap matter if I have two 8-core instances or one 16-core instance on the same physical machine?
edit:
https://cloud.google.com/docs/compare/data-centers/networkin...
"The egress traffic from a given VM instance is subject to maximum network egress throughput caps. These caps are dependent on the number of cores that the VM instance has. Each core is subject to a 2 Gbps cap for peak performance. Each additional core increases the network cap, up to a theoretical maximum of 16 Gbps for each instance. The actual performance you experience will vary depending on your workload. All caps are meant as maximum possible performance, and not sustained performance."
Judging by the ratio it's likely something such as 2x 56gbe network ports on the host (=112Gbps), which has 56 vcpu (2x xeon 14c/28T)
They only use live migration to do host maintenance iirc.
We hate idle resources. :)
But yes, they probably strive to keep below 10% vacancy on all hardware.
https://www.intel.com/content/www/us/en/products/processors/...
Its amazing how much hardware you can pack into a single machine for 10k€. Last year our group bought two additional high-memory (768GB) nodes for around that price each (including support for a couple years from the vendor).
A few years before we bought 40 nodes with 128GB RAM each, for a similar price to last years high-memory nodes (and a fast interconnect and a lot of storage).
If you are at a larger research institution, you probably also have an IT department that can co-locate your hardware for next to nothing (compared to cloud). There you also will save a lot of ingress/egress, storage, backup, etc. costs.
Regarding the per student costs, even with cloud instances I would consider running a traditional HPC job system (grid engine, lsf, torque, ...). The MIT had a nice solution with Starcluster [1] to easily deploy a SGE on AWS. It looks a bit dead now though.
Isn't that already the case for 1 month? Bare metal doesn't mean own data centre or colocation. If you go with a hosting provider most offer dedicated hardware on a monthly contract. As long as you need them longer than 1-2 months that should be significantly cheaper than Google/AWS.
I'm surprised no benchmarks are mentioned.
I can't be too specific about it, but if involved creating a very large tree structure and updating, pruning and transversing the tree a lot
If the algorithm is updating and reading a large data structure a lot it's only practical from a speed point of view to hold the whole structure in RAM
The code was written in C
(We also maintain a Hadoop and Cassandra cluster, and I use Spark for distributed computation - but those are different projects)
My advisor's company straddles the public / private divide, but we've definitely done some simulations for private clients on NERSC, and I assume we weren't misusing hours allocated for some other purpose.
[0] https://cloud.google.com/compute/pricing#machinetype
[1] https://cloud.google.com/compute/docs/instances/preemptible
As for the RAM itself - take 12 DDR4 128GB modules, for example https://www.heise.de/preisvergleich/crucial-lrdimm-kit-128gb... , and fit them into three channels of four modules each to get 1536 MB.