Train an AI model once and deploy on any cloud
developer.nvidia.com
developer.nvidia.com
So your skillset is reusable.
There is no free lunch. But if you learn k8s, moving from AWS EKS to Google GKE to DigitalOceans hosted k8s is easy.
If you are the one maintaining it it's a full time job handling all these edge cases, it's completely miserable and I wouldn't recommend it to anyone.
Are you using AKS, EKS, GKE on those providers, or deploying your own k8s on top of the compute those providers offer? It sounds to me like the former.
But I enjoy working with Kubernetes.
The post I was replying to seemed to be saying (by analogy) “Linux is hard to manage because I run into all sorts of trouble trying to support a mixed environment of SuSE, Ubuntu and RHEL, therefore Linux is just too complicated”.
For example, the consistent use of labels as a way to identifying groups of resources that need to coordinate with each other is very useful for any distributed system. I find myself looking for them in say, CI/CD systems (in the form of agent tags), or at the application level in say, matching players to game servers.
I enjoy working with Kubernetes, but forcing a complex domain into something legible is a recipe for catastrophe. There are quirks, across cloud providers, and this is just another day in Ops, with or without Kubernetes. (See: https://www.ribbonfarm.com/2010/07/26/a-big-little-idea-call... )
"Kubernetes has been essential in making 8-figure monthly cloud spend happen"
The best though was when I ran across someone in the org trying to run a single container to run a periodic job in its own cluster. They spent half the day trying to get it to work with ingress.
You can imagine how it came to a head when the company realized they were spending hundreds of thousands per month on idle clusters in AWS.
So instead of learning how to deploy on GCP, AWS, and Azure, which is only 3x more complicated than deploying to a single cloud, you should learn K8s, which is 10-15x more complicated, in addition to still having to learn about all the various ingress controllers and weird quirks that are completely different on each cloud provider. Doesn't really track for me.
You can learn k8s in a day. It's really simple.
> various ingress controllers and weird quirks that are completely different on each cloud provider
Which are thoroughly documented and not that hard to implement or understand. You'd be reading about each cloud's nonstandard ingress even without k8s.
The beauty of k8s is you can run your software locally and have a much easier time lifting and shifting to another cloud.
Fitting to the shape of a cloud provider is a great way to never leave.
Another benefit of k8s is that you treat your services as cattle you can easily spawn and kill. Adoption of k8s naturally leads to anti-fragility, anti-brittle best practices.
But I think most developers don’t care, and instead, should interact with a platform built with Kubernetes as a foundation.
Probably the biggest one is understanding you don’t ever do anything directly with Kubernetes.
I use gpt-4 through the API where you can set your own system prompt. I developed one that basically instructed it to give me kubectl commands to solve my problems and then wait for me to give it the result before continuing. Through this I learned the practical techniques and which kubectl commands you use on a daily basis, which is so much more helpful than reading the documentation which just gives all commands equal weight.
EDIT: Oh, and definitely watch a few "TechWorld with Nana" videos on YouTube. She does a great job of explaining the architecture, terminology, and philosophy of k8s which I think it very helpful to know.
That mindset of not doing anything directly comes from understanding that you are not setting up pods up, but rather, you are setting automated processes up. These processes knows what the desired state is, and if there is something that changes things, it takes actions to get back to that desired state.
So you don’t create pods. You create deployments that maintain a set of pods. You don’t assign pods to nodes. You define pod affinities, and use node selectors on nodegroups. You use pod priority. You make graceful startups and shutdowns work correctly. And so forth.
If you understand that, you then know you can add processes (such as operators or the cluster autoscaler).
It’s a mindset shift. Thinking of it as “using” kubernetes, as if it is a monolithic thing that you directly control, will greatly increase the difficulty in understanding and reasoning through what’s going on within a Kubernetes cluster.
The complaints are real, because in practice a company needs both aspects and when a small company struggles to setup and manage the kubernetes infrastructure correctly, the application operators are suffering the consequences (e.g. log collection infrastructure doesn't work, it's hard to provision nodes with the right capacity, things like that). They see the infrastructure operation team struggle and they partake in that struggle because what is advertised to be easy, it's not.
That said, K8s is a very good way to build an "API" between infrastructure and application "teams". It can work very well if the people involved set things and processes up correctly. It can be a nightmare if botched up
Agreed.
Someone has to actually build it with a product mindset — including product-market fit. I’m one of those oddballs that have set up and scaled up infra, and have put together and delivered applications before.
In one gig, I had people joke about naming what I had setup after Heroku. I didn’t realize its significance until I came to a place where it was not done this way. Many on the application team have expressed dissatisfaction… but it’s like the infra team is oblivious to that.
Switching your database, just like switching your cloud provider, rarely happens in practice.
As I understand it this new Nvidia VM image comes with Kubernetes on the inside so to speak, perhaps microk8s with nvidia extension enabled.
BTW this is how I’ve started running my own little AI experiments. Sure, there’s some overhead. But compared with constantly downloading new versions of drivers it’s quite lightweight. Also K8S is turning into the ligua franca of sodtware platforms, so well worth learning and paying the overhead on IMHO.
Uh huh.
> Nvidia will pay $5.5 million to settle charges that it unlawfully obscured how many of its graphics cards were sold to cryptocurrency miners...
And
> The CMP HX is a pro-level cryptocurrency mining GPU that provides maximum performance...
Just a quick google away.
Nvidia will develop and sell whatever will make Nvidia more money. They just think the world of AI is two or three orders of magnitude more lucrative than mining ever was. Hence the maximum push on the AI front.
Nvidia knows their biggest revenue sources today, which are growing, and is investing into their business units based on that data.
It’s just smart business.
This is because they didn't serve the market... so they didn't understand how many buyers were coming from crypto.
[1] https://www.nvidia.com/en-us/cmp/ [2] https://arstechnica.com/gadgets/2021/05/nvidia-will-add-anti...
I have no idea if they were successful at achieving that goal, just thought it was interesting that market differentiation wouldn't just be useful for marketing but also for corporate accounting. They would even risk alienating the crypto market and possibly lose revenue, if it would mean they'd get a better handle on what they were selling to whom.
And most of the time, the person isn’t even breaking any rules. In this case, I’m pretty sure they were making a joke and didn’t actually think that a longtime HN user was astroturfing
Isn’t this very recent?
Even if they believe in a technology because they believe they can deliver a profitable product (and reject something else because they think there’s no long term gains), I still prefer that to a company which would blindly try to profit from everything short term.
Given some of the commentary in launch reviews for the 4000 series, I wouldn't be surprised if the overwhelming opinion was that they already are.
But a monopoly can be harmful for a market without anyone doing anything illegal.
Cloud customers can afford to pay more for those GPUs than gamers because they generate revenue with them, gamers don't.
So it make sense to have some product segmentation in place to prevent one market completely cannibalizing the other while leaving Nvidia with less profits.
The current situation is still caused by manufacturing constraints at TSMC for the cutting edge nodes which both the consumer and data center parts occupy so it makes sense for Nvidia to prioritize the higher margin parts.
There have been great points made that Nvidia should split into Nvidia, the general compute company oriented to data center customers with deep pockets, and in GeForce, the gaming GPU company with access to all the cutting edge tech of Nvidia but seeks to be more scrappy and optimize designs for rasterization performance rather than generic compute and chases smaller die sizes on cheaper nodes to be price competitive. This way the data center compute market will stop cannibalizing consumer gaming one and we'll be back to having better GPUs at competitive prices.
But the real issue is physical form factor and power. As has been noted in the press, etc, something like an RTX 3090 (and more so 4090) is literally designed to push frames as fast as possible power and heat be damned. They're multi-slot (which results in poor density), have card design/cooling challenges, power configuration issues, etc.
There's a story out there about the only dual-slot RTX 3090. Gigabyte came up with one (I have several - they're great) but supposedly Nvidia put pressure on them to pull them from the market[0] because people were putting them in x8 server configurations and using them instead of their much more expensive datacenter products.
[0] - https://www.tomshardware.com/news/gigabyte-rains-partners-pa...
This is how they also came out on top from the crypto craze without destroying their gaming market.
An individual developer is happy to charge a higher salary for its services from a larger corporation in comparison to working for an SME, simple because in a large org its services generate more value, allowing it to capture more of it.
Things can still get a lot worse: The fight isn't over until all roads are toll roads and you have to pay for the oxygen you consume.
NVAIE license is what nvidia wants enterprises to pay for using their bespoke cards in shared VRAM configuration by knee capping consumer cards which can very well do the same job better with more cuda cores but lesser memory.
And don't even get me started on RIVA stack
FP8 emulation is also never going to get backported instead only H100 & 4090s can make use of it
RIVA: NVIDIA® Riva, a premium edition of NVIDIA AI Enterprise software, is a GPU-accelerated speech and translation AI SDK
FasterTransformer: https://github.com/NVIDIA/FasterTransformer an highly optimized transformer-based encoder and decoder component, supported on pytorch, tensorflow and triton
TensorRT, custom ml framework/ inference runtime from nvidia, https://developer.nvidia.com/tensorrt, but you have to port your models
Thanks, I came here to see whether anything had changed since I last did ML stuff on Nvidia GPUs, and it looks like things are still the same.
An RTX 4090 has over 16,000 cores and 1 TB/s of memory bandwidth. From what I understand (not really my thing) DDR5 tops out at 51 GB/s per module.
CPUs and GPUs are so fundamentally different architecturally but for extremely parallel tasks GPUs are designed for CPU is very, very far behind.
When I've done performance tests between CPU and GPU for my applications (speech) a $100 six year old GTX 1070 is 5x faster than a AMD Ryzen Threadripper PRO 5955WX[0] while consuming a fraction of the power and cost. If you look at the table the RTX 3090 and RTX 4090 are 17x and 27x respectively. The H100 benchmark of 12x is from a very early access benchmark with some driver and other issues.
[0] - https://github.com/toverainc/willow-inference-server/tree/wi...
I have 20 interviews coming down the pipe -- all of which have highly tactical / near term valuable ideas like this.
If you don't see the potential of the tech and the rapid advances, I can't help you. But the issue around deployment is more legal (and perhaps not enough GPUs to go around).
Not sure it matters. We’re very much in do first ask permission later territory here and nobody is putting the genie back in the bottle.
The legal will have to bend towards reality
major: chatgpt for answering questions, explaining topics and helping with coding has brought me personally massive ROI
minor: a lot of companies are integrating LLMs to upgrade their offerings and a lot of small SaaS now exist due to LLMs. I have to guess at least some of those have a positive ROI
I was in an 800-level (PhD) course last semester, and the professor made a fun lecture where each student had to present a paper from the last 5 years that’s been completely outdone by GPT4. You wouldn’t believe how it casually outperforms the state of the art from just 5 years ago. My paper was about natural language to bash commands. GPT4 is lightyears ahead of the previous state of the art. You could probably make a business off just a natural language interface to the Linux operating system.
There are lots of use cases, we seem to only talk about LLM's recently.
Indeed, I was thinking mostly about LLMs, as it seems to me that this type of news presented in the article is mostly targeting that field.
But this particular data is air gapped.
Is the cost AWS level of waste - or something reasonable?
I can get an A4000 with 16GB vram which can run some models for 140$ per month.
I can't say the setup is anything special really but not having to do that has some value