HNHacker News
TopNewBestAskShowJobs

cheptsov

190 karma · joined June 20, 2017

Building https://github.com/dstackai/dstack
submissionscomments
cheptsov··on That's a Lot of YAML
That’s the problem people IMO don’t understand - flexibility is one of many requirements; simplicity and ecosystem around are two more.

Hating YAML is same as being childish - yo using like something but can’t come up with anything better. Worse than YAML can be only complaining about YAML. Hating YAML is a cliche.

Its much easier to come up with jokes like this website than to change the ecosystem for better.

And I definitely think the websites like this is opposite to the engineering approach.

Not defending YAML though in particular.

Edit: grammar

cheptsov··on That's a Lot of YAML
Next time the author would like to write an article like that, it’s better that they offer a new standard that is as flexible while also doing the work YAML is doing today
cheptsov··on The turbulent AI era is here
What surprises me even more is why Bill Gates writes this kind of bizarre, low-quality nonsense in the first place. What purpose does it serve? It’s so obviously shallow that it influences nothing and adds nothing to the conversation.
cheptsov··on The turbulent AI era is here
Is it me or does it read like pure satire?

An AI ministry, “human-reserved” jobs, and taxes on AI tokens.

cheptsov··on Show HN: Skill that lets Claude Code/Codex spin up VMs and GPUs
Nice, we build something similar at dstack

We recently also added support for agents: https://skills.sh/dstackai/dstack/dstack

Our approach though is more tide-case agnostic and in the direction of brining full-fledged container orchestration converting from development to training and inference

cheptsov··on Agentic Development Environment by JetBrains
Finally a step in the right direction. This brings the best of two worlds: the lightweightness of Fleet and agents battle-tested with Junie/IntelliJ.

Congrats to the team. Can’t wait to try it.

cheptsov··on Simulating hand-drawn motion with SVG filters
Looks very cool but my iPhone got really hot after just playing with it for one minute!
cheptsov··on TPU Deep Dive
It’s so ridiculous to see TPUs being compared to NVIDIA GPUs. IMO proprietary chips such as TPU had no future sure to the monopoly on the cloud services. There is no competition across the cloud services providers. The only way to access TPUs is through GCP. As the result nobody wants to use them regardless of the technology. This is the biggest fault of GCP. Further the road, the gap between NVIDIA GPUs and Google TPUs (call it „moat“ or CUDA) is going to grow.

The opposite situation is with AMD which are avoiding the mistakes of Google.

My hope though is that AMD doesn’t start to compete with cloud service providers, e.g. by introducing their own cloud.

cheptsov··on Show HN: AIButton – Like AI Pin – but only one button to press, Made in Germany
Wow, such a cool idea! Where can I read more about the product, the release ETA, price, etc?
cheptsov··on Ironwood: The first Google TPU for the age of inference
Yea, makes sense
cheptsov··on Ironwood: The first Google TPU for the age of inference
I believe my original sentence was accurate. I was expecting the article to provide an objective comparison between TPUs and their main competitors. If you’re suggesting that El Capitan is the primary competitor, I’m not sure I agree, but I appreciate the perspective. Perhaps I was looking for other competitors, which is why I didn’t really pay attention to El Capitan.
cheptsov··on Ironwood: The first Google TPU for the age of inference
Are you suggesting NVIDIA is not a competitor?
cheptsov··on Ironwood: The first Google TPU for the age of inference
I think it’s not misleading, but rather very clear that there are problems. v7 is compared to v5e. Also, notice that it’s not compared to competitors, and the price isn’t mentioned. Finally, I think the much bigger issue with TPU is the software and developer experience. Without improvements there, there’s close to zero chance that anyone besides a few companies will use TPU. It’s barely viable if the trend continues.
cheptsov··on The Llama 4 herd
https://www.llama.com/llama4-reasoning-is-coming/
cheptsov··on Nomadic infrastructure design for AI workloads
A nice project—I hadn’t seen it before. I’m working on a similar concept but as an open-source project [1].

We see it as an alternative to K8s/Slurm and would be interested in hearing thoughts from other ML engineers.

[1] https://github.com/dstackai/dstack

cheptsov··on Orchestrating GPUs in data centers and private clouds
Thank you! Yup, dstack can also work over K8S too if required. But of course there are many advantages to use dstack’s native orchestrator.
cheptsov··on Orchestrating GPUs in data centers and private clouds
Hey HN, founder of dstack here. We've been working on this over three months and pretty excited about this release. Basically, the main point is that dstack is an open-source AI-native alternative to Kubernetes, designed to be more lightweight, and focusing just on AI workloads on both cloud and data-centers. With this release we are adding the critical feature that allows to run containers concurrently on same host slicing its resources incl. GPU for a more cost-efficient utilization. Another new thing is the simplified way to run things on private clouds where clusters are often behind a login node. There are many more cool things on our roadmap to ensure dstack is a streamlined alternative to both K8S and Slurm. Our roadmap can be found in [1] Super excited to hear any feedback.

[1] https://github.com/dstackai/dstack/issues/2184

cheptsov··on Exploring inference memory saturation effect: H100 vs. MI300x
In this one we were only using 3.1 405B FP8. We took one model to simplify the setup and were mostly looking at the memory saturation effect. So basically we compared inference metrics of the same model. I suppose comparing 3.1 and 3.2 will be difficult as they are different models entirely. But open to ideas
cheptsov··on Exploring inference memory saturation effect: H100 vs. MI300x
Yes, we should have included the price too—thanks for pointing that out. I forgot about it. Appreciate you adding the links here! And yes, Lambda's and Hot Aisle's prices were used in the calculation.
cheptsov··on Dstack: An alternative to k8s for AI/ML tasks
Thank you for the question!

> What are the differences in opinions between dstack and SkyPilot?

SkyPilot is great. I think there are many tiny details though. At dstack, we try to provide out-of-the-box and more high-level experience.

Examples:

1. Authorization built-into services

2. Dev environments with IDE integration

3. HTTPS out of the box with an ability to set up own domains

4. Projects for team management and resource isolation

5. Hardware metrics tracking

Also, we try to distance from Kubernetes and improve our own orchestrator that natively integrates with cloud providers

> Same question could be posed to Modal.

Modal is great too. Modal's strengths is Python decorators and their focus on cloldstarts/serverless kind of experience.

dstack here is more about flexibility/multi-cloud/on-prem/etc. For example, I personally prefer being able to run any code with dstack without changing my code. Otehr people may prefer Python decorators.

> Aren’t you worried about the amount of product development that goes into Kubernetes and its ecosystem?

A very good question. I think its both a strength and a weakness of Kubernetes. So far we see that for us it's a lot easier to bring AI-native experience, simplify it, and make it more out of the box.

We of course respect K8S though. But we think the community to deserve options!

> What about democratic-csi?

Haven't seen it yet. Will look into it. At dstack, we support volumes for both cloud and on-prem.

> NVIDIA GPU Operator, AMD’s k8s device plugins

We aim to support any accelerators out of the box.

> Calico versus Cilium?

Haven't heard of it yet.

P.S.: Whould love to hear your opinion too!

cheptsov··on Dstack: An alternative to k8s for AI/ML tasks
I guess it depends on the use case. For example with dstack, we focus on AI.

Our abstractions include:

1. Dev environments - you need them often and need an easy way to get one with tight GOU resources - either using already provisioned resources or provision on-demand

2. Tasks. For example, in AI you may want to run distributed tasks over a cluster using your favorite framework like pytroch

3. Services - very close to Docker Compose. And you can use it with dstack. But for you may want to also manage GPU requirements; and of course auto-scaling

4. Managing clusters. As an AI user you may want to provision them on-demand. This is what we call fleets with dstack.

5. Ingress for public endpoints. Dstack also handles authorization and OpenAI endpoint mapping - as it’s important for AI.

6. Finally you need to manage tenancies - isolate resources across projects or teams. With dstack, we call it projects.

cheptsov··on Dstack: An alternative to k8s for AI/ML tasks
> Any explanation for how it simplifies orchestration.

There are many examples in the docs. Here are few:

1. ALl accelerators are supported natively (no operators are required).

2. Distributed task work out of the box [1]

3. For inference, OpenAI compatible gateway is provided automatically for any models deployed; along with authentication. [2]

4. Cluster management is a lot more convenient for AI compared to K8S [3]

More than anything else, dstack is more lightweight; no one need to know K8S.

I'm biased but we talk to a lot of users that choose dstack because they don't want to deal with K8S

>dstack supports NVIDIA GPU, AMD GPU, and Google Cloud TPU out of the box.

AWS accelerators are in the plan of course!

[1] https://dstack.ai/docs/reference/dstack.yml/task/#distribute...

[2] https://dstack.ai/docs/services

[3] https://dstack.ai/docs/concepts/fleets/

cheptsov··on Dstack: An alternative to K8 for AI/ML tasks
Well,

1. No changes in the code required; Works out of the box with any Docker image; Incl. support for distributed training

2. Out-of-the-box support for AMD/NVIDIA/TPU

3. Multi-cloud, incl. Neoclouds such as Lambda, RunPod, TensorDock, and more to come

4. 5 min to set up your own on-prem cluster

5. Easye to combine multi-cloud + on-prem

6. On top of task, you get dev environments and model inference

cheptsov··on Dstack: An alternative to k8s for AI/ML tasks
Oh, excited to see dstack featured here. Founder and core contributor to dstack here. Yes, we aim to simplify container orchestration for AI and build an alternative to both K8S and Slurm - for both multi-cloud and on-prem.

Would love to hear feedback!

cheptsov··on We're Leaving Kubernetes
I can completely relate to anyone abandoning K8s. I'm working with dstack, an open-source alternative to K8s for AI infra [1]. We talk to many people who are frustrated with K8s, especially for GPU and AI workloads.

[1] https://github.com/dstackai/dstack

cheptsov··on Notable AI Models
Yes, to me, the report is interesting not because of the insights it offers, but because it raises more questions:

1. What about new AI chip vendors?

2. How will the price of compute change?

3. How will the demand for compute change?

4. How will the overall supply of chips change?

cheptsov··on Benchmarking Llama 3.1 405B on 8x AMD MI300X GPUs
Unfortunately, neither VLLM nor TGI support FP8 on AMD yet. But once they do, we will look into it.
cheptsov··on Benchmarking Llama 3.1 405B on 8x AMD MI300X GPUs
TL;DR

- We explore how the inference performance of Llama 3.1 405B varies on 8x AMD MI300X GPUs across vLLM and TGI backends in different use cases.

- TGI is highly efficient at handling medium to high workloads. In our tests on 8x AMD MI300X GPU, medium workloads are defined as RPS between 2 and 4. In these cases, it delivers faster time to first token (TTFT) and higher throughput.

- Conversely, vLLM works well with lower RPS but struggles to scale, making it less ideal for more demanding workloads.

- TGI's edge comes from its continuous batching algorithm which dynamically modifies batch sizes to optimize GPU usage.

If you have feedback, or want to help improve the benchmark, please let me know.

cheptsov··on Thoughts on the Durov Arrest
“Summing up: for the time being, if you run a social media company, or if you provide encrypted messaging services, which are accessible in France, and you’re based in the United States, get out of Europe.”
cheptsov··on Cerebras launches inference for Llama 3.1; benchmarked at 1846 tokens/s on 8B
Very interested in playing with their hardware and cloud. Also I wonder if it’s possible to try cloud without contacting their sales.
Page 1 of 5Next →