HNHacker News
TopNewBestAskShowJobs

cheptsov

190 karma · joined June 20, 2017

Building https://github.com/dstackai/dstack
submissionscomments
cheptsov··on Show HN: Dstack – an open-source engine for running GPU workloads
Thanks for sharing dstack on HN. I'm the creator of dstack and would love to hear your feedback. We're also working on creating a new standard for GPU orchestration, similar to Kubernetes.
cheptsov··on Learnings from fine-tuning LLM on my Telegram messages
We are building dstack.ai, an open-source tool that helps run anything on vast.ai and TensorDock. Happy to hear your feedback.
cheptsov··on Fine-tune your own Llama 2 to replace GPT-3.5/4
How do you fit Llama2-70b into V100? V100 is 16GB. Llama2-70b 4bit would require up to 40GB. Also, what do you use for inference to get 300+tokens/s?
cheptsov··on Databricks Strikes $1.3B Deal for Generative AI Startup MosaicML
Yup, we're already working on an experimental support for Lambda Labs! Ping us if you'd like to test it out. Perhaps we could show something already next week.
cheptsov··on Databricks Strikes $1.3B Deal for Generative AI Startup MosaicML
Congrats to the MosaicML and wish all the luck! If anyone is interested in MosaicML alternatives, check out dstack, we are building its OSS and cloud-agnostic alternative: https://dstack.ai. Disclaimer: I'm the founder and CEO at dstack.
cheptsov··on Show HN: Oblivus GPU Cloud – Affordable and scalable GPU servers from $0.29/hr
Looks really cool. Congrats on the launch! A quick question. Do you allow to build custom public images? We’d love to integrate dstack.ai with Oblivus.
cheptsov··on Dstack Hub
Not yet, but we are excited to announce that we will be recording a live hands-on session for DataTalks Club Open-Source Spotlight next week together with Alexey Grigorev. To stay in the loop, you can subscribe to either our Twitter [1] or DataTalks Club [2]. We'll let you know as soon as it's available!

[1] https://twitter.com/dstackai [2] https://twitter.com/DataTalksClub

cheptsov··on Dstack Hub
Hey everyone, I'm happy to release dstack Hub, an open-source tool that helps teams manage their ML workflows more effectively without vendor lock-in.

dstack Hub extend dstack [1] with workflow scheduling capabilities and user management. Here's how it works: run dstack Hub via Docker, use its UI to configure projects and cloud credentials, then pass the URL and personal token to the dstack CLI. Now, you can run workflows through the CLI and Hub will orchestrate them in the cloud on your behalf.

This is a beta release and we plan to continuously improve it. We'd love to hear your feedback and answer any questions!

[1] https://github.com/dstackai/dstack

cheptsov··on Launch HN: DAGWorks – ML platform for data science teams
Hey, the founder of dstack here. If I put it shortly, dstack allows you to define ML workflows as code and run them either locally or remotely (e.g. in a configured cloud). ML workflows here mean anything that you may want to do when you're developing a model - prepping data, training or finetuning a model, etc. The value - basically it automates running your workflows, without being dependant on any particular vendor. At the same time, you don't have to rewrite your Python scripts to use a particular API (because dstack is using YAML).

We aim to build the most easy tool to run ML workflows - without making you use the UI of any vendor, or hustling with Kubernetes, custom Docker images, etc.

MLFlow doesn't do what dstack does (automatic infrastructure provisioning) - unless you use Databricks.

cheptsov··on Ask HN: Co-Founder? Seeking Co-Founder?
SEEKING CO-FOUNDER | Dev tools/MLOps | Tech co-founder | US/EU/Worldwide | Seed

We aim at simplifying the MLOps stack. The first OSS version is out: https://github.com/dstackai/dstack. Plan to raise a seed round shortly.

Drop me a message at andrey at dstack.ai, if you’re interested.

cheptsov··on Ask HN: Do you like the #BuildInPublic culture?
I personally don’t think it’s a culture. I’d call it a marketing strategy. I can be wrong of course. Curious to hear what others think.
cheptsov··on Ask HN: Can you get started on GPU programming without having one?
There is a lot of options:

Large clouds (rather expensive): AWS, GCP, Azure - consider using spot instances that are cheaper

Free GPU platforms: Colab, Kaggle

ML platforms (cheap GPU): Jarvis Labs, Lambda Labs, TensorDoc, Genesis Cloud

In case you decide to go with AWS, check out dstack.ai, an open-source CLI to run anything on AWS GPU

cheptsov··on Tell HN: A Demo Day without the investors
An amazing idea! Applied as a speaker.
cheptsov··on Python 3.11.0 final
Would be great to have its supported by Conda. Anyone has an idea on when this might happen?
cheptsov··on Ask HN: Which are the most interesting books you have read in 2022?
I’m listening to Superfans by Pat Flynn now. It is about practical approaches to growing a fanbase or a user community.
cheptsov··on Ask HN: Where do you host images for your blog or landing pages?
Just curious, what do you use for hosting the blog or landing page itself?

For example, we use GitHub Pages for hosting the landing page. So for hosting images, we use GitHub itself and it mostly works.

For the blog, we currently use Substack, but thinking of maybe moving it to GitHub Pages too.

FYI, here is our repo: https://github.com/dstackai/dstack

cheptsov··on Show HN: Open-Source ReadMe Alternative
Sounds intriguing. Building beautiful docs is a super important and challenging tasks, especially for open-source dev tools! I currently use Mkdocs Material but I spent days customizing it for better appearance!

Our docs: https://docs.dstack.ai

cheptsov··on JetBrains invites developers to join the Fleet Public Preview Program
Already saw it in Toolbox App and installed ;-) Gonna give it a try.

Congrats on the launch, and kudos to the Fleet team!

cheptsov··on Show HN: Dstack – a command-line utility to provision infra for ML workflows
Hey! Sorry for late reply. Yes, it does destroy the resources once finished. Also, it allows you to use interruptible (aka "spot") instances efficiently.
cheptsov··on Show HN: Dstack – a command-line utility to provision infra for ML workflows
Yes, by default, it will copy the files. Or you could also tell dstack to use FUSE.

As I said in the comment above, we think of providing a high-level Python API. THis API will allow to avoid copying the artifacts as an initial step.

But of course would love to hear what you think! ;-)

cheptsov··on Show HN: Dstack – a command-line utility to provision infra for ML workflows
Thank you! That's a great question. First and foremost, dstack treats artifacts are 1st class citizens.

Here's the basic way of using them:

A workflow may produce output files to a local folder. In your workflow declaration, you can mark what folders to treat as output artifacts. Then, dstack would save output artifacts automatically, and you'll be able to reuse them via the unique name of the run, or a user tag assigned to this run.

All artifacts are stored in the S3 bucket that is configured for dstack. dstack is capable of syncing artifacts at start//end of the workflow or mount artifact folders via FUSE (of course if that is needed).

Each artifact is stored using the following path: <s3 bucket>/artifacts/<run name>/<job name>

A run can have multiple jobs, e.g. if it's a distributed workflow.

In future, we also think of providing a high-level Python API for accessing/storing artifacts.

Please share your thoughts and feedback!

cheptsov··on Show HN: Dstack – a command-line utility to provision infra for ML workflows
Thank you for the comment! Exactly, this is one of theses tasks, dstack is designed to help with. You basically write a simple Python script that does the job, tests that it works locally, and then run via dstack – asking for a more heavy machine.

The entire setup will take roughly 5-10 mins. Promise to add a tutorial with exact steps and code snippets soon.

In the future, we also plan to support distributed workflows. Actually, this is already supported by the design, we just need to implement a provider!

cheptsov··on Show HN: Dstack – Train models and manage data effortlessly
Hi there! I’m one of the developers of dstack. In case you are curious about the story behind dstack, why we decide to build it and how we made certain design decision, please make sure to read our blog post: https://blog.dstack.ai/p/how-dstack-works?s=w

Happy to answer questions!

cheptsov··on Spyder – a free and open source scientific environment written in Python
Thank you so much. I totally understand what you mean. And this is exactly what we try to improve. Anyway, always happy to hear feedback and pass it over to the team.
cheptsov··on Spyder – a free and open source scientific environment written in Python
Cool, I see the request! Will send you all the info the next week already.
cheptsov··on Spyder – a free and open source scientific environment written in Python
Thank you! Will send all the info within a couple of days.
cheptsov··on Spyder – a free and open source scientific environment written in Python
Julia is on the radar. I'd love some time in the future to include it into scope.
cheptsov··on Spyder – a free and open source scientific environment written in Python
Our current approach to is very similar to the native Jupyter experience where notebook contains multiple separate cells. You can manipulate these cells with shortcuts or mouse. It looks very similar to Jupyter or JupyterLab. At the same time, if you open a Python script, you'll be able to use cell delimiters (aka `#%%`).
cheptsov··on Spyder – a free and open source scientific environment written in Python
Hey, a PM of PyCharm here.

Justin case someone will find it useful and relevant. Our team is currently working on a brand new IDE (based on PyCharm) for data scientists.

A big part of what we're doing is convenient support for Jupyter notebooks within the IDE. This new IDE is currently under development but already provides the support for notebooks similar to how the native web Jupyter notebooks work – cells and outputs under each cell. The IDE's functionality will include SQL, Jupyter, Python, and R support.

If you'd like to try the early EAP builds of this new IDE with the new notebooks support, please fill out this short form: https://pages.jetbrains.com/pycharm-data-science-insiders

Once you’ve confirmed your participation, you’ll get a detailed email with instructions on how to download the early builds and details about how to share your feedback.

Will be also happy to answer questions if any.

cheptsov··on Show HN: dstack.ai – publish, track, and share data visualizations
Hi, a co-founder here. My name is Andrey. Actually, Plotly is our source of inspiration in a way. However, I think, the problem we are trying to solve lies a different direction. Just a few points to show how we are different: 1. We would like to support Plotly as one of the plot libraries (which is super great by the way). In addition we'd like to support many other libraries – anything that the user is using, incl. ggplot2, Matplotlib, Bokeh, etc. 2. We would like to combine the approaches Plotly offers in two different products (Chart Studio and Dash) into one and simplify the process of reporting data and the process of organizing dashboards. 3. Ideally, we'd like to help companies also organize and simplify their security around accessing the data. 4. We'd like to help build dashboards easier – by eliminating the needs of development entirely.

Just also want to mention, that our current offering and all other basic functionality that we will add, such as permission control, dashboards, etc will remain free. Anyway, it's just a start for us to really learn what our potential customers and users are doing and build that things that are currently missing – not blindly copying anyone's features. Having said that, I'd still like to thank you for your feedback. In case you have an idea of how we can help your teams collaborate around data, please let me know andrey at dstack.ai

← PreviousPage 4 of 5Next →