190 karma · joined June 20, 2017
[1] https://twitter.com/dstackai [2] https://twitter.com/DataTalksClub
dstack Hub extend dstack [1] with workflow scheduling capabilities and user management. Here's how it works: run dstack Hub via Docker, use its UI to configure projects and cloud credentials, then pass the URL and personal token to the dstack CLI. Now, you can run workflows through the CLI and Hub will orchestrate them in the cloud on your behalf.
This is a beta release and we plan to continuously improve it. We'd love to hear your feedback and answer any questions!
We aim to build the most easy tool to run ML workflows - without making you use the UI of any vendor, or hustling with Kubernetes, custom Docker images, etc.
MLFlow doesn't do what dstack does (automatic infrastructure provisioning) - unless you use Databricks.
We aim at simplifying the MLOps stack. The first OSS version is out: https://github.com/dstackai/dstack. Plan to raise a seed round shortly.
Drop me a message at andrey at dstack.ai, if you’re interested.
Large clouds (rather expensive): AWS, GCP, Azure - consider using spot instances that are cheaper
Free GPU platforms: Colab, Kaggle
ML platforms (cheap GPU): Jarvis Labs, Lambda Labs, TensorDoc, Genesis Cloud
In case you decide to go with AWS, check out dstack.ai, an open-source CLI to run anything on AWS GPU
For example, we use GitHub Pages for hosting the landing page. So for hosting images, we use GitHub itself and it mostly works.
For the blog, we currently use Substack, but thinking of maybe moving it to GitHub Pages too.
FYI, here is our repo: https://github.com/dstackai/dstack
Our docs: https://docs.dstack.ai
Congrats on the launch, and kudos to the Fleet team!
As I said in the comment above, we think of providing a high-level Python API. THis API will allow to avoid copying the artifacts as an initial step.
But of course would love to hear what you think! ;-)
Here's the basic way of using them:
A workflow may produce output files to a local folder. In your workflow declaration, you can mark what folders to treat as output artifacts. Then, dstack would save output artifacts automatically, and you'll be able to reuse them via the unique name of the run, or a user tag assigned to this run.
All artifacts are stored in the S3 bucket that is configured for dstack. dstack is capable of syncing artifacts at start//end of the workflow or mount artifact folders via FUSE (of course if that is needed).
Each artifact is stored using the following path: <s3 bucket>/artifacts/<run name>/<job name>
A run can have multiple jobs, e.g. if it's a distributed workflow.
In future, we also think of providing a high-level Python API for accessing/storing artifacts.
Please share your thoughts and feedback!
The entire setup will take roughly 5-10 mins. Promise to add a tutorial with exact steps and code snippets soon.
In the future, we also plan to support distributed workflows. Actually, this is already supported by the design, we just need to implement a provider!
Happy to answer questions!
Justin case someone will find it useful and relevant. Our team is currently working on a brand new IDE (based on PyCharm) for data scientists.
A big part of what we're doing is convenient support for Jupyter notebooks within the IDE. This new IDE is currently under development but already provides the support for notebooks similar to how the native web Jupyter notebooks work – cells and outputs under each cell. The IDE's functionality will include SQL, Jupyter, Python, and R support.
If you'd like to try the early EAP builds of this new IDE with the new notebooks support, please fill out this short form: https://pages.jetbrains.com/pycharm-data-science-insiders
Once you’ve confirmed your participation, you’ll get a detailed email with instructions on how to download the early builds and details about how to share your feedback.
Will be also happy to answer questions if any.
Just also want to mention, that our current offering and all other basic functionality that we will add, such as permission control, dashboards, etc will remain free. Anyway, it's just a start for us to really learn what our potential customers and users are doing and build that things that are currently missing – not blindly copying anyone's features. Having said that, I'd still like to thank you for your feedback. In case you have an idea of how we can help your teams collaborate around data, please let me know andrey at dstack.ai