HNHacker News
TopNewBestAskShowJobs

covi

201 karma · joined September 5, 2012

https://github.com/skypilot-org/skypilot/
submissionscomments
covi··on Scaling Karpathy's Autoresearch: What Happens When the Agent Gets a GPU Cluster
This feels like the chimpanzee with a power drill. An agent is honestly just brute-force search, but guided.
covi··on Migrating from Slurm to Kubernetes
The post says Slurm supports gang scheduling, k8s doesn't (out of the box).
covi··on Cloud Run GPUs, now GA, makes running AI workloads easier for everyone
Take a look at SkyPilot. Good for running these batch workloads. You can use spot instances to save costs.
covi··on Cloud Run GPUs, now GA, makes running AI workloads easier for everyone
To massively increase the reliability to get GPUs, you can use something like SkyPilot (https://github.com/skypilot-org/skypilot) to fall back across regions, clouds, or GPU choices. E.g.,

$ sky launch --gpus H100

will fall back across GCP regions, AWS, your clusters, etc. There are options to say try either H100 or H200 or A100 or <insert>.

Essentially the way you deal with it is to increase the infra search space.

covi··on Bulk Object Storage data migration with SkyPilot
Related: https://skyplane.org/en/latest/ (mentioned in OP)

From what I know this idea underpins a few FAANG-level companies' data transfer systems. OP's value = a simple implementation of the idea that's OSS and applied to AI.

covi··on Show HN: Sonauto API – Generative music for developers
Congrats on the API launch (from SkyPilot)!
covi··on Show HN: Attaching to a virtual GPU over TCP
If you want to use your own GPUs or cloud accounts but with a great dev experience, see SkyPilot.
covi··on California suspends Cruise's autonomous vehicle deployment
Now just need a Waymo invite code :)
covi··on Mistral 7B
Cloud deployment docs: https://docs.mistral.ai/cloud-deployment/skypilot/
covi··on Hugging Face raises $235M from investors including Salesforce and Nvidia
https://www.forbes.com/sites/alexkonrad/2023/07/13/ai-startu...

> Its revenue run rate has spiked this year and now sits at around $30 million to $50 million, three sources said — with one noting that it had more that tripled compared to the start of the year.

covi··on Vicuna v1.5 series, featuring 4K and 16K context, based on Llama 2
Part of the Vicuna team wrote a guide on finetuning Llama2: https://blog.skypilot.co/finetuning-llama2-operational-guide...
covi··on Run Llama 2 uncensored locally
Related ongoing thread: https://news.ycombinator.com/item?id=36975245
covi··on Cookbook: Finetuning Llama 2 in your own cloud environment, privately
You could finetune then add on a retrieval step, which has the advantage of citing sources. Jury's probably still out on which, or which combination of, methods work best. Likely use case and data size-dependent.
covi··on Meta’s Microservice Architecture [pdf]
Meta's infra and system teams have been on a roll. Feels like Google in the 2000s.

Check out quite a few interesting papers from them in OSDI 2023 (top systems conference, co-located with ATC where the OP URL links to): https://www.usenix.org/conference/osdi23/technical-sessions

covi··on UC Berkeley's open-source Vicuna LLM chatbot released new improved model weights
Twitter announcement: https://twitter.com/lmsysorg/status/1646618380114468864
covi··on Free Dolly: First truly open instruction-tuned LLM
Kudos to Databricks! Anyone has insights into benchmark & real-world quality?

From https://huggingface.co/databricks/dolly-v2-12b#benchmark-met..., it seems like dolly-v2-12b's benchmark results are actually slightly worse than dolly-v1-6b.

A commercially viable instruction-tuned LLM is still a huge deal.

covi··on Show HN: Tabby – A self-hosted GitHub Copilot
This is awesome, and glad to see that SkyPilot is useful in distributing Tabby on any cloud! https://github.com/TabbyML/tabby/blob/main/deployment/skypil...
covi··on Vicuna: An open-source chatbot impressing GPT-4 with 90% ChatGPT quality
From the post:

> The training was done with PyTorch FSDP on 8 A100 GPUs in one day.

> We employ SkyPilot managed spot to reduce the cost by leveraging the cheaper spot instances with auto-recovery for preemptions and auto zone switch. This solution slashes costs for training the 7B model from $500 to around $140 and the 13B model from around $1K to $300.

So, this is using for example a2-ultragpu-8g (8x A100-80GB) on GCP using spot instances. You can use SkyPilot to quickly see the price is $12.8 per hour (~$307 for a day):

» sky launch --gpus A100-80GB:8 --use-spot

Check out detailed CLI instructions and SkyPilot YAMLs here if you want to give it a try:

- https://github.com/lm-sys/FastChat#vicuna

- https://github.com/lm-sys/FastChat/blob/main/scripts/train-v...

covi··on UC Berkeley launches SkyPilot to help navigate soaring cloud costs
We don't anticipate this will happen for a few reasons. When the usage of SkyPilot (or a SkyPilot-like "intercloud broker" system) is small, it probably doesn't warrant the dominant clouds' attention.

When the usage gets bigger, I'm not sure how providers can restrict access anyway (curious if there are precedents). There are quite a few large multicloud platforms like Snowflake or Databricks heavily utilizing AWS/GCP/Azure already. (Granted, these platforms are not meta-cloud, in the sense of moving their customer workloads transparently across clouds.)

Ultimately we see such a system to grow the pie for the whole cloud market. The incumbents' relative shares may drop, but their absolute volume will grow.

covi··on UC Berkeley launches SkyPilot to help navigate soaring cloud costs
SkyPilot devs here, happy to answer any questions.

GitHub repo (Apache 2 license): https://github.com/skypilot-org/skypilot

Getting started is easy:

$ pip install "skypilot[aws,gcp,azure]" # Pick your clouds

$ sky check

$ sky launch

covi··on UC Berkeley launches SkyPilot to help navigate soaring cloud costs
One big reason people use multiple regions or clouds: higher resource availability.

Allocating scarce resources can be very hard on the cloud. These are things like high-end GPUs (both on-demand and spot; the latter being much harder to get) and also beefy CPU-based instances. We've seen 10s of hours of waiting times or longer.

The natural solution is to have the flexibility to allocate resources in multiple regions (and ultimately, clouds) to increase the total pool size.

covi··on UC Berkeley launches SkyPilot to help navigate soaring cloud costs
+1. We've heard from some heavy users that Cloudfare R2 is saving them $$$ on egress costs: https://www.cloudflare.com/products/r2/

As outlined in the position paper (linked by another commenter) we believe such tailwinds are increasingly helping foster the "Sky" and making workloads moving between clouds much easier.

covi··on UC Berkeley launches SkyPilot to help navigate soaring cloud costs
(SkyPilot dev) Can't agree more on autostop / autodown. Subjective to measure but save our own costs massively.
covi··on UC Berkeley launches SkyPilot to help navigate soaring cloud costs
I'm one of the creators of SkyPilot. Thanks for the thoughtful questions and let me try to take a stab:

SkyPilot is not just for multi clouds. It's useful for all of these scenarios:

- using a single region of one cloud

- using multiple regions of one cloud

- using multiple clouds

Data transfer between zones/regions within a cloud is much cheaper than across clouds. We see many users falling in the "one cloud" category and they frequently read 10s of TBs of data across regions to do ML training.

Finally, saving money is one of several key problems we aim to solve, and there are quite a few ways to save other than lots-of-compute-on-small-data. Other reasons why you may want to use a system like SkyPilot include

(1) improving resource availability (big pain point for GPUs/TPUs)

(2) use one interface and know that your jobs can migrate across regions or clouds

More rationale in the intro blog post: https://medium.com/@zongheng_yang/skypilot-ml-and-data-scien...

covi··on UC Berkeley launches SkyPilot to help navigate soaring cloud costs
Indeed! Having lower-cost GPU clouds in the "Sky" is on our immediate roadmap: https://github.com/skypilot-org/skypilot/blob/master/ROADMAP...

In fact, as we speak we're working with folks at Lambda Labs to add support for their cloud. If other providers are interested, we'd be happy to chat.

(SkyPilot dev here)

covi··on Volkswagen's U.S. diesel emissions settlement to cost $15B
Now's better than ever to buy a new Volkswagen. Because of the scandal, I've found crazy deals (~30-40% off MSRP) for new VWs.
covi··on Unicorn: a simple and flexible abstraction of BigTable-like databases
This was my very first reaction as well. The Unicorn paper was published in a high-profile conference, so as a N=1 sample I'd say it's famous in the systems community.
covi··on Taskwarrior – intelligent TODO list
You can access it on desktop/laptop too.
covi··on Taskwarrior – intelligent TODO list
As a Google Inbox user, I've found its reminder features to be very good for these kinds of day-to-day tasks.
covi··on TensorFlow: Large-Scale Machine Learning on Distributed Systems (2015) [pdf]
More C++ support is on our roadmap:

https://github.com/tensorflow/tensorflow/blob/master/tensorf...

Page 1 of 3Next →