HNHacker News
TopNewBestAskShowJobs

vishnukool

101 karma · joined February 16, 2019

A geek
submissionscomments
vishnukool··on Show HN: Cheapest Managed OpenClaw hosting – 30s setup (agent37.com)
Love it!
vishnukool··on Show HN: Claude Skills Marketplace – search and try Claude skills instantly
Thanks!
vishnukool··on Show HN: Claude Skills Marketplace – search and try Claude skills instantly
I think github stars is a reasonable signal, but it would be nice to augment it with a signal of usage of this skill on a sandbox to rank certain one's higher.
vishnukool··on [dead]
Yeah, vercel does something similar with install data and ranks them higher based on it. I currently don't have such data but maybe down the line, we can introduce metrics around, quality signals from the users of this app.
vishnukool··on Show HN: Claude Skills Marketplace – search and try Claude skills instantly
Well, we rely on github stars as a signal of quality on top of that, we don't consider exactness match since it can be pretty misleading. A better signal is that we vector embed the name and description of a skill and find proximity match.
vishnukool··on Show HN: Claude Skills Marketplace – search and try Claude skills instantly
Thank you, let me know of any feedback once you try it out :)
vishnukool··on Show HN: Agent37 – Monetize your Claude skills with shareable links
So far the creators I've worked with are all non-technical. They really just wanted to solve a problem in their niche. In the process of building it for them I kind of naturally biased towards using skills + Claude Agent SDK, to solve it for them since in my personal experience, it's by far the most capable Agent as of today.

From the creator's perspective, it's just a more powerful agentic CustomGPT that they can easily monetize. One of them has made more than $1600 so far, so I think she's happy.

vishnukool··on Show HN: Agent37 – Monetize your Claude skills with shareable links
Well, my hypothesis is, that this is a significant enough pain point. And in a way the distribution problem is amplified by the fact that currently there's no easy way to share Claude skills to your audience without a lot of setup friction on the clients part.

Time will tell of course how big a pain point this really is, but I've seen this being useful for a handful of creators already.

vishnukool··on Show HN: SpotML – Pip Package for Managed ML Training on AWS Spot Instances
Deep learning is expensive and even if you have AWS/GCP credits, you'll quickly run out of it. We built the tool to make training cheaper and to have these credits longer. SpotML is a command line tool that automatically manages ML training on AWS spot instances which are ~3X cheaper. It lets you handle spot interruptions by resuming training using the latest checkpoint.

Documentation link to try it out - https://docs.spotml.io/getting-started

Looking for feedback from early testers. You would be an ideal candidate if you have a side project that you're spending your own money to train.

Acknowledgement: - SpotML is built on top of existing open source library Spotty: https://github.com/spotty-cloud/spotty

vishnukool··on Show HN: Seamless ML training on AWS Spot instances
SpotML is a command line tool that automatically manages ML training on AWS spot instances. It lets you handle spot interruptions by resuming training using the latest checkpoint.

Documentation link to try it out: https://docs.spotml.io/getting-started

Looking for feedback from early testers. You would be an ideal candidate if you have a side project that you're spending your own money to train.

Acknowledgement: - SpotML is built on top of existing open source library Spotty: https://github.com/spotty-cloud/spotty

vishnukool··on Show HN: SpotML – Managed ML Training on Cheap AWS/GCP Spot Instances
Haven't used cortex.dev, but looking at the docs, I'd say primarily simplicity and ease of getting started quickly (<3 mins) to get it up and running. Also with cortex, It's not clear to me yet if spot instance fails, cortex can wait for next spot/replace it with onDemand, to keep training going automatically. If not, that would be the second difference.
vishnukool··on Show HN: SpotML – Managed ML Training on Cheap AWS/GCP Spot Instances
We'll be updating the documentation with examples in a few weeks, once we release it. No, it doesn't support distributed training right now.
vishnukool··on Show HN: SpotML – Managed ML Training on Cheap AWS/GCP Spot Instances
Thanks, Makes sense. It doesn't use any "code" from nimbo. The documentation and the design simplicity of the tool were the things that was appealing to us and adopted. The project itself was forked from spotty which has an MIT license.

The biggest advantage which was missing in the Open source options was monitoring on the training job and auto recovery from spot interruptions which spotML does.

vishnukool··on Show HN: SpotML – Managed ML Training on Cheap AWS/GCP Spot Instances
It uses a mounted EBS(Elastic Block Store) so all the checkpoints, data etc. is already in the persistence storage. This is simply be re-attached to the next spot/onDemand instance after interruption.
vishnukool··on Show HN: SpotML – Managed ML Training on Cheap AWS/GCP Spot Instances
Interesting, thanks, we weren't aware of Metaflow.

I've read through the docs, the one difference that comes to my mind is the automatic fallback to on-Demand and resume back to spot when available. I can't readily see a way to do this yet in Metaflow, but it's possible I've missed something.

vishnukool··on Show HN: SpotML – Managed ML Training on Cheap AWS/GCP Spot Instances
Thanks for the suggestion, yes makes sense.
vishnukool··on Show HN: SpotML – Managed ML Training on Cheap AWS/GCP Spot Instances
Thanks for pointing it out, We realize our mistake here. We also should've done proper attribution. Will be correcting this.
vishnukool··on Show HN: SpotML – Managed ML Training on Cheap AWS/GCP Spot Instances
That's right it will need aws credentials with access to create EBS volumes, S3 Bucket, spawn instances, etc. In addition to the "cli" a cloud service constantly monitors the progress of the Jobs(by registering the pid when launching it), and the instance states. So the billing will be based on the hours of training run and $ saved.
vishnukool··on Show HN: SpotML – Managed ML Training on Cheap AWS/GCP Spot Instances
Thanks for this, I wasn't aware of it. From reading the docs, I don't see anywhere if Grid automatically handles spot interruptions to resume from last checkpoint which was the main focus of our internal tool.

Have you used it btw ? and what has your experience been with Grid ?

vishnukool··on Show HN: SpotML – Managed ML Training on Cheap AWS/GCP Spot Instances
Yes we liked the elegance of both the tool and the docs. So it's very much inspired from it. I must also give credit to another great tool https://spotty.cloud/ from which this project was adopted.
vishnukool··on Show HN: SpotML – Managed ML Training on Cheap AWS/GCP Spot Instances
Thank you! Interesting, we actually tried aws batch ourselves. 1) How were you able to handle spot interruptions and resuming from the latest checkpoint ? 2) Not to mention fallback to OnDemand on spot interruptions 3) then switching back to spot from onDemand would also need additional process to be setup.

Also i'm not sure how straightforward it is to detach/attach persistent volumes to retain data across different spot interruptions ? The latter can be done but it's just the same rote each time you wanna train something new.

Also thanks for the suggestions ! We're a team of 2 right now, I used to be in the bay area but in Mexico temporarily.

vishnukool··on Show HN: SpotML – Managed ML Training on Cheap AWS/GCP Spot Instances
Sure, in our own startup, we used to spend roughly $1000 for training a StyleGAN model for a class and then additional latent space manipulation models. In the early days we recklessly wasted a lot of our AWS credits during experimentation. But later on with spot instances were able to bring it down to $250 to $300 per category class training which was more bearable as cost.
vishnukool··on Show HN: SpotML – Managed ML Training on Cheap AWS/GCP Spot Instances
We tried to keep it minimal. All you need to do is specify the format to resume last checkpoint in the spotml.yml file. So let's say your checkpoint files are saved as ckpt00.pt, ckpt01.pt, ckpt03.pt and so on. You can configure checkpoint regex file format ^ckpt[0-9]{2}$ and spotml resumes by picking the latest of it.

For detecting if the training process is still running or errored out it registers the training command pid when launching the task and then monitors for the Pid for completion. It also registers and monitors the instance state itself to check for interruptions and resuming.

vishnukool··on Show HN: SpotML – Managed ML Training on Cheap AWS/GCP Spot Instances
This is interesting, from what I've read AWS no longer recommends setting spot instance prices so that they manage it themselves. I wonder, if there's an actual advantage of avoiding spot interruptions by setting the spot price even higher than onDemand.
vishnukool··on Show HN: SpotML – Managed ML Training on Cheap AWS/GCP Spot Instances
Thanks for the feedback. The biggest upside is, if you have a long running training(say hours, days) and the spot training is interrupted. You probably don't want to manually monitor to check for the next available spot instance and kick start the training. SpotML takes care of that part. Also optionally, you can configure it to resume with an onDemand instance on interruption until the next available spot instance. In essence we try to do it make i) creating buckets/EBS ii) code to save to S3 in loop iii) Monitor for interruptions and resume from checkpoint parts easy.
vishnukool··on Show HN: SpotML – Managed ML Training on Cheap AWS/GCP Spot Instances
Hmm.. makes sense, yes. thanks!
vishnukool··on Show HN: SpotML – Managed ML Training on Cheap AWS/GCP Spot Instances
Thanks for the feedback, yeah at this stage we really just wanted to find out if we can build something useful for the community. Agreed on the pricing suggestion.

Also, interesting point about inference. I'm not sure though how common it is for companies to need GPUs for inference. Because if you can have a CPU based inference model, which I thought was most common, it's probably not a big usecase?

vishnukool··on Show HN: SpotML – Managed ML Training on Cheap AWS/GCP Spot Instances
Curious, did you eventually start using Sagemaker with spot instances or did you give up on it ? Also what would you say were the biggest pain points with Sagemaker?
vishnukool··on Show HN: SpotML – Managed ML Training on Cheap AWS/GCP Spot Instances
Thank you!
vishnukool··on Show HN: SpotML – Managed ML Training on Cheap AWS/GCP Spot Instances
Thanks, yeah. At this stage we really just want to validate if this was a real problem in the ML community. Down the line I suppose as we scale to handle multi instance training and other use cases, we could probably charge more, say a % of the cost savings in training.
← PreviousPage 2 of 3Next →