101 karma · joined February 16, 2019
From the creator's perspective, it's just a more powerful agentic CustomGPT that they can easily monetize. One of them has made more than $1600 so far, so I think she's happy.
Time will tell of course how big a pain point this really is, but I've seen this being useful for a handful of creators already.
Documentation link to try it out - https://docs.spotml.io/getting-started
Looking for feedback from early testers. You would be an ideal candidate if you have a side project that you're spending your own money to train.
Acknowledgement: - SpotML is built on top of existing open source library Spotty: https://github.com/spotty-cloud/spotty
Documentation link to try it out: https://docs.spotml.io/getting-started
Looking for feedback from early testers. You would be an ideal candidate if you have a side project that you're spending your own money to train.
Acknowledgement: - SpotML is built on top of existing open source library Spotty: https://github.com/spotty-cloud/spotty
The biggest advantage which was missing in the Open source options was monitoring on the training job and auto recovery from spot interruptions which spotML does.
I've read through the docs, the one difference that comes to my mind is the automatic fallback to on-Demand and resume back to spot when available. I can't readily see a way to do this yet in Metaflow, but it's possible I've missed something.
Have you used it btw ? and what has your experience been with Grid ?
Also i'm not sure how straightforward it is to detach/attach persistent volumes to retain data across different spot interruptions ? The latter can be done but it's just the same rote each time you wanna train something new.
Also thanks for the suggestions ! We're a team of 2 right now, I used to be in the bay area but in Mexico temporarily.
For detecting if the training process is still running or errored out it registers the training command pid when launching the task and then monitors for the Pid for completion. It also registers and monitors the instance state itself to check for interruptions and resuming.
Also, interesting point about inference. I'm not sure though how common it is for companies to need GPUs for inference. Because if you can have a CPU based inference model, which I thought was most common, it's probably not a big usecase?