the dataset is open source and we plan to train an aesthetics picker on it but obviously have to do proper evals (with at least 1M data) to come to a reasonable conclusion.
1,949 karma · joined December 27, 2020
the dataset is open source and we plan to train an aesthetics picker on it but obviously have to do proper evals (with at least 1M data) to come to a reasonable conclusion.
Is there a clear example of this? E.g. is kubernetes inherently unable to start a pod (assuming the same sequence of events, e.g. warm/cold image with streaming enabled) under 500ms, 1s etc?
I am asking this as someone who spent quite a bit of time and wasn't able to bring it down 2s< mark, which eventually led us to rewrite the latency sensitive parts to use Nomad. But we are currently in a state where we are re-considering kubernetes for its auxilary tooling benefits and would love to learn more if anyone had experiences with starting and stopping thousands of pods with the lowest possible latencies without caring for utilization or placement but just observable boot latencies.
Which has a simple solution, release the model weights with a license which doesn't let anyone to commercially host them (like AGPL-ish) without your permission. That is what Stability.ai does it.
Test the example on Stable Cascade as well (latest open-weight stability model), and yeah, even that is not great at it https://fal.ai/models/stable-cascade?share=eab44060-690b-497....
[0]: https://twitter.com/EMostaque/status/1760660709308846135
It might work well if you have a single model with lots of customers, but as soon as you need more than a single model and a lot of finetunes/high rank LoRAs etc., these won't be usable. Or for any on-prem deployment since the main advantage is consolidating people to use the same model, together.
[0]: https://wow.groq.com/groqcard-accelerator/
[1]: https://twitter.com/tomjaguarpaw/status/1759615563586744334
This is one of the things we (at https://fal.ai) working very hard to solve. Because of ML workloads and their multiple GB environments (torch, all those cuda/cudnn libraries, and anything else they pull) it is a real challange just to get the container to start in a reasonable time frame. We had to write our own shared Python virtual environment runtime using SquashFS distributed thru a peer-to-peer caching system to bring it down sub-second mark.
After the container boots, there is the aspect of storing model weights, which IMHO less challenging since it is just big blobs of data (compared to Python environments where there are thousands of smaller files where each might be sequentially read and incur a really major latency penalty). Distributing them once we had the system above was super easy since just like squashfs'd virtual environments, they are immutable data blobs.
We are also starting to play with GPUDirect on some of our bare metal clusters and hopefully planning to expose it to our customers, which is especially important if your models is 40GB or higher. At that point, you are technically operating at the PCIE/SXM speeds which is ~2-3 seconds for a model of that size.