Show HN: Panini AI – A platform to serve ML/DL models at low latency
panini.ai
panini.ai
Given that you can horizontally scale model prediction infinitely the only sensible way to compare is to include price.
I agree that this looks compelling while it is free! But will it be price competitive later?
And if price competitiveness is claimed, then how is it possible? Yes, you can do the whole spot instance thing, but that is difficult to make reliable enough at scale.
You can always download the entire panini in your own private server and not pay anything. Ie. used Helm to install in your own kubernetes or DockerHub. For now, We're making it free for models under 2GB. Our main goal is to make it usable and we don't want cost to be a factor.
This seems surprising. What makes it so much faster?
Edit: Unless of course you are hitting the cache for a lot of the predictions?
In other words, if I had a model that previously took three seconds to get a response from would this platform respond in one second?
It is very hard to believe that deploying your models on GKE is going to be cost saving for anyone involved.
Google has also dumped a lot into tensorflow serve, so if you are outperforming it by that much it would be great to know how.
The site isn't very upfront about it, which is the sketchy part. Other than that, it looks much more straight forward than other options (I did watch the youtube tutorial). I like the idea, just question the motives.
If a user downloads panini to their private server and use it that will always be free since there is not infrastructure cost for us. If you're deploying it in our website we will be charging you to pay for the infrastracture cost.
Our main goal currently is to find out if people find this product useful and if it's worth for us to spend more time working on it. Thanks for watching the YouTube tutorial and if you have further questions, please contact us. Thanks
Then, you have caching. I actually fail to see how any caching at all is useful on a CPU bound task when you have unique inputs each time. This is just not something that is cacheable!
Batching may be one thing that can be helpful --- but typically requires deep modification of the model itself to support it, and no mention is made of that. Furthermore batching may help throughput but may make latency WORSE as you need to wait for multiple inputs before firing off a batch of computation.
Then you fail to specify whether your model will run on a GPU or CPU, and what type / core count thereof.
So, a lot of this just doesn't make much sense from a computer science perspective. Add in the free pricing with no limits and you've got a eyebrow-raising product!
2. It's up to you if you want to use it in GPU or CPU. Benchmark was done in a CPU but you're free to download panini via Helm and use GPU in your private kubernetes.
3. For now, during beta testing, we're offering free inference and there is a limit of model size cannot exceed over 2GB.
Hope this was helpful.
We have hundreds of models, across many domains, real estate, energy prediction, time series crypto, video analytics, molecular modelling.
I would bet money that across the millions of predictions that we make weekly, over all of the models, no two inputs are the same.
That’s kind of the point of Deep Learning - high dimensional noisy input
Caching will not help you here