It helps us figure out what got done and where we are in our roadmap
87 karma · joined February 24, 2016
It helps us figure out what got done and where we are in our roadmap
> If you need a few dozen inferences per second per server, this is the cheapest way. And you're not depending on a proprietary solution whose parent company could go out of business in a year.
Definitely the cheapest way.
We've been in business for more than a year already actually :)
The model size is the zipped size of your model that is uploaded to Inferrd (either through the SDK or the website).
I'll fix the full screen problem right away, thank you for reporting.
We only have servers in the United States at the moment but are looking to have servers all around NA and EU very soon.
We don't have any cold start delay! In our custom environment, you can do exactly what you are describing (running both CPU and GPU code). We provide you with access to the GPU and the CUDA libraries installed. It's basically lambda (minus the cold start) with GPU access.
We can scale a lot very quickly depending on how much you need.
I'll send you some documentation soon.
2. Yes! We will be releasing tokens soon so that anyone can interact with the API.
3. Models can be up to 1GB, each model gets their own server , we only do real-time predictions for now. Meaning the models are constantly running waiting for a request
It's a simple drag and drop to deploy tensorflow, scikit, spacy, keras or pytorch.
We stopped at 50 request/s in the pricing table because that seemed like a reasonable number for most use cases.
If you have a more custom pipeline, we have a custom environment where you can deploy any custom code with specific package versions!
In addition we support custom pre and post processor via custom environments! Simply write your inference code into a predict.py and we take care of the rest.
(Btw I am really into CML.dev, great idea)
This was in January right after the 16" came out.