P.S. I am not the author: just found the link in the Llama 4 thread.
8 karma · joined January 18, 2015
P.S. I am not the author: just found the link in the Llama 4 thread.
@kkielhofner thanks a lot! now I realize it. I see, there is even GRPC support in Triton, so it make sense.
one NVIDIA T4 GPU, 16 GB RAM, and, this is an EC2 instance, it means "install anything" all this for $0.526 /Hour
do you see any hidden gotchas?
Let us discuss it here -> https://news.ycombinator.com/item?id=41768088
Well, the categories I use - do not overlap at all with the list of 1092 categories in Google Content Categories.
> it handles other classifications as well
hm... I highly doubt that. First of all - I do not see API to upload list of MY categories. Second: Can somebody with Google Cloud account try it? I have no account and when creating it - it asks for credit card...
> you can always run a zero shot pipeline in HF with a simple Flask/FastAPI application.
Yeah, sometimes things that are right in front of your nose, you don't see them. you mean this? https://huggingface.co/docs/api-inference/index
Is ONMX runtime + OpenVINO [2] a good idea ? Seems easier to install and to use: Pre-built Docker image and Python package... Not sure about performance (the hardware-related performance improvements - they are in OpenVINO anyway, right?).
[1] https://github.com/triton-inference-server/openvino_backend
[2] https://onnxruntime.ai/docs/execution-providers/OpenVINO-Exe...
Learning a new programming language or database? let's write a commenting system! )))
Probably I have to rephrase the question to something like "Alternatives and side effects of Hetzner CX11 ?"
And, absolutely unexpected, I saw "Russian". 3 times.
Would you please explain?
P.S. I'm a native Russian speaker.