Cloudflare launches new AI tools to help customers deploy and run models
techcrunch.com
techcrunch.com
Regular Twitch Neurons (RTN) - running wherever there's capacity at $0.01 / 1k neurons
Fast Twitch Neurons (FTN) - running at nearest user location at $0.125 / 1k neurons
Neurons are a way to measure AI output that always scales down to zero. To give you a sense of what you can accomplish with a thousand neurons, you can: generate 130 LLM responses, 830 image classifications, or 1,250 embeddings.
Who came up with this? This is ridiculous. I understand the underlying issues but would still prefer a metric like seconds of utilization multiplied by the size of worker.
Besides this, the expected pricing doesn't talk about the expected pricing but just the pricing model. Have the feeling that this is not going to be competitive to platforms like Vast.ai
Quality will likely be heaps worse than chatgpt3.5, given it's llama 7b
It's 0.96$ per 100 fast chat responses It's 0.0076$ per 100 slow chat responses
Chatgpt 3.5 with 50 tokens input, 50 tokens output will give you 0.02$ per 100 fast responses If the llm responses are 500 tokens in and 500 tokens out then you get 0.2$ per 100 fast responses
I presume people will flock to the cheap version for when they can't afford the price and quality of chatgpt3.5.
>“Currently, customers are paying for a lot of idle compute in the form of virtual machines and GPUs that go unused,”
I'm definitely looking forward to having a lot more competition in the "pay as you go LLM AI" space. Especially services that use models one can download and run on your own hardware once a good use case has been developed.
We have models that are crucial but do not require dedicated hosting. We are looking for an aws lambda type of service, but for a fine tuned llama2-13b. any suggestions? would try out Cloudflare AI too.
It's often cheaper and far more powerful in quality and latency to pay for a full server funnily enough.
For running custom fine-tuned models on serverless, you could look into https://beam.cloud which is optimized for serving custom models with extremely fast cold start (I'm a little biased since I work there, but the numbers don't lie)
Any proof?
Cloudflare doesn't currently have a "not edge" worker, so anything they offer has to be "edge".
Claude 2 is very fast too...
but they're also offering more than just LLMs but also image models, sometimes it takes 190 seconds or more on playgroundai.com and 40 seconds on leonardo.ai, and about same on tensor.art.
I'm trying to get an ai Etsy store off the ground and faster gen times would be greatly appreciated.