Workers AI: Serverless GPU-powered inference
blog.cloudflare.com
blog.cloudflare.com
It had a cold boot and run on a 8 word STT time of 45 seconds and warm never got past 15 seconds.
This does not work for STT, where it has to be much faster turnaround.
Can anyone give feedback on if whisper and any of its model sizes can work well on serverless?
Do most AI serverless solutions suffer from significant cold boot delays?
The cheapest persistent GPU cloud instance I saw on G was ~$160 a month. Is that roughly the kind of money people need to be prepared to spend to have a model ready to go at all times as a service to another product?
(Disclaimer: I am the author)
And possibly has coverage in only 65% of the desktop browser market. [1] Does that roughly conform to how you understand the current penetration of this browser api?
Presuming coverage for a given user, I don’t have a good answer for why to consider remote.
It seems like it would be worth testing compatibility for webgpu and attempting to run on the client if possible, but then have a remote instance available otherwise.
Does that make sense to you?
Can you tell me another reason why someone would want a remote instance of whisper given 20x realtime potential at client in your project?
This means that the primary usecase for whisper-turbo and my upcoming libraries is Electron/Tauri apps. For users that don't have WebGPU supported for whatever reason, we will still hit OAI/other server deployment. In the ideal case there should be a 90% cost reduction and same or improved UX.
Someone will still want a remote instance today as there is still engineering to be done. I need more aggressive quantization, better developer experience and more features in order to get people off of the OAI API.
I see WebGPU is available in safari tech preview 92, but still an "experimental" feature there.
It looks like Webgpu is kind of been around the block for a while now. I wonder what the hold up is at firefox and safari. It would be much preferable to run more ops on the client. (Complete speculation but I could see this, and the battery use implications possibly making Apple hesitant)
My goal is to provide a browser-based experience first. This is to get at the largest potential user base and there's no friction from install. So, at least for now a electron app is not in plan.
Firefox has a working implementation, but it lags behind the Chromium one.
Browser based still works great! Check out the whisper turbo demo: https://whisper-turbo.com/
Is this because the users are streaming audio in a more conversational style?
For example, when you give siri a command, it is stated, and then you stop speaking.
For most of ChatGPT‘s life, in openAI’s iOS app, if you wanted to speak to input text, you would tap the record button, and then tap it off, either using the app’s own Speech to text capability or siri’s input field speech to text.
Conversational speech to text is more ongoing, though, which would make a 10 second cold start OK, because you don’t sense as much lag because you’re continuing to speak.
Or perhaps people in general record input longer than 10 seconds, And you are sending the first chunk as soon as possible to get whisper going.
Then follow up chunks are handled as warm boots? Then the text is reassembled? Is that roughly correct?
Anything you can provide on sort of the request and data flow that works with a longer cold boot time in the context of single recording versus streaming, and how audio is broken up would be helpful.
Correct me if I'm wrong, but that's how I interpreted it when they first started with ML models on their CPU's.
"Neurons are a way to measure AI output that always scales down to zero (if you get no usage, you will be charged for 0 neurons). To give you a sense of what you can accomplish with a thousand neurons, you can: generate 130 LLM responses, 830 image classifications, or 1,250 embeddings."
130 LLM responses of that length? 1,250 embeddings what size of text?
“neuron” is a cute name but there’s too much conceptual overlap with floating point ops, layers, model parameters etc which are time independent. Should just call them inference credits or something. When some large model runs on multiple GPUs it’s even more confusing what neurons / dollars per second might be.
The post says that 1000 neurons will give you 130 LLM responses - but of what length?
(LLMs are generally priced by input and output tokens. The longer the tokens the longer the compute time. Without an idea of what you mean by a response it's hard to understand.)
Likewise: 1,250 embeddings – How big is the text size in the example?
I'm VERY excited to see you doing this and understand it's early stages, but I wan't wrap my head around the pricing without context.
Error 1102
Worker exceeded resource limits
Did I mess up the config or is it just not intended to be tried out without being on a paid plan already?Also, we're still figuring some things out, but current limits are here: https://developers.cloudflare.com/workers-ai/platform/limits...
Honestly, I'm beginning to think that we need some kind of documentation-first style development. Like TDD, but DDD....
I have integrated Facebook APIs, Instagram (well, same thing), Google APIs, Stripe APIs, Mailchimp APIs, etc. etc.
And the only thing common among all of them? The documentation is _terrible_..., like, _terrible_.
I also run a few products online and I spend so, so much time trying to get the documentation right. It's incredibly boring and tedious, but I really feel that if you want to set yourself apart from the big players, make good documentation. It can't be that hard.
The worst documentation are the jargon filled abstract vibes-based ones where the authors basically typed it with one hand on the keyboard. It's like "ok, you're amazing. Now how do I resolve this error and what are your command line flags?"
{'errors': [{'code': 'invalid_union', 'unionErrors': [{'issues': [{'code': 'invalid_type', 'expected': 'object', 'received': 'string', 'path': ['body'], 'message': 'Expected object, received string'}], 'name': 'ZodError'}, {'issues': [{'code': 'invalid_type', 'expected': 'object', 'received': 'string', 'path': ['body'], 'message': 'Expected object, received string'}], 'name': 'ZodError'}], 'path': ['body'], 'message': 'Invalid input'}], 'success': False, 'result': {}}
Also happy to help if you run into any other issues - pwittig at cloudflare dot com.
But I have a question - why not make inference as easy as the translation? Why do I have to run that in a worker rather than just as a simple API call? That would be much simpler.
Is there a technical reason or is it that people would want to have logic before making the call to llama?
docs: https://developers.cloudflare.com/workers-ai/get-started/res...
(llama specific example here too under curl: https://developers.cloudflare.com/workers-ai/models/llm/ )
Very curious if you can elaborate on Vectorize. More than edge GPU's, entering the Vector DB marketplace and a CF proprietary integration is interesting (and a bit scary) to build on.
- Will Vectorize ever get OSS'd?
- If you want to migrate either direction from some other Vector DB(Milvus, Weaviate, Qdrant, Pinecone, etc), what should you expect in terms of level of effort and features?
- What inherent advantages(latency? features?) would you get exclusively from Vectorize?
Their pricing for their products always seems so much more affordable than something like AWS or GCP. I'm using R2 for storage for a client project, and from my calculations we would have to pay almost 3 times more if we hosted the files with AWS on S3.
I really hope they keep adding all the latest open-source AI models to their platform. If their pricing will be as cheap as they state in this blog post, then I would rather use this service than install models on my computer. To get a good inference speed for open source models on my PC right now, I have to let usage spike to up to 100%...
https://replicate.com/ is much better with an average 10 requests per second but would still very much prefer unlimited (where they can monitor high volume users to see if it's legitimate), or add new pricing model where e.g. 0-10k/rpm is at $0.001/sec, 10-100k/rpm is at $0.010/sec, 100k+/rpm is at $0.100/sec (pricing would of course need to be fine-tuned, just a quick example).
> (Coming Soon) The Llama 2 13B and 70B parameter models by Meta will soon be available via Amazon Bedrock’s fully managed API for inference and fine-tuning.
https://twitter.com/eastdakota/status/1707056412575023352?t=...
Are the models run in JavaScript/WebAssembly behind the scenes?
(bias: am Banana CEO)
A cheap offering like this can make it a lot more reasonable for self-hosters.
OT but if cloudflare fixes their false positives on showing random captchas when I'm trying to browse the net on VPN they will surely be one of my favorite companies.
It's region-less, and runs your inference task on the Cloudflare network, near your end users. Though that's not entirely true yet - we'll be in 100 sites by EOY '23, and nearly everywhere by EOY '24.
It was built to work alongside our new vector database, Vectorize, out of the box.
It's accessible to all developers, regardless of where you deploy (via API), but we wanted to offer a seamless option for developers already building on Cloudflare - Workers, Pages, etc.
Because of how they operate, they can't just release in a single DC like other providers. It's a partial "all or nothing" scenario.
Also, this strikes me as really weird phrasing:
> Llama 2 now available for global usage on Cloudflare’s serverless platform, providing privacy-first, local inference to all
"privacy-first & local inference" would be to run it locally, on your own hardware, isn't that exactly what should be referred by when using "local"? Or has the definitions of words completely gone out the window as of late?
It would be similar as their partnership with Huggin Face, I suppose.
What data would be useful to capture? Outside of a human feedback loop, I don't think there is an actual use-case for it.
Note: not certain
> What data would be useful to capture?
Pipe the prompts straight to Meta and I'm sure they'll be able to extract a ton of useful data.
But off the top of my head, classifying the data into various categories and mentions of Facebook/Meta could allow them to derive sentiment about the company based on geographical location, and know where to invest more in changing the sentiment.
This would be incredible hard to do without IP.
Cloudflare won't forward it and it can be setup as an API. Additionally, IP's on mobile are re-used across users.
That doesn't even sound remotely to a valid indicator for "investments". Definitely not considering the costs to create and update such a model.
>Models you know and love
>We’re launching with a curated set of popular, open source models, that cover a wide range of inference tasks:
>Text generation (large language model): meta/llama-2-7b-chat-int8
>Automatic speech recognition (ASR): openai/whisper
>Translation: meta/m2m100-1.2
>Text classification: huggingface/distilbert-sst-2-int8
>Image classification: microsoft/resnet-50
>Embeddings: baai/bge-base-en-v1.5
> This is just a press release, and not a great source of in-depth info
Not sure I'd call outright lying/getting the most basic points wrong "not a great source of in-depth info", like the "local-first" part.
I think it means local as in "nearby" - so if you're in the UK it isn't processed in us-east-1, for example. You can pick one local to you.
https://blog.opensource.org/metas-llama-2-license-is-not-ope...
Thus you will have to read the license and judge whether it is compatible with your business or use cases.
One should give them credit for making it available, which is a lot better than plenty of others. However, actually open models are starting to appear, so perhaps we will soon see Facebook and the likes making theirs open as well? Who knows.
It's hard to take Cloudflare's commitment to security seriously when they still ship such terrible default settings.
You can install certificates by cloudflare and then the only one that can connect to your server is from cloudflare.
No one can intercept it then.
If you're talking about flexible SSL. Sure, you can use it purely as a https proxy for your SEO score of your blog. But securing it is not much effort.
If it's just for a static blog, I'm not sure what you would though.
To the less experienced sysadmin everything looks like it's working fine and users also don't notice any difference, which is why it's a terrible default.
Sure you _can_ configure Cloudflare securely, but it should be secure out of the box. But that adds friction when the origin doesn't have a valid SSL certificate which probably hurts someone's KPIs.
Cloudflare drops this! Sweet! Now, we have BYOAI.
Our whole offering runs on their network https://RTEdge.net
Our playground lets one live-code user workers that deploy to Cloudflare Worker for Platforms!
- https://sw.rt.ht/?io (in-browser)
- https://sw.rt.ht/ (region-Earth)
Once we raise funding, and establish strong market presence, we will revisit this decision, and dedicate developer resources to sharing our in-house tech with the world more openly. This will take a full position. If you liked what you see, and want to work with us, send us an email at work@elefunc.com and we will get in touch once we have open positions!
By way of comparison, I know exactly what I'm getting as a developer on Replit's homepage: "Make something great. Build software collaboratively with the power of AI, on any device, without spending a second on setup"