Transformers.js – Run Transformers directly in the browser
github.com
github.com
- CLIP in a browser: https://observablehq.com/@simonw/openai-clip-in-a-browser
- Image object detection with detra-resnet-50: https://observablehq.com/@simonw/detect-objects-in-images
The size of the models feels limiting at first, but for quite a few applications telling a user with a good laptop and connection that they have to wait 30s for it to load isn't unthinkable.
The latest releases adds binary embedding quantization support which I'm really looking forward to trying out: https://github.com/xenova/transformers.js/releases/tag/2.17....
I’ve made an npm package of transformers.js v3, which I should update (not sure if I include this yet).
Mostly, I’ve had to have a fork so it runs on bun. V3 when released will support bun just fine. Although the webgpu won’t work, but that’s optional.
[edit: dm if you want to use it, I don’t want to promote a fork]
It’s only 384 dimensions but it works surprisingly well with a paragraph of text! It also ranks better than text-embedding-ada-002 on the leaderboard
https://syntax.fm/show/740/local-ai-models-in-javascript-mac...
I made a small web app with it that uses it to remove backgrounds from images (with BRIA AI's RMBG1.4 model) at https://aether.nco.dev
The fact you don't need to send your data to an API and this runs even on smartphones is really cool. I foresee lots of projects using this in the future, be it small vision, language or other utility models (depth estimation, background removal, etc), looks like a bright future for the web!
I'm already working on my next project, and it'll definitely use transformers.js again!
My plan is to test embedding and retrieval strategies for different RAG strategies to be used by a server or electron app.
1. Large downloads on every visit to a website.
2. Large downloads and high storage consumption for each website using large models. (150 websites x 800 MB models => 120 GB of storage used)
Both of those options seem terrible.
I think it might make sense for browsers to ship with some models built in and be exposed via standardized web APIs in the future, but I haven't heard of any efforts to make that happen yet.
- (44MB) In-browser background removal: https://huggingface.co/spaces/Xenova/remove-background-web. (We also put out a WebGPU version: https://huggingface.co/spaces/Xenova/remove-background-webgp...).
- (51MB) Whisper Web for automatic speech recognition: https://huggingface.co/spaces/Xenova/whisper-web (just select the quantized version in settings).
- (28MB) Depth Anything Web for monocular depth estimation: https://huggingface.co/spaces/Xenova/depth-anything-web
- (14MB) Segment Anything Web for image segmentation: https://huggingface.co/spaces/Xenova/segment-anything-web
- (20MB) Doodle Dash, an ML-powered sketch detection game: https://huggingface.co/spaces/Xenova/doodle-dash
… and many many more! Check out the Transformers.js demos collection for some others: https://huggingface.co/collections/Xenova/transformersjs-dem....
Models are cached on a per-domain basis (using the Web Cache API), meaning you don’t need to re-download the model on every page load. If you would like to persist the model across domains, you can create browser extensions with the library! :)
As for your last point, there are efforts underway, but nothing I can speak about yet!
I'm keen to do more stuff with WebGPU, so very interested to learn about challenges and limitations here.
- WebGPU embedding benchmark: https://huggingface.co/spaces/Xenova/webgpu-embedding-benchm...
- Real-time object detection: https://huggingface.co/spaces/Xenova/webgpu-video-object-det...
- Real-time background removal: https://huggingface.co/spaces/Xenova/webgpu-video-background...
- WebGPU depth estimation: https://huggingface.co/spaces/Xenova/webgpu-depth-anything
- Image background removal: https://huggingface.co/spaces/Xenova/remove-background-webgp...
You can follow the progress for full WebGPU support in the v3 development branch (https://github.com/xenova/transformers.js/pull/545).
To answer your question, while there are certain ops missing, the main limitation at the moment is for models with decoders... which are not very fast (yet) due to inefficient buffer reuse and many redundant copies between CPU and GPU. We're working closely with the ORT team to fix these issues though!
Really glad to hear the last part. Some of the new capabilities seem fundamental enough that they ought to be in browsers, in my opinion.
Works on mobile network, though, so might just be my internet connection.
The Unreal 3 demo using Flash is still on YouTube.
And this is why most game studios are playing wait-and-see with streaming instead, proper native 3D APIs, with easier to debug tooling (Web still hasn't anything better than SpectorJS), and big size assets.
I don't think it would have the same issues, because the files could be stored in a user specified location outside the browser's own storage area.
Browser vendors can't just delete stuff that may be used by other software on a user's system. And they cannot put a cap on it either, because users can store whatever they like in those directories, bypassing the browser entirely.
But I have never used this API, so maybe I misunderstand how it's supposed to work.
So the web app could ask the user to pick a directory for model storage and henceforth store and load models from there without further interaction.
Even then I think cloud hosted models will probably always be far better for most tasks.
Separately, having to rely on preinstallation very likely means stagnating on overly sanitized poorly done official instruction-tunes. With the exception of mixtral7x8, the trend has been the community overtime arrives at finetunes which far eclipse official ones.
It might depend on just how good you need it to be. There are lots of use-cases where an LLM like GPT 3.5 might be "good enough" such that a better model won't be so noticeable.
Cloud models will likely have the advantage of being more cutting-edge, but running "good enough" models locally will probably be more economical.
I'd look to see Apple doing some stuff here.
Lots of cool demos and real world applications you can build with it. Eg we powered an AR card ID feature for Magic: The Gathering, built a scavenger hunt for SXSW, a test proctoring assistant (to warn you if you’re likely to get DQ’d for eg wearing headphones), and a pill counter for pharmacists. Really powerful for distribution to not make users install an app or need anything other than their smartphone.
Maybe just some sort of api to give the website fine grained access to the filesystem might be enough. You'd specify a directory or single file the website can read from at any time.
However at some point you will have to download large files. I feel when done implicitly it's bad user experience.
On top of that the developer should implement a robust downloading system that can resume downloads, check for validity, etc. Developers rarerly bother with this, so the user experience is that it sucks.
https://developer.mozilla.org/en-US/docs/Web/API/File_System...
And in any case it's easier to direct users to install Chrome (or preferably Chromium) or instruct drag&drop than doing the brittle and error and bitrot prone virtualenv-pip-docker-git song and dance.
Js/browser based solutions seem to be very often knee-jerk dismissed based on decade old understanding of browser capabilities.
As I ask it seems wrong to me but just to confirm?
For reference I have a 31mb vision transformer I run in my browser. Building the inputs, running inference, and parsing the response takes less than half a second.
I can understand that but where time is not a factor and solely a question of data, can a model be streamed?
I don't see streaming helping anything besides maybe Time-To-First-Inference, but regardless, you're still not getting any output until the entire weights are downloaded.
Its impressive for what it is, but training would be painful at those latencies (fp16, batch 32, sequence length 512 generates a ~500ms forward pass with a 22M param model)
Like for example
- did the user tap the wrong location on the screen because their device was physically jolted, and can you correct for that, considering you have access to accelerometer in HTML5
- does the user keep repeating an action (checking every box in a list of e-mails) and can you extrapolate the rest of what the user wants to do
- did the user bounce because you popped up a stupid intercom box or newsletter popup, and did you learn anything about what you need to do if you want to retain this particular user in the future
these kinds of things could be done with hundreds or thousands of parameters or less
EDIT: (I responded before your full edit with the bullet list). This next comment is orthogonal to the slow training performance topic I think, but the use cases you reference there don't seem to be well-suited at all for an autoregressive decoder-only model architecture.
That certainly also has to open up possibilities for on-demand predictions?
for example, I used the image upscaling playground on hugging face all the time. But I do it manually here: https://huggingface.co/spaces/bookbot/Image-Upscaling-Playgr...
would transformers.js allow me to somehow executive that in my own local or online app programmatically?