Some were asking: a/s/l hahahaha
4,219 karma · joined October 11, 2007
* OpenJam, vibecoding for PMs who feel stuck in the discovery process and want to iterate fast: https//openjam.ai
* AI job search tool, Command Jobs: https://github.com/nicobrenner/commandjobs
* Wheelless, experimental RC toy that moves using vibration: https://www.tiktok.com/@wheelless?_t=8prRonmSXyV&_r=1
* Online JS image processing filters and animation, Blursort: https://jsfiddle.net/nicobrenner/a1t869qf/
Founder
* ClickFono (cloud phone platform for businesses) https://clickfono.com
* AutopilotReviews (customer satisfaction surveys that don't suck) https://autopilotreviews.co
* Lemontech (software for lawfirms) https://lemontech.com
_
Some were asking: a/s/l hahahaha
I lowkey hate that I knew exactly where this reference is from :'(
It's true that almost no-one RTFM
The quality of the output/work seems the same, but the speed at which it gets stuff done is a lot slower, because it's asking for permission so much more
I don't have any numbers/stats, just my impression. However, I imagine that if Anthropic could make the models ask for permission more often, it could be an interesting way to throttle access, without degrading quality of the output
This is already happening. Write code with an AI agent, agent creates PR, other agents review, coding agent iterates, PR gets merged, and hopefully real humans use the application and provide feedback
The latter step is also now staring to be replaced by AI (user directs agent to access app, get data, perform actions)
Similar thing with email
> A protected workspace for each dot
> Each dot has its own cloud computer, where it can browse, analyze information, create files, and run tools. Dots can keep making progress in these workspaces, even when you are not actively engaged.
> Within each dot’s protected workspace, sandboxing restricts what code and tools that dot can access, helping contain the impact of harmful code or a mistaken command. We also isolate users’ cloud environments from one another and maintain the underlying Linux operating system and Chrome browser
> Each dot’s cloud workspace brings together its computer and the tools it can use. You choose which apps to connect and whether to connect your personal computer. Auto-review checks actions that need review before they run
So it seems like it runs on a Linux container on OpenAI’s cloud infra, but can get access to your local env through ChatGPT’s/Codex on your computer if you give it access
Yes, one general classifier would be very hard to train. However, you can create a sort of ensemble of classifiers, each trained in different tasks
I’m currently experimenting with this. So far I’ve combined classifiers for 13 different datasets, my target is 95 (the ones Laya used for training)
One way: separately embed sender, recipients, subject, body - then use the embedding vectors as input to a logistic classifier
With that setup, I get 95% accuracy on email classification, training on 50-100 base examples. The model trains on CPU in under 1min, and it does inference in under 20ms (most of it is running the embeddings, so you can make it faster if you train your own embeddings model)
Here’s a gist with some sample code: https://gist.github.com/nicobrenner/056a5aaff5d0119c0032ecda...
That code applies the embeddings + classifier setup on the Banking77 dataset. It gets 93-94% accuracy depending on the embeddings you use (SOTA for this is ~95%, with much bigger and slower models)
Depending on how much overfit, you can go from routing deterministically based on features/shape of the input data, all the way to training a routing model (which could be a classifier too). I’ll need to experiment to find the best approach
For completely unseen/unexpected, I’ve also experimented routing to a local LLM: request comes in, if there’s a marching classifier, send it there, otherwise send to LLM+training. As the system learns more tasks, the % of requests that go to the LLM go down over time
Here's a gist with code you can use to test the Banking77 dataset: https://gist.github.com/nicobrenner/056a5aaff5d0119c0032ecda...
The gist uses BAAI/bge-large-en-v1.5, which is 1.2GB approx. You can replace it for all-MiniLM-L6-v2 (91 MB @ fp32 or 45 MB quantized fp16) small enough for mobile/edge. With all-MiniLM-L6-v2 it still gets 93.0% on Banking77, only 1.3 points behind bge-large at 15x smaller
I haven’t compared different ways of sending requests to Jev
The data to train the classifiers comes from the datasets used to test them (not from Jev)
I’ve run some benchmarks. Using embeddings + logistic classifier, the architecture matches or beats Jev and Laya in all basic classification tasks (datasets tested: AG News, Emotion, MASSIVE Intent, Banking77) The type of task in which it does really well, especially against Laya, is classification with >50 classes
The classifiers also run in <1ms, so they can be very fast and precise at the same time
But this architecture has no “reasoning”, so it performs rather poorly on tasks that require it, like the ones from the XLNI dataset (Jev/Laya do a lot better on this one)
For the latter cases, you could use add a local lightweight LLM, something like a Gemma model. Or even some basic MLP, depending on the tasks/data
The Von numbers have led me on a rabbit whole of getting a classifier to play Doom
I got it to average 22 kills (the max is 26) on that same scenario that Jeff and Von are testing on (it’s called Defend Center)
Now I’m having it play a more advanced scenario, and it’s doing about 45 kills (SOTA is ~59 kills)
It’s amazing what you can do with small classifiers if you can collect some data. These models I’m testing train on CPU in seconds (what takes the longest is running the game, doing test runs and collecting data), they are <1MB in size and do inference in <1ms on CPU
Edit: after looking at Jeff's numbers more in detail, the 6.5 kills number is not that bad, but it can definitely be better ;)
Also, it only works on Macs
Btw, DJ_Dave will be performing in LA this Thursday, and SF on Saturday : https://linktr.ee/dj_dave
For example, a typical/stock LLM can’t really play Doom in real time, but a Jev-like model can. Just because of latency
Of course, if you want the best Doom player, there are way better and faster adhoc models
This is partly the appeal of Jev et al; having a quick model for simple tasks, that doesn’t require that much thinking
It’s amazing all the workflows that models like that can unlock. And yes, classifiers and other ML models have been around for a while for these types of tasks, but Jev has made it easy and cheap to play and experiment. This in turn, is incentivizing people to try them for a bunch of stuff, unlocking creativity and producing a lot of new cool (and eventually potentially very useful) applications
So, if someone else is going to develop something very harmful, is the best solution to develop something equally or more harmful? Or are there maybe other better solutions to this problem?
Very cool proof of concept!
For example:
Going 600 km for a family event. How do you travel? * Overnight train * Flight * Drive * Sleeper bus
To me the answer depends. It depends on things like what kind of family event it is, the urgency, who is going, who is coming with me, availability and timing. So essentially, there's a lot of missed context, that is not being captured or represented well enough to produce good predictions
Where do you see the biggest potential improvement gains?
https://gist.github.com/nicobrenner/056a5aaff5d0119c0032ecda...
The <10MB does not include the embeddings encoder
The gist uses BAAI/bge-large-en-v1.5, which is 1.2GB approx. You can replace it for all-MiniLM-L6-v2 (91 MB @ fp32 or 45 MB quantized fp16) small enough for mobile/edge. With all-MiniLM-L6-v2 it still gets 93.0% on Banking77, only 1.3 points behind bge-large at 15x smaller
Haven’t tried with more
Do you have a specific use case?
Amazing, thank you for sharing your setup. Very cool applications
Also curious about if you plan on doing some sort of routing for the requests. Like detecting the type of task to decide which model to route the request to
The type of task in which it does really well, especially against Laya, is classification with >50 classes
But this architecture has no “reasoning”, so it performs rather poorly on tasks that require it, like the ones from the XLNI dataset (Jev/Laya do a lot better on this one)
For the latter cases, you could probably enhance the architecture with a lightweight LLM, something like a Gemma model. Or even some basic MLP
For emails, I get 95% accuracy with this method, with only 50-100 examples for training
Training the model takes less than 5 minutes on a CPU
The resulting model is <1MB, and inference is sub 100ms
Some other cool things about this approach:
* the model doesn’t train on some “ideal” or general classification, instead it learns your preferences
* the model runs on pretty much any mobile device and can be retrained online on the device
* privacy, the whole training and inference is 100% local, no data goes anywhere (except whatever you feed codex/claude while building the model)
Note: to do a more general test, I made a classifier for the Banking77 dataset. The model is <10MB, trains in <30s on CPU and gets 94.5% accuracy, which puts it in the top 5?models by accuracy for that set (the best one is at 94.86%, but it’s 350MB in size and takes hours to train on a GPU).
The other day a neighbor asked me about AI. I said I wasn’t really up to date with things anymore. They asked: like what things? And then I said: like the Astra model that OpenAI released yesterday, I know nothing about it. And they were like: “bro, yesterday?! And you feel you’re not up to date?! Pfff”
> "Sometimes I genuinely don't know what to do next without asking another model. That made me wonder whether I spent a year building a product, or partly building the appearance of one: something sophisticated enough to work, but which I don't yet understand deeply enough to truly own
This is always the reality for a sufficiently complex system. We only have an illusion of understanding
Now, more specifically about this feeling, it’s the way a lot of managers feel as well. They can only ask others to fix/change things, and they don’t really understand how/why things break in the code. Even if they lead the whole team to build the product
Very cool project
I was wondering if the generated people “fill up a spot” and temporarily change the stats for generating new lives. Eg, if the app already generated 10M stories, is it still using the same initial distribution for the new ones, or is generation conditioned on the stats of the already generated?
The random life I got: https://anyhumanever.com/life/1323665477-1682995380-16357802...