1-Bit Bonsai Image 4B Image Generation for Local Devices
prismml.com
prismml.com
There are many problems I want to work on which require billions of tokens. These are completely inaccessible without corporate project sponsorship at the moment. An asic generation machine which can pump out a few 10s of thousands of tokens per second at opus4.6 quality is more than sufficient.
I would very easily find ways to hit that level of token usage if it was cheaper/faster.
This is a hobby version of a project, but you can imagine commercial versions of the same prompt for new databases, genomics studies, material analysis, operating systems etc.
Thats the thing - these models see and predict tokens. For any real engineering you get more bang for your buck using math.
There might be some interesting side effects from making simulation software, which is currently either an expensive niche or quirky university project (SPICE always has that feel).
The TL;DR though is that a 10-15b param model baked into an ASIC with the latest fab tech would take around 62W of power draw when active. At ~10k+ t/s though it likely would only be active for short bursts of time. It'd fit perfectly fine within the thermal envelope of a laptop.
The approach makes a lot of sense. Once you get to those speeds, latency of the network becomes one of the bigger bottlenecks, so local has a real advantage over a subscription.
This isn't true.
Anthropic is making an operating profit including the loss making subsidised subscriptions (but excluding training).
Your normal inference provider is doing great. Do the math on a H100 rental and you can see the margins.
That said, if you are doing always on agents and you spend $3k-$4k on a GB10 or, $5+ k on Apple Silicon as your sunk cost, you will probably come out ahead.
I've got 5 agents running a purely experimental social experiment. AThey operate in an evennia mud (a familiar sounding city called "gothmud). I've built a channel, idle prompts, sleep schedule. I feed in real world news, weather. There's a character up in a clock tower that reads evennia's audit logs every 20 minutes to surveil the city, and a cast of people wandering around, investigating things, having coffee, repairing robots. This is all hitting qwen3.6-35-A3B on the Asus GB10, which cost me $3k.
Over the last 30 days, I've hit 394M input tokens, 1.6B output tokens. I would have spent between $1600 to $1700 if I was using openrouter. Not calculated - I also have comfyui running in the spare space, and the agents "take photos" of the rooms they're in, selfies, workshop photos, etc.
How much did I spend on electricity? I don't have a meter on my box. My total electric bill for the last 30 days was $220, so I know it's less than that. My rate to compare is 11.7/kwh, but it's closer to 15c/Kwh total. The Asus GX10 has a 240W power supply, and it's probably only pulling 180. I estimate $15-$20/month. But worst case red-lining. 240 Watts, 720 hours = 172KWH , and at $0.20, I come to $35
Here's the kicker thought - that github copilot subscription I mentioned? I have another agent running on that, reading all my other agent logs, managing my obsidian notes, doing research, sending briefings. And all by itself, it used almost the same amount of claude-opus tokens for that $39/month subscription. I was actually a bit shocked when I pulled a recent report and saw that. I'm working to migrate functionality away from copilot subscription to the local model. A lot of the initial setup might have needed it, but not the ongoing review style work it does.
What is the experiment? What are you hoping to learn from all this?
Or do you just mean you've made a dynamic dollhouse that you think is cool? The Sims on your own terms?
For A:
The learning is in building agent harnesses that aren't just cron jobs reading a file like HEARTBEAT.md. I have some serious tools for my own use. One main assistant/coordinator agent, one SRE/coder agent (with sub-agents of its own).
I originally just started last year with the AI assistant (Jane from enderverse). Along the way building scheduled systems, hand offs to other agents, etc. As I ran into problems, I'd be rewriting and refactoring. So I spent some time making a low-stake hatbot with history and routines. Instead a from-scratch golang harness, I built it around pi and extensions. Time of day prompt splices (extensions can inject into or modify prompts on the fly, wake up reminders. Things that you do in the main session vs spinning up an ephemeral session. Self improvement daydreaming (modify your own skills and AGENTS.md) A lot of that went back into rebuilding Jane to something more useful for me.
For B:
The "dynamic dollhouse" as you put it was seeing where I could take that living chatbot next. There's a lot of projects pointing agents at slack, discord, message boards. I figured why not a mud with rooms, weather, and props. Lots of interesting challenges. How to keep bots from nesting in their own room, how to keep them from yes-anding each other all day long. How to slow down 3 bots talking at each other so a human can get a word in edge-wise.
Different levels. There's plain old NPCs that have dice roll random responses. There's LLM driven NPCs that only remember the last 5-10 messages. And the main ones are bot agents. Full agent harness, moving around the environment. Long lived context windows. One character (a nurse at the hospital) gets into arguments with an NPC receptionist that treats her as another patient. Complains about it to other characters, they remember and the word spreads.
The agents get prompted to write down notes, the head home for sleep (and session compaction). Next time they enter a room with that person after compact, their notes get loaded automatically. This kind of behavior can feed back into the more productivity based agents.
Memory recall:
Lots of systems out there to give agents memory. I've used a bunch and written a couple. Storing memories is easy, but getting an agent to recall them, no matter how much you mention it in your AGENT/CLAUDE.md is a bit of an uphill battle. I've even watched claude make useful project memories and never refer to them again.
In my agent ant farm - agents go "to sleep" at night. They get nudged to head home, once there they get prompted to make notes about their day, about other characters. Then we do a compact with custom instructions. After compact/sleep cycle, if they enter a room with one of the characters in their notes, that gets loaded back into context automatically.
That all boils down to hooks in Pi like before_agent_turn. You can intercept a prompt, check it against code/flat files, and smartly inject more information into context. You can have a long running main session with compacts that discard procedural bits and offload the rest to memory.
Time Awareness:
Agents have no concept of time. You can send them a message at 5am, then at 10pm, and it's been 2 turns for them. For coding, this is fine. But for assistant level stuff, adding a message like "It's 3PM. It has been 3 hours since the last interaction with the user" goes a long way. Without me saying something like "new topic", it knows now that time has passed, i'm probably onto something new. If I left something hanging, it will remind me about it, or maybe go check on things that should have happened during the day.
Inner Thoughts/Idle nudges:
I can have an extension run every 5 seconds, check a a schedule, check activity level of the main session and fire off nudges on the main session. These look like the user sent it, but I generally prefix it with [inner thought]. For my social bot, I tested this along the lines of "[inner thought] it's been 3 hours since you last talked with user, why not reach out, let him know what's new, maybe send a selfie or a photo of where you are". For my assistant bot, it's an 8am, 3pm, and 7pm nudge along the lines of "[inner thought] put together an activity report of work things that has changed since the last report". This all runs in the main context, they get the thought, have historical context, can run skill to check on vault updates, open beads, anything observed from ingesting other agent sessions, and sends me a summary. It take into account my idle factor. If I'm heavily engaged in conversation at 3PM, the report might get delayed 15 minutes or an hour, or skipped altogether.
Logically five people pooling their resources beats one guy.
therefore datacenters will always win because they get higher time utilization.
so forget it.
I always wonder the same but i let logic tell me its a fantasy, on average you cant outspend a whole group of people making better use of the hardware.
you will get better hardware though, cutting edge will always be cloud
Don't think anyone was refuting that?
And of course when you pool resources you have access to more resources.
Upgrading local hardware will remain the more expensive alternative to the subscription regardless what the relative cost of running the models themselves are. If the local hardware to do so becomes affordable then the subscription will be even more affordable, not expensive.
At least for these kinds of mega tasks. For more micro task we will always end up with unutilized local compute we already purchased which will be "free" since we already paid for non-AI reasons (e.g. a gaming GPU while not gaming).
The problem is that expectations rise in datacenters, hardware/power/security/availability guarantees cost real money. Then the operator providing these guarantees expects some margin.
You can see this most clearly with "developer desktops", a gcp instance costs about 10x a hetzner instance which costs between 5 and 10x the same hardware sitting in the back of an office somewhere. While all of these premiums matter for 24/7 systems under active development, they don't really matter for ephemeral small scale workloads.
At 10x you have to be at hours per day and 5x you’re at 4h.
Paying 2k for something that you use 100 hours of is quite expensive. Having the capability built into your existing silicon which you would buy for 1k is cheap. Paying 200 dollars a month for 2 years give a present value of $4200 dollars. Meaning that that paying 2k upfront cuts your overall spend in half.
I spent 6k on codex last month which, if repeated, implies a present value of ~144k.
HBM has way higher bandwidth and its not all about flops.
Also the FP4 flops (inference) are so mind bogglingly high on these things.
Lastly what you fail to consider is the chip to chip bandwidth which is critical.
the people running these know that networking is just as critical.
all reduce etc.
they wouldnt pay if they could get something better value.
Which explains why you're using a dumb terminal to access compute services?
I understand the point of distorted facts, but what I’m not sure how things are improved by basically having no trust in any facts?
I personally worry that what that would mean is we are left with little to no institutions to trust, besides Universities and family members, I don't think I would be able to trust governments and corporations, but I guess before internet people also weren't blindly trusting those.
One implication was people were more social and talked about ideas more. Thought had not been outsourced to arbiters in the way that it was in the U.S. People with authority, knowledge, and close family members were definitely inputs into what people thought, but by-and-large people still came to their own conclusions.
You got to see the gradient of thought that people actually had about issues. People would say their insane ideas out loud. You could disagree with people and have them actually engage with your perception of reality.
It was strictly better in my opinion.
Listen, I'm hosting this Telegram channel for people like us, where we can exchange free information without media bias, share the real facts and plan coordinated activities against these poisoning mainstream scumbags.
I also have a 20% coupon code for Wamp® Wolf-Testosteron for you, Wamp® really helped me stay awake and alert in these dire times.
It's true: Wamp, it really whips the Llama's ass!
A high information trust society has regulation in place which tracks such acts of manipulation so that this trust is not abused (e.g. regulation against misleading/false advertising).
A low information society promotes the notion that everyone is lying anyway and everyone is on his own to figure out what's true. So regulation gets dismantled, and the premise becomes "it's not lying if I can make enough people believe it".
“which plays on this general tendency of humans to prefer comfort over challenge, confirmation over rejection.”
This is completely and observably wrong. It reads like it would make sense, but the most ideologically open cultures I’ve seen are LIT.
I lived in societies of high and low information trust, and observed first-hand what happens on transition. If trust in public information is high, people contribute and challenge common sources of information, and there is high expectation and punishment to rectify, improving the quality and transparency of commonly agreed information.
In low trust environments, people start to distrust everything, initially seek out explanations but then largely gravitate towards information sources which confirm their assumptions or own bias, because they lack time, energy and skills for any other path. To stabilize themselves, they then build a mental "fortress" around their belief which is periodically fostered not by challenging it but by seeking out others who confirm it.
Advancing in this direction increases the general consensus that there is no common ground (because this requires common trust in some information source), it gets increasingly difficult to educate such fractured groups, because there is no longer a path for many to accept inconvenient truth in light of a more convenient "alternative truth".
This is poisonous in a democracy, because common agreement on facts is what's so crucial for this process.
Hence my learning that it's not a healthy direction for a society (and also the reason why systems attacking democracies don't aim to gain trust, they build distrust)
> but the most ideologically open cultures I’ve seen are LIT.
There is a huge difference between a ideologically open society and a low information trust society. Accepting other ideologies requires trust.
A low information trust society breaks down trust not just in institutions but also among citizens, which is the fertile ground for polarization and actively prevents open ideology.
You may have the wrong understanding of the terminology. A high-information-trust society doesn't mean that everyone blindly accepts a leader, it describes a system where trust is maintained. In a functioning democracy this happens because society constantly challenges institutions, as it DEMANDS to be able to trust them.
In low-information-trust systems, the consensus is that no one can be trusted. So no joint effort is taken to hold power accountable, everyone plays to his own fraction which erodes common grounds, solidarity and citizen power.
Low information trust societies get destroyed by pandemics of both physical viruses (due to anti-vax and medical distrust; we can see this happening again with Ebola) and destructive memetic lies (see 20th century fascism).
Moving there, I was shocked at how "conspiratorial" everybody seemed about everything. Living there, I was shocked out how often they were right. But it didn't impact people's likelihood to do things. I think they are actually orthogonal in a way that is unintuitive.
Here's a local story published after I made my comment, about tour operators using AI images to misrepresent destinations in the area: https://www.abc.net.au/news/2026-06-01/ai-videos-spark-conce...
Increasing the availability of fake image generators directly enables more harms like these.
Now images have gone back to being like a book. People are beginning to assume anyone can write anything, make up quotes, etc. We still get to keep the ones we just like or the ones we know were really that way (e.g. your wedding vows/wedding photos) but we've dropped this silly notion image=fact just because it's an image. It's not all good that faking an image has gotten consistently easier over the last 100 years, but it's also not bad that it's fallen apart enough people are seeing through it.
Slate published this about it in 2012: https://slate.com/technology/2012/03/narrative-science-robot...
For as long as we've had computers people have tried to make them sound human. It's not a new thing that people are concerned about knowing if they're talking to (or reading) a robot imitating a person.
I was surprised how well it worked, even then.
Back in the day, baseball commentators sometimes did this for live games they couldn't see based on very limited information they were being passed. One such commentator was .. Ronald Reagan.
This is seriously underplaying it. It's become trivial to generate and inundate the internet with fake content (either for laughs, for internet points, or for more nefarious purposes). Manipulating photos required a lot of skill to make something plausible. We're reaching a point (if we're not there yet) where most content produced on the internet is fake.
Not saying that tech is inherently evil, these could have been prevented, but to me it seems we have underestimated the social risks and failed to regulate accordingly.
That said, while I am generally skeptical of these effects, fake news are a real problem that social media has exacerbated.
If the internet saves 1,000 lives per day, and hurts 1, then the roi is there.
Its seems to be hidden from view. Global health is getting worse, not better
Covid pushed back the longevity gains, but there hasn't be a decline or flatten since the internet grew in popularity (2000 - 2019).
https://www.who.int/data/gho/data/themes/mortality-and-globa...
Obesity continues to increase everywhere, and is a leading indicator for poor health outcomes over time.
What are the positives? Before, you had to go to a doctor to do stuff, now you do too. Maybe people are more likely to get tested. Maybe scheduling appointments is a bit better. What am I missing?
There are plenty of documented cases of reddit posts encouraging deeper investigations or legally authorized tele-health providers (where you can communicate directly with a licensed medical professional).
The number of people using telehealth services is significantly less than the number of people that have bad health outcomes from internet usage. If it were otherwise, ChatGPT, Reddit, and telehealth companies would be managed very differently.
98point6 - 50k+ reviews
https://apps.apple.com/us/app/98point6-by-transcarent/id1157...
doctor on demand has millions of users
Still can't. "ChatGPT can make mistakes." People still trust it, doesn't mean they should. Wiki's not as bad of a tertiary source as it used to be, but it's still a tertiary source and you had a research assignment. Even official authoritative sources can be un/intentionally wrong.
> You should never date someone you met through an app or website because they are 100% murderers.
This remains sound advice that teachers should continue giving children. Even (or morso) now that online dating has been normalized in the meantime. Do I have to explain?
Motivated actors have been able to doctor, fake, or spin media content since time immemorial. But peoples default mode was to trust what they saw. Now that fake imagery is ubiquitous, maybe we'll all get a bit more skeptical.
All of which are under serious threat from social media, buyouts by billionaires, and simple smear campaigns. Not dependent on AI, but effects which will be magnified by AI.
Boy do I wish that were the case. Investigative journalism is rare now and instead favours activist journalism, public debate is hard (but getting better), and institutional trust is at all time lows, for various reasons.
People will muddle through regardless, we're not as fragile as most assume.
I think the inability to see the freedom AI gives people is one of the saddest things I've seen.
I remember when the internet was young people would complain about how it was becoming read-only.
Now we have a tool to let people express themselves and people complain the fact there are fake pics on AirBNB means the collapse of society. Please!
We’re in an era now where every image and video (and for that matter audio) is potentially fake; where knowing what’s real and true is no longer possible.
This was always the case. Spin and propaganda are not new, the way it's conveyed has just become a bit easier. People are not as susceptible to misinformation as most assume, they recalibrate how much stock they put into the things they see based on the quality of the information environment. Basically everyone knows now that the internet has a low signal to noise ratio.
20 years ago your misinformation came from television, radio, and print. All of those things were expensive to produce and there was an implicit need for them to be at least vaguely believable and reliable, because their existence depended on it to continually generate revenue.
- Today, a single person produce 100% AI-generated media for basically the cost of their time.
- That media is as high quality as anything else out there.
- Social media platforms provide the delivery system for free.
- There's no real way to tell if there's even a real person behind the name/pseudonym used for posting it. It might be a person, it might be an algorithm, it might be a nation-state. You have no way to know.
Coincidentally, this is at the top of HN right now: https://news.ycombinator.com/item?id=48355751
It's more about using LLMs to impersonate someone, but the point stands.
Misinformation on Misinformation: Conceptual and Methodological Challenges, https://journals.sagepub.com/doi/10.1177/20563051221150412
> It's more about using LLMs to impersonate someone, but the point stands.
I personally knew someone that fell for the Nigerian prince scam 20 years ago. Same old tricks, just recycled in a new medium.
That's not about technology but is a fundamental moral issue.
It's that the powers-that-be have and always will have more resources to bring to bear than the individual will have to combat them. And right now we're moving at absolute break-neck speed to invent technologies that can be used undermine us individually as well as collectively at ease and scales heretofore never seen.
I believed for a long time that the information age would be the great liberator - the great balancer. But we're on the precipice right now of governmental and societal collapsem and it has everything to do with the massive proliferation and preponderance of either misinformation, or agenda-aligned (shaped) information.
Anti-vax, Antifa, ACAB, conspirituality, QAnon, etc., etc., have all had an enormously negative impact, and we're still in misinformation infancy, and lucky that the leadership in goverment is so old that they're not particularly great at bending these technologies to bear on us. But that's not always going to be the case.
Have you not seen the old Soviet photos where people who fell out of favour were erased from photos?
Media literacy has always been about working out if you can trust something.
Now the ability to edit has been democratised. I choose to celebrate that.
No one’s failing to see the good things, hypothetical or not. Most of us are aware just fine, we just don’t all agree that the negative trade-offs are worth it.
You could never trust ANYTHING you read or see on the internet, this isn't new. There are thousands of old hoaxes that many people still believe.
and so now I'm wondering how cool /fast / compressed a diffusion image generator could be if the images it was trained on / space it worked in was limited to 1 bit (Floyd-Steinberg / Atkinson / your favorite algo here) dithered images.
Training would surely be pretty quick and probably fit onto one modern GPU.
IME, the bottleneck when using diffusion models isn't storage space or memory, it's generation time. Lots of models will run on 8-12 GB 1080-generation GPUs onwards, or on Macs with similar memory, which are probably the bottom end from a GPU power perspective anyway. I also note that these models are marginally slower than the small FLUX.2 model they're based on.
Okay, maybe this allows running a local model on something that has a reasonably powerful GPU and limited memory, like an iPhone, but is that really a common requirement?
I can make this into a 5-lines Python program. I’m not saying the images will match the description, but that isn’t part of your spec ;)
Their 1-bit quantized Diffusion Transformer is just under 1 GB. You also need the text-encoder (4-bit quantized) and VAE (unquantized) for inference and their combined weight is ~3.42 GB.
TBF, even at that size it's no less mind blowing.
It does not need to directly solve any particular problem to be overall good for consumers, by putting pressure to all those subscription based solutions… at least it’s private and does not require you to provide all your data…
But, getting remarkably higher density of capability per unit of compute is a big thing. It means the frontier can get better and cheaper to operate and less resource hungry, and it means what can be accomplished at the edge, on personal laptops or phones, becomes a broader spectrum of tasks.
And, for privacy, there are a lot of things that should run on-device and not everyone has big dedicated GPUs.
For speed, no. Draw Things runs on iPhone just fine and generally faster than their implementation on the same model (FLUX.2 [klein] 4B).
Not the bottom end - most people are on laptops or mobile devices that are much lower GPU power than this.
Sure, you could theoretically take a model compressed in this manner and deploy it on an old netbook and run the calculations on the CPU, but each image would probably take an hour…
If this model can run inside of the 4GiB limit, that makes this infinitely more useful than existing models for me.
This quantization has a small order of magnitude improvement on memory and compute requirements, how can it be slower?
And all that while retaining quality.
This is wrong. But they worded it carefully to be not entirely wrong.
FLUX.2 [klein] 4B (the same parameter class, basically the same model) runs on iPhone through Draw Things app, with 8-bit or 6-bit quantization (hence not "directly", I guess, but that is the technicality that sounds fishy enough).
Isn't SD XL 3.5B? And the refiner model is even larger. Those can run on an iPhone 13 Pro.
I do wonder how these compare to existing image generation models. I've tried https://github.com/alichherawalla/off-grid-mobile-ai for a while but I find the image generation models rather lacking.
So everyone acts as a sort of beta tester for obscure posts.
My email is in my profile. Feel free to reach out.
Sadly right now the expensive developer subscription means the few folks willing to hold a forever subscription make something that barely works then move on… or make something with so many ads it is an app. For example Google’s “Model Garden” app has no ads but still has major UX issues and isn’t suitable for daily use, even though the models are amazing.
Raising awareness of how capable today’s phone hardware is will make normal people demand to run what they choose on their phones. It’d be a much stronger way back to general purpose computing than via all legislation that has been tried so far..
https://huggingface.co/spaces/webml-community/bonsai-image-w...
Led me to wonder what happens if a domain gets a new owner, and they want to petition Apple to remove the block.
Is it compatible with Ollama, ComfyUI or are those providers unneeded, compatible with low-end hardware?
Also, where does "./setup.sh/ drop the components in Linux?
Thank you, Sol
Website Not Allowed “prismml.com” is a restricted website.
Here's a generation in your honor: https://peterc.org/img/johndoe.png
NVIDIA Card Firefox wayland
I can think of a lot of positives. The negatives amount to a convoluted argument about the limits of free speech.
Prisoner 2: I made a picture of a nice sunset over the ocean
[1]https://en.wikipedia.org/wiki/AACS_encryption_key_controvers...
having trouble loading the webgl browser demo on my phone but no biggy
I took few minutes to try to make it work on ROCm (AMD's alternative to CUDA), landed in python dependency hell.
What do you mean? They are the ones introducing the matmul extensions to Vulkan, which makes compute like this possible