50 karma · joined July 30, 2025
I have been thinking about better and faster ways to communicate with agents for UI related changes and this is the right direction I believe.
Another approach is to run OCR on 1FPS screenshots. Everything runs locally without draining the battery like an LLM would.
It did not change the text on a hat (ended up changing 1 of 3 words).
On one occasion it regenerated the same image again, ignoring my instructions to edit.
I get the feeling that this model is optimised for images with people in it than objects or drawings etc
What platforms have you been using for this? Did you face any challenges in getting the backend setup?
Hyperclay has given me some ideas. What I want is something like [3] but that the user only needs to install once. One electron app that can load our mini-apps.
[1] - https://news.ycombinator.com/item?id=44930814
After building this, I realise that we need a standalone electron that can download and run offline apps that are more complex.
How are you planning to market this?