Next top things:
* Continue to build/extract into a yaml+shellscript agentic framework/tool
* Continue exploring pre-segmenting or other methods of spacial awareness
* Write a reSvg backend that sends actual pen-strokes instead of lots of dots
Next top things:
* Continue to build/extract into a yaml+shellscript agentic framework/tool
* Continue exploring pre-segmenting or other methods of spacial awareness
* Write a reSvg backend that sends actual pen-strokes instead of lots of dots
For example, maybe I'm taking notes involving words, simple math, and a diagram. Underline a key phrase and "the device" expands on the phrase in the margin. Maybe the device is diagramming, and I interrupt and correct it, crossing out some parts, and it understands and alters.
Sorry, I know this is vague, I don't know precisely what I mean, but I do think that the combination of text (via some sort of handwriting recognition), stroke gestures, and a small iconography language with things enabled by LLMs probably opens up all sorts of new user interaction paradigms that I (and others) might be too set in our ways to think of immediately.
I think there's a "mother of all demos" moment potentially coming soon with stuff like this, but I am NOT a UX designer and can't quite imagine it clearly enough. Maybe you can.
Thinking about it as a product, I’d want a way to easily slip in and out of “LLM please respond” so it wasn’t constantly trying to write back the moment I stopped the stylus - maybe I’d want awhile to sketch and think, then restart a conversation. Or maybe for certain pages to be LLM-enabled, and others not.
Does it require any sort of jailbreak to get SSH access to the device?
It is triggered right now by finger-tapping in the upper-right corner, so you can ask it to respond to the current contents of the screen on-demand. I think it would be cool to have another out-of-band communication, like voice, but this device has no microphone.
Also right now it is one-shot, but on my long long TODO list is a second trigger that would _continue_ a back and forth multi-screenshot (like multi-page even) conversation.
I’m curious if this is becoming something that you are using in your own day-to-day, or if your focus right now is on building it?
The context for my question is just a general interest in the transition to AI-enabled workflows. I know that I could be much more productive if I figured out how to integrate AI assistance into my workflows better.
The one use-case that is _close_ to ready-for-useful: I often take business meeting notes. In these notes I often write a T in a circle to indicate a TODO item. I am going to add a bit of config in there, basically "If you see a circle-T, then go add that to my todo list if it isn't already there. If you see a crossed-out circle-T then go mark it as done on the todo list" .
I got slightly distracted implementing this, working instead toward a pluggable "if you see X call X.sh" interface. Almost there though :)