Ghostwriter – use the reMarkable2 as an interface to vision-LLMs
github.com
github.com
Next top things:
* Continue to build/extract into a yaml+shellscript agentic framework/tool
* Continue exploring pre-segmenting or other methods of spacial awareness
* Write a reSvg backend that sends actual pen-strokes instead of lots of dots
Thinking about it as a product, I’d want a way to easily slip in and out of “LLM please respond” so it wasn’t constantly trying to write back the moment I stopped the stylus - maybe I’d want awhile to sketch and think, then restart a conversation. Or maybe for certain pages to be LLM-enabled, and others not.
Does it require any sort of jailbreak to get SSH access to the device?
It is triggered right now by finger-tapping in the upper-right corner, so you can ask it to respond to the current contents of the screen on-demand. I think it would be cool to have another out-of-band communication, like voice, but this device has no microphone.
Also right now it is one-shot, but on my long long TODO list is a second trigger that would _continue_ a back and forth multi-screenshot (like multi-page even) conversation.
I’m curious if this is becoming something that you are using in your own day-to-day, or if your focus right now is on building it?
The context for my question is just a general interest in the transition to AI-enabled workflows. I know that I could be much more productive if I figured out how to integrate AI assistance into my workflows better.
The one use-case that is _close_ to ready-for-useful: I often take business meeting notes. In these notes I often write a T in a circle to indicate a TODO item. I am going to add a bit of config in there, basically "If you see a circle-T, then go add that to my todo list if it isn't already there. If you see a crossed-out circle-T then go mark it as done on the todo list" .
I got slightly distracted implementing this, working instead toward a pluggable "if you see X call X.sh" interface. Almost there though :)
For example, maybe I'm taking notes involving words, simple math, and a diagram. Underline a key phrase and "the device" expands on the phrase in the margin. Maybe the device is diagramming, and I interrupt and correct it, crossing out some parts, and it understands and alters.
Sorry, I know this is vague, I don't know precisely what I mean, but I do think that the combination of text (via some sort of handwriting recognition), stroke gestures, and a small iconography language with things enabled by LLMs probably opens up all sorts of new user interaction paradigms that I (and others) might be too set in our ways to think of immediately.
I think there's a "mother of all demos" moment potentially coming soon with stuff like this, but I am NOT a UX designer and can't quite imagine it clearly enough. Maybe you can.
I made a little app for reMarkable too and I shared it here some time back: https://digest.ferrucc.io/
Basically authentication with devices is "all-access" or "no-access". I would've liked it if a "write-only" or "add-only" api permission scope existed
https://news.ycombinator.com/threads?id=memorydial
" @dang " isn't a thing, he doesn't watch for it - take credit and email him direct.
FWiW I mostly read HN at it's deadest time (I'm GMT+8 local time) and I see a lot of mechanical turk comments, especially from new (green coloured) accounts.
I always look for a response (eg: yours) before flagging them as spam bots . . .
AltR(hold) - - -
(The discoverability of these functions is way too low, on GNOME/Linux; I really dislike the direction of modern UX, with its fake simplicity, and infantalization of users. Way more people would be using —'s and friends if they were easily discoverable and prominently hinted in their UX. "It's documented in x.org man pages" is an unacceptable state of affairs for a core GUI workflow).[0] https://news.ycombinator.com/item?id=35118338#35118598 (On "Punctuation Matters: How to use the en dash, em dash and hyphen" (2023); 356 comments)
Retrieval is tricky as Algolia doesn't index '@' symbols:
https://hn.algolia.com/?query=%40dang%20by%3Adang&sort=byDat...
edit: found the official developer website https://developer.remarkable.com/documentation
It's one of my favorite pieces of hardware and wish there were more apps for it.
I wanted to try to implement this for months. You did a really good job.
Main limitation is that the reMarkable drawing app is very very minimal, it doesn't let you place text in arbitrary screen locations and is instead sort of a weird overlay text area spanning the entire screen.
I’ve been playing with the idea of auto creating tasks when I write todos by emailing the PDF and sending it to an LLM.
This just opened up a whole realm of better ways to accomplish that goal in realtime.
Rust binary so should be easy to install. In theory :)
I don’t use discord much but I’ll find you somewhere around here!
"proof" to partner of tablet investment value based on interactive fiction conversation == excellent strategy and nothing could go wrong
The other way to go would be to make a specific app. I just picked up an Apple Pencil and am thinking of porting the concepts to a web app which so far works surprisingly well ... but for a real solution it'd be better for this Agent to interact with existing apps.
I’ve been working on a different angle - in place updating of PDFs on the Remarkable, so it’s cool to see what you’re working on. Thanks for sharing it.
But also -- the main thing that might be different is the screenshot algorithm. I'm over on the reMarkable discord; if you want to take up a bit of Rust and give it a go then I'd be happy to (slowly/async) help!
That said, I based the memory capture on https://github.com/cloudsftp/reSnap/tree/latest which is a shell script that slurps out of process space device files. If you can find something like that which works on the rPP then I can blindly slap it in there and we can see what happens!
I like it.
[0] https://fonts.google.com/?categoryFilters=Calligraphy:%2FScr...
But one of the early deep learning papers from Alex Graves does this really well with LSTMs - https://arxiv.org/abs/1308.0850
Implementation - https://www.calligrapher.ai/
If nothing else it could use an SVG font that has handwriting; you'd need to bundle that for rendering via reSVG or use some other technique.
But if I ever make a pen-backend to reSVG then it would be even cooler, you would be able to see it trace out the letters.
you know, the diary that wrote back to you and possessed your soul? that cursed diary?
edit: https://viwoods.com/ (based in Hong Kong)
edit 2:
It's a blatant copy of the Remarkable 2 for sure :/ LLM integration is interesting --> Remarkable are you listening?