Obsidian-Copilot: A Prototype Assistant for Writing and Thinking
eugeneyan.com
eugeneyan.com
There would be a simple system for rooms and the AI / program would edit them and things to them which when clicked on could lead to new "places"
Hmm good insight there. I've done some experimenting formerly by chunk length and it's been pretty troublesome due to missing context.
https://unstructured-io.github.io/unstructured/bricks.html#p...
Define a custom recursive text splitter in langchain, and do chunking heuristically. It works a lot better.
That being said, it is useful to maintain some global and local context. But, I wouldn't use overlapping windows.
When working with more extensive documents, the process gets a bit more intricate. In this case, your embedding database might need to hold more information per entry. Ideally, for each document, the database should store identifiers like the document ID, the starting token number, and the ending token number. This way, even if a document appears more than once among the top results from a query, it's possible to piece together the full relevant excerpt accurately.
If you use local models then it's a fantastic idea.
And then I found Mem.ai and dove into that instead, and i've been extremely happy with it. It accomplishes this aspect he's offering here (where it uses your knowledgebase to assist in your writing). However, it's also got built in chat with your knowledgebase, and helps with auto-sorting and all of that.
For those that want their data on their computer, I totally see why Obsidian is the most desirable. So this sort of addition would be the best of both worlds for them.
Mem.ai has integrated many aspects into it - I love that I am now unconcerned about tags or folders or categories.
Obsidian/Logseq are already great for thinking by way of their "show a random note" feature. Usually pulling up unfinished thoughts from past days will give me an idea for extending it.
When you say "upload some of the vectorized data" do you mean in a numerical embedding form or that it will embed the original text from original similar-seeming notes directly into the prompt? I've only ever done the latter, is there a way to build denser prompts instead? I can't find examples on Google.
Those documents are then injected into your prompt and sent to some kind of LLM completion system such as GPT.
So yes, you will be sending chunks of your actual notes over the wire.
I turned that into cli tool, but it isn't ready to release. Mine works like this:
qwoo index .
qwoo qa "What have I been up to?"Edit: Also, why is the search bar for searching all notes so buried, requiring so much effort to open? Is that because it works so poorly?
[1] https://github.com/brianpetro/obsidian-smart-connections
uses GPT 3.5, requires an OpenAI API KEY
It has prompt templates.
For bulk processing of files I still use my own python scripts.
In the video the user chooses the 'Copilot: Draft' action, and wow, it generates code...
...but, the 'draft' action [1] calls `/get_chunks` and then runs 'queryLLM' [2] which then just invokes 'https://api.openai.com/v1/chat/completions' directly.
So, generating text this way is 100% not interesting or relevant.
What's interesting here is how it's building the prompt to send to the openai-api.
So... can anyone shed some light on what the actual code [3] in get_chunks() does, and why you would... hm... I guess, do a lookup and pass the results to the openai api, instead of just the raw text?
The repo says: "You write a section header and the copilot retrieves relevant notes & docs to draft that section for you.", and you can see in the linked post [4], this is basically what the OP is trying to implement here; you write 'I want X', and the plugin (a bit like copilot) does a lookup of related documents, crafts a meta-prompt and passes the prompt to the openai api.
...but, it doesn't seem to do that. It seems to ignore your actual prompt, lookup related documents by embedding similarity... and then... pass those documents in as the prompt?
I'm pretty confused as to why you would want that.
It basically requires that you write your prompt separately before hand, so you can invoke it magically with a one-line prompt later. Did I misunderstand how this works?
[1] - https://github.com/eugeneyan/obsidian-copilot/blob/bdabdc422...
[2] - https://github.com/eugeneyan/obsidian-copilot/blob/bdabdc422...
[3] - https://github.com/eugeneyan/obsidian-copilot/blob/main/src/...
[4] - https://eugeneyan.com/writing/llm-experiments/#shortcomings-...
I like the "open" aspect and using it to pull from my docs rather than being a cloud based thing.
So it may be "just this one plugin", but Obsidian is so important that I'm just not willing to risk it.