I used GPT to build a search tool for my second brain note-taking system
reasonabledeviations.com
reasonabledeviations.com
There's a big up-front cost of building a notes database for this application, but it illustrates the point nicely: encode a bunch of data ("memories"), and use an AI like GPT to retrieve information ("remembering"). It's not a fundamentally different process from what we do already, but it replaces the need for me to spend time on an automatable task.
I'm excited to see what humans spend our time doing once we've offloaded the boring dirty work to AIs.
For similar reasons, it doesn’t really matter that computers are better at chess than us.
In chess the first "new" (unplayed in a high level game) move in a game is called the novelty, or theoretic novelty. In times before computers this would not infrequently be an objectively strong move that simply had not been played in a given position before. And this continued for time after computers became quite strong with players using computers to find interesting strong ideas in all sorts of positions. Each time these sort of novelties would be sprung, positions would become redefined and our broader knowledge of the game continued to stretch on outward.
But then something fun happened - the metagame shifted. Now it's no longer really about founding some really strong move as your novelty - but often about finding a technically mediocre, if not simply bad, move that gives you good practical chances. So you're looking for moves that your opponent probably has not considered because they look bad (and the computer would agree that they're bad) but you're much more prepared and comfortable in than he is.
The big difference now also is that instead of a novelty redefining a position in a positive way, it's often something you spring once or maybe twice - and then never touch again. And this is now happening regularly at the absolute highest levels of chess. So rather than having humans just desperately trying to emulate machines, those machines became yet another tool to exploit and improve our practical results with.
It's kind of funny watching a game when this happens and less experienced players will immediately begin shouting "BLUNDER!" when the computer evaluation of a position suddenly drops, without realizing the player who just "blundered" is still well within his preparation. But the other guy is now probably out of his. Even the players themselves, there's often a sort of "u srs?" type response. This [1] is a fun one from the always emotive Ian Nepomniachtchi during the most recent world champions candidates event. He is now playing for the world championship. In any case, it's at that point that the game begins!
Automation has been here for quite a long time now, if it took people out of the work pool for them to become entertainer we'd know about it.
It's always the same issue in fact, replacing workers by machines is good, but if your goal is still to have a "full employment" society you have to make them work somewhere else. It's not even a new concept but we seem to rediscover it every now and then apparently
> Automation, the most advanced sector of modern industry as well as the model which perfectly sums up its practice, drives the commodity world toward the following contradiction: the technical equipment which objectively eliminates labor must at the same time preserve labor as a commodity and as the only source of the commodity. If the social labor (time) engaged by the society is not to diminish because of automation (or any other less extreme form of increasing the productivity of labor), then new jobs have to be created. Services, the tertiary sector, swell the ranks of the army of distribution and are a eulogy to the current commodities; the additional forces which are mobilized just happen to be suitable for the organization of redundant labor required by the artificial needs for such commodities.
Guy Debord, 1967
It creates a natural language search assistant for your second brain. Search is incremental and fast. You notes stay local to your machine.
There's also a (beta) chat API that allows you to chat with your notes[2]. But that uses GPT, so notes are shared with OpenAI if you decide to try that.
It is not ready for prime time yet but maybe something to check out for folks who are willing to be beta testers. See the announcement on reddit for more details[3]
Edit: Forgot to add that khoj works with Emacs, Org-mode as well[4]
[1]: https://obsidian.md/plugins?id=khoj
[2]: https://github.com/debanjum/khoj#chat-with-notes
[3]: https://www.reddit.com/r/ObsidianMD/comments/10thrpl/khoj_an...
[4]: https://github.com/debanjum/khoj/tree/master/src/interface/e...
I'd love a system where I can just point a search engine at my brain. I tried really hard for a while, but I just didn't have the discipline or memory to exhaustively document everything.
An AI that can do this kind of thing in the background would be an absolute godsend for ADHD and ASD people.
but I've seen more detailed capability... I can't remember if It was under NDA though. I cant seem to find it with a quick search though
There’s a whole bunch of “crimes” that society just kinda ignored at scale. Eg underage drinking in college. People knowingly and willingly speed when it’s against the law. Imagine if your smart car automatically recorded and stored its speed and gps coordinates at all times? It’d be so easy for the government to automatically subscribe to that data and start sending automated tickets… nevermind all the worse things that can happen.
This data can be manipulated and abused by stalkers and hackers, abusive partners controlling their wives or kids, churches trying to guilt you into behaving differently, etc.
Which is why the very first sentence of my post includes the phrase "self-hosted"
Alexa used to be better, but only yesterday I asked it to flip a coin.
"I've added flip a coin to your basket"
???
I asked again, and she actually flipped a coin.
That's why siri and Google are driving their assistants into the ground with "by the way"
If someone would nail voice assistants that would be a bit selling point for their platform and it's a good way to get people into their paid services too.
Not really sure what you are referring to with driving the assistants in the ground, from what I can tell it's just not very good tech ("one moment", "working on it", "there was an problem answering this query",...) and not related to any monetization strategies.
(I am not affiliated.)
Also, it’s helpful not to cut paragraphs into separate pieces, but rather to use a sliding window approach, where each paragraph retains the context of what came before, and/or the breadcrumbs of its parent headlines.
I like to stand up for the little guy. I hear Pinecone this and Pinecone that. And nobody seems to pay any attention to the awesome dude who made hnswlib.
And yes, both he and HNSW are awesome.
Similar questions for civil trials, divorce proceedings, child custody....
Humans can be required to testify too if they're immunized.
If a cell phone is recognized as having a higher expectation of privacy than a mere passive document, then it stands to reason courts will also recognize a personalized machine that is even more earnest to help as my infernal "by the way..." Alexa home device.
The fifth amendment has very different criteria - with a proper warrant, your most private things are admissible evidence. You can't be compelled to testify but all your most private notes in a safe can be "interrogated"; in many situations you can't be compelled to testify against your spouse but any writings or recordings of what you said about him/her are valid evidence. There is no debate that all the contents of your computer or phone can be used, they definitely can, the cases you quote are disputed only because they are "worthy of the protection for which the Founders fought" which is the requirement for a warrant.
I'd say that fifth amendment is not a privacy law (which is the 4th, stating that you have all the privacy without a warrant and no privacy with one), the 5th is essentially an "anti-torture" law to prevent coerced confessions, and there is no reason why it would apply to some physical evidence like the data for a trained model on your device.
You can read a letter. You can play an audio recording. You can look at a photograph. You can print a word-processing document. Those feel like static records. Each tells a single story, no matter how many times it's read.
An ML model, on the other hand, is just a long sequence of floating-point numbers. There is no meaningful model "viewer" that spits out a text file or JPEG. The only thing a non-engineer human can do with an ML model is interact with it when it's running. Thus, any output is the product of both the human and the model. It doesn't feel like a record. It feels more like a performance -- and a performance by the interrogator, at that.
If an ML-model "record" is a special kind of record that needs manipulation to produce human-readable output, then it feels like we're back in the 5th amendment department. No, chatting with a chatbot is not torture. But it produces a kind of evidence that is an unwelcome collaboration between the interrogator and the record (unwelcome by the owner of the record). It's one thing to let a jury look at a screen full of numbers. It's another thing to say that an expert witness used those numbers in a chat session that printed "Yes, it was my owner, in the living room with the candlestick."
If you lock up a person and ask them the same question over and over again, eventually you'll get the answer you want. The same will probably turn out to be true for that person's chatbot. Should society allow that second case?
Your phone is already more or less an extension of your brain, and whether or not you can be forced to unlock and surrender it for inspection is already a contentious topic.
IANAL, but phone privacy would probably set the precedent for AI assistant privacy.
Always keep your phone encrypted and be aware what your local laws are. Some places will force you to provide biometric authentication, but not provide a PIN or password. Check if your phone has a duress lockdown mode: some phones lock and/or wipe if you press the power button five times or something like that.
I worked on a ‘Semantic Search’ product almost 10 years ago that used a neural network to do dimensional reduction and had inputs to the scoring function from the ‘gist vector’ and the residual word vector which was possible to calculate in that case because the gist vector was derived from the word vector and the transform was reversible.
I’ve seen papers in the literature which come to the same conclusion about what it takes to get good similarity results w/ older models as a significant amount of the meaning in text is in pointy words that might not be included in the gist vector, maybe you do better with an LLM since the vocabulary is huge.
Because of attention mechanisms, we no longer so heavily depend on the existence of those "pointy words," so generally, Transformers-based semantic search works quite well.
Nils has written a lot of papers
and I think the Medium post you are talking about is
https://medium.com/@nils_reimers/openai-gpt-3-text-embedding...
and that SBERT is a Siamese network over BERT embeddings
https://arxiv.org/abs/1908.10084
which one would expect to do better than cosine similarity if it was trained correctly. I'd imagine the same Siamese network approach he is using would work better than cosine similarity with GPT-3.
There's also the issue of what similarity means for people. I worked on a search engine for patents where the similarity function we wanted was "Document B describes prior art relevant to Patent Application A". Today I am experimenting with a content based recommendation system and face the problem that one news event could spawn 10 stories that appear in my RSS feeds and I'd really like a clustering system that groups these together reliably without false positives.
I'd imagine a system that is great for one of these tasks might be mediocre for the other, in particular I am interested in some kind of data to evaluate success at news clustering.
I played around with chatgpt and it worked pretty well. I have a lot of other things in my plate to get around first (including starting a math curriculum) but it's definitely an exciting direction.
I think LLMs and AI are not anywhere near actual intelligence (chatgpt can spout a lot of good sounding nonsense ATM), but the semantic analysis they can do is by itself very useful.
It's a very good idea in theory but takes almost as much work to verify that the flashcards and curriculum that it generates is accurate and not a hallucinogenic nightmare.
The biggest danger is that the target audience are not experts in the desired subject domain, so they have no way of sanity checking the generated curriculum.
I agree that using the training data would probably generate more garbage. But it's the semantic analysis part that I think it's useful. In general, I think VCs and OpenAI are overhyping it by calling it "intelligent" and obscuring the very good use cases of the technology. AFAIK, no one involved has explained how a statistical model running on a Turing machine magically develops agency and awareness, which are requirements for actual intelligence (under my definition, at least).
I'd say the most popular applications are Knowt in the US for now, and Saveall.ai + Revision.ai (my company) in the UK, all been around with BERT/T5 etc long before this GPT trend.
The flashcard accuracy varies wildly amongst current solutions, that's for sure.
For now, this is the most general way to create the exercises: https://trane-project.github.io/generated_courses/knowledge_.... The JSON file thingy is just a script that automates creating these files given the specification, so they will be interchangeable.
Can I train it on 5 years of stream of consciousness morning brain dumps and then say "write blah as me"?
Before I do that, I'd love to know if training data becomes part of the global knowledge base available to everyone..
Privacy on the side of model servers would be good. Open source models that can be run locally would be better.
This is popular because it's much much easier to do effectively than fine tuning and the OpenAI model is very capable of integrating kb snippets into a response. What I have heard is that it's easy to overdo fine tuning with OpenAI's model and makes more sense when you want a different format of response rather than just pulling in some content.
Having said all of that, they do have a fine-tuning endpoint and I am guessing if you find the right parameters and give it a lot of properly formatted training data then it will be able to do an okay job. I have the impression it is not easy to do either of those things quite right though.
As far as privacy, no they will not share your data when you use the API. ChatGPT is different, they ARE using the inputs to train the model.
Unfortunately, the fine-tuning API cannot be used to add knowledge to the model. It only helps condition the model to a certain response pattern using the knowledge it already has.
Put that in a database or some files that you could use Python data science tools with or something.
Then use text completions to translate the natural language query into some short Python program or SQL query etc.
There are already data-focused tools for using OpenAI's newest models for doing this. Search for 'ChatGPT/GPT/OpenAI' data query, SQL, datatable, etc. See the OpenAI Discord #api-projects Discord, I have seen one or two like that.
[1] https://www.pinecone.io/learn/hnsw/
Imagine, I could ask it questions about myself, my friends, and my business. It would in many ways know me better than me from reading all my journal entries.
How many years are we away from something like this?
If offline static markdown exist to be input for AIs to learn, then OK.
But for me, long before Obsidian, they're just notes, because the act of writing them causes you to process them "outbound" which is the same process you need to recall them later.
I am much excited to do more with my Second Brain, but one concern, as you point out, is to use chatGPT or similar; we'd need to upload all our private and sometimes sensitive notes, which is a no go for me. So happy that you do everything locally. I wonder what the equivalent would be to train the model to search and ask questions based on our second brain (plus the already trained information). That's also where Obisidan will win in the long run, as other tools do not have the data locally. Obviously, it's already in the cloud; they could train on them, but training on customer-sensitive data would be a big problem. Something I will follow closely.