107 karma · joined November 1, 2022
A single AI chat message can consume 0.34 watt-hours of energy (1). So, let's say a hundred messages in an hour (quite an aggressive session) would be 34 watt-hours of energy.
An LCD TV running for an hour consumes about 100 watt-hours of energy, depending on size, LED, vs. OLED etc. (2).
I think AI does help people do better research faster, which is a significant uplift to humanity, while I do not see anyone specifically curbing their TV usage. We should probably focus our effrots on helping people use AI better and meanwhile build more nuclear energy plants, imo.
(1): https://epoch.ai/gradient-updates/how-much-energy-does-chatg...
(2): https://santannaenergyservices.com/how-many-watts-does-a-tv-...
---
And then consider the amount of energy traditionally required by one human to do the same research tasks. Also quite significant.
I think we should be focused on making the more efficient, for sure! But I don't buy that the arguments based on energy consumption are very strong.
Major points of interest for me:
- In the "Main capabilities evaluations" section, the 120b outperform o3-mini and approaches o4 on most evals. 20b model is also decent, passing o3-mini on one of the tasks.
- AIME 2025 is nearly saturated with large CoT
- CBRN threat levels kind of on par with other SOTA open source models. Plus, demonstrated good refusals even after adversarial fine tuning.
- Interesting to me how a lot of the safety benchmarking runs on trust, since methodology can't be published too openly due to counterparty risk.
Model cards with some of my annotations: https://openpaper.ai/paper/share/7137e6a8-b6ff-4293-a3ce-68b...
Whoa, that's really cool.
You can see the paper along with figures & regional breakdowns here: https://openpaper.ai/paper/share/1d0c6956-4820-4ee2-ac1e-12c...
Isn't it very notable that the latency improvement didn't have a performance loss? I'm not super familiar with all the technical aspects, but that seems like it should be one of the main focuses of the paper.
When it comes to things I am not good at at, it has given me the illusion of getting 'up to speed' faster. Perhaps that's a personal ceiling raise?
I think a lot of these upskilling utilities will come down to delivery format. If you use a chat that gives you answers, don't expect to get better at that topic. If you use a tool that forces you to come up with answers yourself and get personalized validation, you might find yourself leveling up.
See: https://arxiv.org/pdf/2305.04388
On a related note, if anyone here is also reading a lot of papers to keep up with AI safety, what tools have been helpful for you? I'm building https://openpaper.ai to help me read papers more effectively without losing accuracy, and looking for more feature tuning. It's also open source :)
Example: Do the problem sets yourself. If you're getting questions wrong, dig deeper with an AI assistant to find gaps in your knowledge. Do NOT let the AI do the problem sets first.
I think it was similar to how we used calculators in school in the 2010s at least. We learned the principles behind the formulae and how to do them manually, before introducing the calculators to abstract the usage of the tools.
I've let that core principle shape some of how we're designing our paper-reading assistant, but still thinking through the UX patterns -- https://openpaper.ai/blog/manifesto.
My approach is for user-driven highlights with AI help. As in, you can upload your paper and directly ask questions ("Why did they only include the HotPotQA dataset?", "What scores did they achieve with the fine-tuned model?", etc). When the AI responds, it provides citations inline to reference text. You can then click on the reference text to take you there in the doc.
Might be easier to visualize using some of the demo screenshots on the README: https://github.com/sabaimran/annotated-paper.
Hopefully that answers your question?
Features I'm thinking about:
- Dynamic podcast generation (journal -> audio summary)
- Model switching for diversity of responses
- Paper search tool for finding relevant papers
- Improve note-taking feature so people can work on writing their own papers in-app
- Improve the citation protocol to make it more reliable
- Diversify input types beyond just PDFs (include audio files, multiple documents, plaintext documents)
How's your experience on Zotero?
You can make it as 'fancy' as you want, and use speech-to-text, image generation, web scraping, custom agents.
Let me know if you run into any issues? I'd love to get this setup for senior citizens! You can reach me at saba at khoj.dev.
Khoj will allow you to plug in your Obsidian vault or any plaintext files on your machine or Notion workspace. After you share the relevant data, it creates embeddings and uses it for RAG, so you get appropriately contextual responses with your LLM.
This is the best place to start for self-hosting: https://docs.khoj.dev/get-started/setup
I responded in the issue, but I'll paste here as well for those also curious:
Khoj does not collect any search or chat queries. As mentioned in the docs, you can see our telemetry server[1]. If you see anything amiss, point it out to me and I'll hotfix it right away. You can see all the telemetry metadata right here[2].
[1]: https://github.com/khoj-ai/khoj/tree/master/src/telemetry
[2]: https://github.com/khoj-ai/khoj/blob/master/src/khoj/routers...
Configuration with the `docker-compose` setup is a little bit particular, see the issue^ for details.
Thanks for the reference points for GPU integration! Just to clarify, we do use GPU optimization for indexing, but not for local chat with Llama. We're looking into getting that working.
I'm particularly interested in your OS/build environment.
For example, if I stayed at an Airbnb last year in Houston and needed to lookup the address for some reason, I'd be going either to gmail and running some keywords searches ("Houston", "Airbnb"), or going to my Airbnb app.
Really, I want a single endpoint where all my personal data can be made available to me, ideally without sacrificing my privacy. Location's a cool use case.
We use it for understanding usage -- like determining whether people are using markdown or org or more.
Everything is collected entirely anonymized, and no identifiable information is ever sent to the telemetry server.
To opt-out, you set the `should-log-telemetry` value in `khoj.yml` to false. Updated the docs to include these instructions and what we collect -- https://docs.khoj.dev/#/telemetry.
For now, local LLMs take up an egregious about of RAM, totally agreed. But we trust the ecosystem is going to keep improving and growing and we'll be able to make improvements over time. They'll probably become efficient enough where we can run them on phones, which will unlock some cool scope for Khoj to integrate with on device, offline assistance.
I'll provide my insight from experimentation integrating Llama V2/GPT4All into Khoj -- Falcon 7b is probably the runner up in models that can be supported on consumer hardware, and it really wasn't good enough (for me) on my machine to be useful. The token consumption with personal notes context is too large, and the content too variable for a small model like that to be able to understand it. It's fine if you're just doing normal question-answering back and forth, but you don't need Khoj for that.
Yeah, I ran into a couple of funny edge cases using Llama v2 with my personal notes. For example, if I ever asked it anything remotely personal (as I would with a personal assistant), it would often start telling me that asking for personal data is unethical. I get it, you have to be careful with the open source LLMs, but still a bit funny. It does work with enough coaxing though.
But that would allow you to access Khoj from the web.
We determine note relevance by using cosine similarity between the query and the knowledge base (your note embeddings). We limit the context window for Llama2 to 3 notes (while OpenAI might comfortably take up to 9). The notes are ranked based on most to least similar and truncated based on the context window limit. For the model we're using, we're still limited to 2048 tokens for Llama v2.