1. AI video watermarks that carry over even if a video of the AI video is taken
2. Cameras that can see AI video watermarks and put an AI video watermark on the videos of any AI videos they take
346 karma · joined August 9, 2016
1. AI video watermarks that carry over even if a video of the AI video is taken
2. Cameras that can see AI video watermarks and put an AI video watermark on the videos of any AI videos they take
Not saying we're in necessarily the same situation. But it remains difficult to evaluate effort required for actual progress.
[1]: https://www-formal.stanford.edu/jmc/history/dartmouth/dartmo...
The 7B version was a decent enough starting point in terms of what it can answer (and way fewer folks can run a 13B on their machine).
If you really want you can just replace the 7B model file with the 13B one under the ~/.cache/gpt4all directory on your device and it should just work.
We also have some fixes and perf improvements we plan to release later today in version 0.10.1
We have fixes for the seg fault[1] and improvement to the query speed[2] that should be released by end of day today[3].
Update khoj to version 0.10.1 with pip install --upgrade khoj-assistant later today to see if that improves your experience.
The number of documents/pages/entries doesn't scale memory utilization as quickly and doesn't affect the search, chat response time as much
[1]: The seg fault would occur when folks sent multiple chat queries at the same time. A lock and some UX improvements fixed that
[2]: The query time improvements are done by increasing batch size, to trade-off increased memory utilization for more speed
[3]: The relevant pull request for reference: https://github.com/khoj-ai/khoj/pull/393
Khoj works more like an incremental, natural language version of org-agenda-search (or projectile) rather than isearch.
You've to configure which files it should index first. You can then use natural language to search those files with a search-as-you-type experience.
This is not the same as isearch that just searches the current file for keyword matches.
Second, khoj doesn't index source code and the default search models don't work well with code files.
But yeah someone should implement a natural language isearch as a standalone tool (as suggested by parent comment). It'd be super-useful
And Search can be configured to work with 50+ languages.
You'll just need to configure the asymmetric search model khoj uses to paraphrase-multilingual-MiniLM-L12-v2 in your ~/.khoj/khoj.yml config file
For setup details see http://docs.khoj.dev/#/advanced?id=search-across-different-l...
Yes, we don't do optimizations on the query encoding yet. So SBERT just re-encodes the whole query every time. It gets results in <100ms which is good enough for incremental search.
I did create a plugin system, so that a data plugin just has to convert the source data into a standardized intermeditate jsonl format. But this hasn't been documented or extensively tested yet.
You'll just need to configure the asymmetric search model khoj uses to paraphrase-multilingual-MiniLM-L12-v2 in your ~/.khoj/khoj.yml config file
See http://docs.khoj.dev/#/advanced?id=search-across-different-l...
Khoj and your other apps need more RAM themselves, so practically 8GB of System or GPU RAM should suffice.
Khoj has been tested with CUDA and Metal capable GPUs. So Nvidia and Mac M1+ GPUs should work. I'm think it'll work with AMD GPUs out of the box too but let me know if it doesn't for you? I can look into what needs to be done to get that to work.
[1]: The calculation is [params] * [bytes] GB RAM, so 7 * 0.5 = 3.5Gb
Having something that indexes all your digital travels and makes it easily digestible will be gold. Hopefully Khoj can become that :)
Ideal: 16Gb (GPU) RAM
Less Ideal: 8GB RAM and CPU
2. The web UI isn't required if you use Obsidian or Emacs. That's just a convenient, generic interface that everyone can use.
Of course, having it be stable enough to not `rm -rf /` soon after is definitely not part of the warranty
Khoj is using the Llama 7B, 4bit quantized, GGML by TheBloke.
It's actually the first offline chat model that gives coherent answers to user queries given notes as context.
And it's interestingly more conversational than GPT3.5+, which is much more formal
Khoj can index directory of PDFs for search and chat. But it does not currently work with scanned PDF files (i.e not with ones without selectable text).
Being able to work with those would be awesome. We just need to get to it. Hopefully soon
If you can collate your notes into markdown or some such, then messy notes can be handled, at least using Khoj with GPT3.5+.
Do let us know how we can help out and what your current biggest pain-points are?