203 karma · joined June 4, 2011
[0] https://pyodide.org/en/latest/usage/packages-in-pyodide.html
That is quite the claim !
Thx !
1. M3 Ultra 512 2. AMD Epyc (which Gen ? AVX512 and DDR5 might make a difference in both performance and cost , Gen 4 or Gen 5 have 8 or 9 t/s https://github.com/ggml-org/llama.cpp/discussions/11733 ) 2. AMD Epyc + 4090 or 5090 running KTransformers (over 10 t/s decode ? https://github.com/kvcache-ai/ktransformers/blob/main/doc/en...)
That would be most useful imho !
I would love to know which are these 3 models, especially if they can perform grounded RAG. If you have models (and their grounded RAG prompt formats) to share, I'm very interested !
Thx.
(It seems to me obvious that a fgrep would sanitize synthetic data obtained from competitors.)
https://www.reddit.com/r/LocalLLaMA/comments/1hqidbs/deepsee...
Epyc Gen4 and 12 memory channels of DDR5 @4800 should give you 7 to 9 t/s.
I must confess that my interest in LLMs is grounded RAG as I consider any intrinsic knowledge of the LLL to be unreliable overfitting. Is DeepkSeek able to perform grounded RAG like Command R and Nous-Hermes 3 for instance ?
Thx for this amazing model and all the insights in your report !