1,084 karma · joined July 24, 2019
https://github.com/NVIDIA/trt-llm-rag-windows https://github.com/NVIDIA/TensorRT-LLM
It's quite a thin wrapper around putting both projects into %LocalAppData%, along with a miniconda environment with the correct dependnancies installed. Also for some reason the LLaMA 13b (24.5GB) and Ministral 7b (13.6GB) but only installed Ministral?
Ministral 7b runs about as accurate as I remeber, but responses are faster than I can read. This seems at the cost of context and variance/temperature - although it's a chat interface the implementation doesn't seem to take into account previous questions or answers. Asking it the same question also gives the same answer.
The RAG (llamaindex) is okay, but a little suspect. The installation comes with a default folder dataset, containing text files of nvidia marketing materials. When I tried asking questions about the files, it often cites the wrong file even if it gave the right answer.
https://github.com/NVIDIA/TensorRT-LLM?tab=readme-ov-file#pr...
Previously there was TensorRT for Stable Diffusion[1], which provided pretty drastic performance improvements[2] at the cost of customisation. I don't forsee this being as big of a problem with LLMs as they are used "as is" and augmented with RAG or prompting techniques.
[1]: https://github.com/NVIDIA/Stable-Diffusion-WebUI-TensorRT [2]: https://reddit.com/r/StableDiffusion/comments/17bj6ol/hows_y...
Given the number of software projects that use github as their canonical distribution platform and the number of supply chain attacks due to hacks, it's no pretty obvious why they've started enforcing 2FA.
> We’re using Coollaboratory’s Liquid MetalPad through their Taiwan-based partner CCHUAN. (https://frame.work/gb/en/blog/framework-laptop-16-deep-dive-...)
It looks like the part is sold with the surrounding protective sheet: https://frame.work/gb/en/products/16-liquid-thermal-pad-amd-...
https://science.nasa.gov/eclipses/future-eclipses/eclipse-20...
At minimum the payload.
https://forum.torproject.org/t/release-candidate-0-4-8-3-rc/...
When under attack, legitimate users will experience a moderate delay whilst attackers will need to scale their compute.
[0]: https://gitlab.torproject.org/tpo/core/torspec/-/blob/main/p...
lim players -> inf
mean wealth -> starting_wealth * (win_chance * (1 + win_gain) + lose_chance * (1 - lose_loss)^rounds)
So in their example case win_change = lose_change = 0.5
win_gain = 1.5
lose_loss = 0.4
(0.5 * (1 + 0.5) + 0.5 * (1 - 0.4)) = 1.05
so on average the mean wealth increases as the number of rounds increases. Yet at the same time for a single player lim rounds -> inf
wealth -> 0
As others have mentioned, this is because the win multiplier is 1 + win_gain = 1.5 whereas the loss multiplier is 1 - lose_loss = 0.6. This is more easily seen by comparing the log of the multiplierln(1.5) = 0.41 ln(0.6) = -0.51
For the game to be profitable, it must satisfy
ln(1 + win_gain) > -ln(1 - lose_loss) <=> 1 + win_gain > 1/(1 - lose_loss)
So for a gain of 50% the loss must be less than 33.3%.