1,106 karma · joined April 8, 2018
[ Prompt: 95.0 t/s | Generation: 26.5 t/s ]
(the test's prompt is very short but with longer ones the I get ~200 prompt processing rate, but I was hoping for a better token generation rate...)Am I understanding correctly that no draft model exists (will never exist or just currently does not exist yet)?
There is no draft file in Huggingface's repository ( https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf/tr... ) and in the file "scripts/download_models.sh" of the demo repository ( https://github.com/PrismML-Eng/Bonsai-demo/blob/main/scripts... ) I see this remark:
if [ "$_family" = "bonsai2" ]; then
the projector ships in the same repo; Bonsai 2 has no dspark drafterI started with "Ollama" (precompiled version) and it worked and was good enough to understand the very basics.
Then I downloaded the sourcecode of "llama.cpp", compiled it with specific compilation options for my GPUs (CUDA/nVidia using proprietary module on Gentoo Linux) & CPU (AMD), and the same model ran twice as fast -> since then I stuck with "llama.cpp" (and "ik_llama.cpp" in very few cases).
I honestly don't know what made "Ollama" (precompiled) so much slower than "llama.cpp" (compiled locally) at that time and I'm too lazy to doublecheck now, in any case I now absolutely love all the knobs that "llama.cpp" has to tune your hardware setup & your workload, which is the reason why I recommend it.
Qwen3.8 comes with official support for reasoning_effort, which can be used to adjust reasoning depth and control cost:
- xhigh (default): for complex tasks demanding thorough analysis
- medium: balancing accuracy and speed
- low: efficient reasoning optimizing for speed and cost
In addition, preserve_thinking is enabled by default for all workloads for the best out-of-the-box experience.
Asking because in my case (OCR of scanned historical "National Geographic" magazines) the LLM trying to merge text split into separate columns was running in circles from time to time and needed a lot of prompt tuning when using Qwen 3.0/3.5/3.6 (still needs from time to time).https://techcrunch.com/2026/08/07/spacexs-terafab-will-rely-...
I use it since ~20 years on most of my notebooks/PCs/root servers mainly because of its flexibility and I often end up wanting to use the very latest versions of software (e.g. ZFS, kernels, etc...) without struggling as most things are available from the official central package repository. I admit that updating/upgrading the OS is a pain from the time perspective (because of the long compilation times). As a programmer/developer the experience is as well fantastic as most of the libraries, compilers, interpreters, etc are very well integrated or easily installable.
I did try to use a few times Linux Arch, but I never fully liked its AUR.
For VMs I basically use always either Debian or Linux Mint with Xfce.
For gaming, mediacenter-PC and webradio-miniPC I use Linux Mint with Xfce.
It's funny that whenever I need assistance with some task (fixing a problem, doing something weird, ...) and I mention to Google Gemini something like "I'm using Gentoo Linux blahblah" then often Gemini starts its reply with something like "Ah, a picky Gentoorian, I see - then let's do this the complicated way..." :o)
ESTA = "Electronic Swatch Timepiece Application", as pun/mockery of the US "Electronic System for Travel Authorization".
https://nmon.sourceforge.io/pmwiki.php
Especially disk throughput and I/O (keys "d" & "D") can be very useful.
Maybe your LAN cables have deteriorated?
Reaching 100 MiB/s should be easy peasy for a PC; in my case (2 10Gbps switches) by using nvme SSDs I reach almost 950 MiB/s and with 4x raidz1 HDDs I reach 450-650 MiB/s.
Until a few years ago I was smoking in my flat and often I had to change lan cables (whatever's in the smoke was sticking to the connectors)
Oh yes, I'd love them too (if you're referring to, in Oracle slang, "...update on commit") - and it would be cool to have as well the option for a lazy update ("on demand" by taking into consideration only the records that have been changed since the last refresh, to handle multiple updates in a single pass - not sure how Oracle can achieve that technically...). This would be in my opinion a fantastic added functionality compared to basically all other (OLTP?) opensource DBs.
And: I'm really curious about the "OrioleDB" project... ( https://github.com/orioledb/orioledb/releases ) as a few years ago I was struggling a lot with "vacuum" of a kind-of-temporary table that had quite high amounts of continuous random inserts & deletes (problem solved by accumulating more changes in RAM before flushing them to the table therefore increasing amount of rows changed per "page", but I had to sweat a lot to find a good balance...).
E.g. when doing text transcription/OCR from images (Qwen 3.6 27B Q4_K_M by Bartowski) with a context size of ~50k I get a pp of ~460 tokens per second and a generation ranging from 35 to 45 tokens per second (using "--spec-type draft-mtp --spec-draft-n-max 2" currently with llama.cpp b6548).
On the other hand when handling code (Qwen 3.6 27B Q5_K_M by Bartowski) with a context size of 128k I get a pp ranging between 500 to 1500 tokens per second and a generation between 25 and 40 tokens per second (using in this case as well "--spec-type draft-mtp --spec-draft-n-max 2" currently with llama.cpp b6548).
Anyway in theory with "--split-mode layer" I think that it's anyway the slowest card that drives the overall performance (I do see in "nvtop" that usually the 5070 is ~25% active, the 5060 ~50% and the 3060 ~75%).
I'm currently running those models using an RTX 5070 12GiB + RTX 5060 16GiB + RTX 3060 12GiB with a 96k context size with MTP/speculative decoding and I'm quite happy (the 5070 is about 4x faster than the 3060, the 5060 is inbetween them so about 2x faster than a 3060).
(you're right - I was wondering the same thing 1h ago :o) )
Right now the biggest threat to their IPO's is that people realize that local models are good enough for whatever they're peddling...
...plus the recent price increases by AI companies, made me actually think the opposite: that there might be another additional "run" for memory and/or GPUs.
Therefore, yesterday I decided to order an additional RTX 5060 with 16 GiB VRAM for the ~500$ that I saved during the last months (to be added to the RTX 5070 12 GiB that I bought last year to play games in 4k + my old RTX 3060 12 GiB which I recycled a few months ago after noticing how nice it is to run llama.cpp locally without having to worry about subscription costs).
The original 24 GiB VRAM were actually quite enough for some of the stuff that I do (e.g. transcribe text of image scans of old magazines, coding with Aider, etc - I usually use Q5_K_M quantizations of Qwen & Gemma by Bartowski as lower ones delivered sometimes weird results and/or looped forever in "thinking"-mode), but I guess that with 40 GiB I should be bullet-proof for my pessimistic view of our future :o)
- esp4 (kernel config "CONFIG_AF_RXRPC")
- esp6 (kernel config "CONFIG_INET_ESP")
- rxrpc (kernel config "CONFIG_INET6_ESP")
Is this correct?
I was running in Gentoo "6.18.18" (amd64) and the exploit worked (and all other shells which I PREVIOUSLY opened could then just execute "su -" without password to become "root") -> doing temporarily a "modprobe -r algif_aead" on-the-fly did not fix it as I was still able to swap to "root" from the unprivileged user by executing just "su -".
"6.18.25" fixed it (module "algif_aead" still running).
- Maybe older Kernel versions that don't contain the fix should be blacklisted?
- FYI in Gentoo I had to recompile "sys-fs/zfs-kmod" after the minor kernel upgrade (I initially skipped it, but after rebooting with the new kernel I could not mount my raidz1) -> the same might be needed for other external modules.
My performance when using an RTX 5070 12GiB VRAM, Ryzen 7 9700X 8 cores CPU, 32GiB DDR5 6000MT (2 sticks):
- "qwen2.5:7b": ~128 tokens/second (this model fits 100% in the VRAM).
- "qwen2.5:32b": ~4.6 tokens/second.
- "qwen3:30b-a3b": ~42 tokens/second (this is a MoE model with multiple specialized "brains") (this uses all 12GiB VRAM + 9GiB system RAM, but the GPU usage during tests is only ~25%).
- qwen3.5:35b-a3b: ~17 tokens/second, but it's highly unstable and crashes -> currently not usable for me.
So currently my sweet spot is "qwen3:30b-a3b" - even if the model doesn't completely fit on the GPU it's still fast enough. "qwen3.5" was disappointing so far, but maybe things will change in the future (maybe Ollama needs some special optimizations for the 3.5-series?).I would therefore deduce that the most important thing is the amount of VRAM and that performance would be similar even when using an older GPU (e.g. an RTX 3060 with as well 12GiB RAM)?
Performance without a GPU, tested by using a Ryzen 9 5950X 16 cores CPU, 128GiB DDR4 3200 MT:
- "qwen2.5:7b": ~9 tokens/second
- "qwen3:32b": ~2 tokens/second
- "qwen3:30b-a3b": ~16 tokens/secondCannot be - there is no competition in Switzerland, but things run pretty smoothly -> in the case of Germany I'd rather say: "lack of oversight, controls, 'konsequent zu sein'" -> in the case of Germany's DB I think that nobody at all levels gives a *hit about its problems.
Maybe - I guess that they must have served that "cached" content from DB-records that had it all saved directly (URL X has contents Y => basically a "mirror" of the terms that they indexed) => not having to store that "mirror" (only the search index) might save quite a lot of storage space (and I/O and CPU to decompress it, as users won't be requesting it anymore) => all in all that might save quite a lot of infrastructure costs $$$.
> Could this be an advantage that Google can use to train their models on but others won't have access?
Maybe (if they decided to just get rid of the I/O related to the user requests), but on the other hand I don't know if previously any "Google-consumer" was ever able to perform mass-downloads of Google's "cached" data - could that be done without being banned by Google's webpage (or API)?