"Affected products: ... Up to (excluding) 152.0.7977.82"
.82 is fixed.
1,693 karma · joined March 16, 2010
web: https://www.thanassis.space
CV: https://www.thanassis.space/cv.pdf
"Affected products: ... Up to (excluding) 152.0.7977.82"
.82 is fixed.
Instructions to reproduce, and benchmarks here: https://forums.developer.nvidia.com/t/deepseek-v4-flash-offi...
But mostly I wanted to raise awareness to readers of your article that no, if you want to do inference, paying 15K for a single 96GB card almost certainly makes no sense. Buy 4 GX10s with the same money, and enjoy dramatically better models and user scalability.
Regardless - thanks for putting the effort to share your findings! I keep postponing doing the same... there's tons of things everyone is re-discovering on their own.
IMHO, the author could have done two things better:
- vllm instead of llama.cpp. With NVIDIA HW, there is huge difference in multi-user loads and caching with vllm; when he was complaining about what happens when more than one user uses the model, and about losing caching, I was "well, duh".
- The budget he used for a single card could have instead be put to far, far better use with SPARKs. I have access to a cluster of 2 x GX10 - total cost less than half what he paid, even today - and I am running vllm and Deepseek v4 Flash. The difference compared to any Qwen is tremendous - I've NEVER seen it loop, and in all my experiments so far, it's the most Sonnet-y model I've ever tried (antirez seems to agree, hence his ds4 fork).
If you're wondering about how I set it up in the 2 GX10s: https://forums.developer.nvidia.com/t/deepseek-v4-flash-offi...
Performance: 2K t/s prefill ( very useful for feeding tons of source code into its massive context window ) and around 50-60 tg/s in my coding sessions in the pi.dev harness. With the money the author paid, he could have bought 4 GX10s, and double both numbers ( vllm basically scales almost linearly with tensor parallelism ).
It's very nice seeing it put to use in actual Spectrum machines - love it :-)
Thank you so much! :-)
But I was curious for your approach... so I asked Claude to convert it to bash: https://claude.ai/public/artifacts/01a49347-1617-4afe-8476-0...
Works like a charm - pinned it to Ctrl-k, which was free in my setup. I guess I don't have to depend on XTerm for this any more :-)
Thanks!
I became so obsessed with the project that I was looking forward to tinkering with it after coming back from work every day; so it was hacked in 5 evenings and a weekend. It was that much fun, to build a Forth.
I highly recommend the process; I think the only other time I felt so enlightened was when I first met Lisp macros (https://www.thanassis.space/score4.html#lisp).
I just changed my server, and uninstalled the - now truly useless - fail2ban. I use SSH keys of course, but without fail2ban my server's logs were constantly flooded with hacking attempts.
No longer - wireguard for the win. Thank you, chlorion!
Many thanks, OP.
Both of you are :-)
Heart=warmed. Happy holidays to both of you!
No, mate - it wouldn't. Don't be confused by the "printf"-dump; the Python script processed it into pairs of (frequency,delay). That is, when you see...
989 Hz @ 15209
989 Hz @ 15213
989 Hz @ 15218
784 Hz @ 15222
784 Hz @ 15226
...the data actually generated are: (989, 15222-15209),
(784, ... -15222),
Simply put: RLE can't do anything on them. The repetition has already been "cleaned out".I wouldn't have gone to Huffman compression if I had a simpler choice.
Well, if the bird in question used an LZ77 followed by Huffman, he could compress all that wood/chuck stuff down to almost nothing. So he could chuck a lot :-D
Live and learn :-)
I am seriously considering playing it again with them this summer :-D
For anyone else interested, the same solution applies to shellcheck - just install shellcheck-bin from AUR.
Direct link to shadertoy: https://www.shadertoy.com/view/ssBGRG
Youtube video with intermediate steps: https://www.youtube.com/watch?v=RJf2nVOkOb4&t=1s
while true ; do sleep 1 ; done
You'll see that after 'fg', the loop ends :-)Simply put: C-z followed by fg is not bulletproof. Not to mention that I had no idea what I was running in there, and how any signal would impact it... So I wanted to find a safer way to dump what was already there, in my shell's memory.
Anyway, I hope you guys enjoyed reading this regardless :-)
But TBH, keeping that script (that I wrote in 60 seconds) nice and clean was the least of my worries... Making the damn toolchain work had far, far higher priority :-)
And after that, booting the LEON cores of course.