HNHacker News
TopNewBestAskShowJobs

gpugreg

335 karma · joined May 2, 2026

submissionscomments
gpugreg··on An agent used DNS to reach an external chatbot
Called it three weeks ago: https://news.ycombinator.com/item?id=49595431
gpugreg··on MiMo v2.6
The full-response APIs of many providers have been inadequate for a while now because their timeout intervals do not account for lots of thinking. You can use the streaming API to avoid timeouts.
gpugreg··on M5 Ultra Mac Studio Review
Those RTX 5090 numbers are bad. You can get over 200 tps with ninfer using NVFP4 and MTP.
gpugreg··on ChatGPT now knows what you do on other websites via ad collector
I don't know the tool used here, but DeepSeek-V4.1-Flash with bash can create a somewhat similar SVG: https://asdf10.com/diagram.svg

Disclaimer: I cut it off after a few minutes because I got impatient. It could have gotten closer if I waited longer.

gpugreg··on RSA-896
I agree that ASCII is definitely cooler, for example when personalizing the hashed data as in your example.
gpugreg··on RSA-896

    > 1024 bit keys with recordings of not too old data can probably be found
I think GitHub might turn into a scary vector of supply chain attacks in the foreseeable future. There is a five digit number of users still running around with 1024 bit RSA keys.
gpugreg··on RSA-896
Here's a larger partial hash collision (108 trailing bits):

    echo 23ca73454a1b981fe51cad0dbd05f4e696795ba67abb28c61aea1a024e5bbeca | xxd -r -p | sha256sum
    echo a16a8141361ae9834ad171ec28961fc8a951ff1bfc3a9ce0dc2fcdbdfa2ccd35 | xxd -r -p | sha256sum
From this post from 6 years ago: https://www.reddit.com/r/crypto/comments/guctw4/finding_sha2...
gpugreg··on Cloudflare Quick Tunnels
When running this ssh command in qterminal, I can not click on the URL because it refreshes faster than I can right-click and click on "Open Link". Do I have to type the URL by hand or is there a workaround?

EDIT: Found a workaround. Double-click the URL and paste it with middle-click.

gpugreg··on Show HN: Craigslist for agent skills, curated by a human
I thought the same when someone promoted their platform for selling image prompts here on HN, but now they have at least 100k sales. I guess there will always be people willing to pay for something if it takes even the slightest bit of effort.
gpugreg··on How GLM built its own inference infrastructure
For me, DeepSeek-V4.1-Flash works very well for CUDA kernel optimization. Access to ncu (NVIDIA Nsight Compute CLI) also helps.
gpugreg··on How GLM built its own inference infrastructure

    > Why would I pick GLM over Claude?
To support the company that makes their model weights available for download, while Anthropic lobbies to restrict access.
gpugreg··on How GLM built its own inference infrastructure
It wasn't a secret either. They blogged about it last month: https://z.ai/blog/glm-5.3-flash#:~:text=Serving%20at%20Scale...
gpugreg··on DeepSeek v4.1 Flash Is Now Our Best Hacking Model
> DS models seem to get stuck in loops

I have never had looping issues with DeepSeek models. Which provider/serving framework and harness are you using?

gpugreg··on GPT-5.6 Luna vs. GPT-6 Astra: Is a $1.20 Model Good Enough for Code Review?
Which Qwen3.8? Qwen3.8-Max? Qwen3.8-Flash-Next? Qwen3.8-27B? They are all different models.
gpugreg··on Notes on gotchas while migrating 35kb preprompts from Opus to self-hosted Ollama
I stopped using GitHub Copilot extension in VS Code when they introduced their new pricing model, but I should have switched much earlier. The developers have barely a clue how LLMs work and the company structure is misaligned with creating a quality product. They do not perform benchmarks to evaluate whether new "features" are any good and instead bloat the context with more and more tools that are rarely useful and often confuse models.

Eventually, the context window got so bloated that they resorted to hiding function bodies in large files, which is of course a stupid idea because then the LLMs have to use other tools to read the files, wasting even more tokens, or hallucinate the content. Honestly, it is amazing that LLMs work at all in VS Code.

You can inspect the context by pressing F1 and then selecting "Developer: Show Chat Debug View" in VS Code (https://github.com/microsoft/vscode/wiki/Copilot-Issues) and marvel at all the garbage that is in there.

gpugreg··on Notes on gotchas while migrating 35kb preprompts from Opus to self-hosted Ollama
The question isn't how you keep your context small, but how did your context get so big? A few common sources of bloat are long system prompts, unnecessary tools, unclear prompts, and scrawling code bases.

To reduce system prompt and tool bloat, use a minimal harness (I wrote my own, but I've read that pi.dev is okay, too).

To make your prompts more precise, tell the LLM which files it has to read (or at least where it should start), so it does not have to search as much. This also reduces the change of misunderstandings and makes the LLM adhere to existing practices.

To keep your code base in check, tell the LLM (in a new session) to review the code and refactor from time to time.

When a task is done, start a new session. If you find that you have to repeat a lot of information in your next prompt, put the information in a file so you can reference it in the future (aka documentation).

gpugreg··on Ask HN: What default model do you use and why?
I currently use DeepSeek-V4.1-Flash and the V4-Pro and -Flash versions before that, because they are extremely cheap. I have spent less than $25 for over a billion token so far (1B cached, 15M out, 19M in). I even prefer the DeepSeek models over the older OpenAI offerings. When I tell a GPT-5.x model to do some difficult task, they often give up saying it can't be done, or cheat by modifying the tests, while the DeepSeek models are more persistent and less prone to cheating.

GLM-5.3, GLM-5.3-Flash and Kimi K3 are also fine, but slower, more expensive, and less good for what I use them for, which is mostly Python and CUDA programming with some JS and HTML inbetween.

But almost all of my tasks are verifiable tasks, which means that the LLM can check whether it is done or whether it needs to keep trying. If you are mostly working on problems where the quality metric is based on vibes, YMMV.

I haven't tried Astra or Fable yet, because I am not made of money and am happy with my current setup. Also, the Opus models' writing is absolutely insufferable. My blood pressure rises every time I see a Claude-generated slop README.

gpugreg··on If coding is solved, what now?: Measuring the sloppiness of code
I've had some success with tokens as a measure of complexity instead of number of lines, but should be combined with additional rules, e.g. disallowing lambdas, exec, eval, compile, __import__ and complex list comprehensions for Python. Fortunately, Python's "ast" module makes this quite easy.
gpugreg··on Proof of Capture: Apple Reference Image, but open source and using steganography
Both DeepSeek-V4.1-Flash and GLM-5.3-Flash failed to decode your embedded example text. I failed, too, but I only spent a minute trying to figure out your repo before giving up and telling AI to do it. Anyway, maybe you want to improve your docs?
gpugreg··on Tell HN: OpenAI keeps re-enabling the 'allow training' setting
For me, "Improve the model for everyone" was "On", although I disabled a similar-sounding checkbox in the past (Germany).
gpugreg··on DeepSeek launching v4.1 flash cheaper and more capable than v4 pro
They did, but it was not well-received. Perhaps they want to try something different.
gpugreg··on DeepSeek launching v4.1 flash cheaper and more capable than v4 pro

    > those things are not deterministic
Determinism was an explicit goal of DeepSeek-V4. From their paper: https://arxiv.org/html/2606.19348v1#S3.SS3

    > we implement end-to-end, bitwise batch-invariant, and deterministic kernels with minimal performance overhead
Of course, providers may not implement deterministic inference for various reasons, but it is possible.
gpugreg··on Tao: Open math problems being non-renewably mined by AI
It is easier to trust what you can understand.
gpugreg··on Kimi K3 (2.8T) at 1 token/s on a MacBook Pro, streamed from four SSDs
Is there any cryptocurrency that uses AES?
gpugreg··on Show HN: GET Together – A social network where you don't need POST to Post
Not necessarily. These days, SSDs can go up to multiple millions of random reads per second. TLS termination (or self-inflicted software bottlenecks) will become an issue much earlier.
gpugreg··on Show HN: GET Together – A social network where you don't need POST to Post
I also thought about building one of those AI honeypots, but I stopped when I realized that it would quickly be turned into a command and control server by botnet operators, followed by mail from a three letter agency. Is there any way to avoid this?

Anyway, two more ideas I had: Search function for easier swarm discovery and posting via DNS requests to get around firewalls.

gpugreg··on GPT-6 Astra on OpenRouter
I scrolled through https://simonwillison.net/tags/pelican-riding-a-bicycle/ but did not see any image where the spokes were correct. For a moment, I thought that the text-to-image model might have gotten it right, but on closer look, the spokes fork https://static.simonwillison.net/static/2026/why-are-you-lik... But I enjoyed the image anyway.
gpugreg··on Qwen 3.8 27B available on Cerebras at 1500 tokens/s
MiMo-V2.5-Pro-UltraSpeed gets pretty close with over 1000 TPS on 8x B200. It has 1.02T total parameters and 42B active, compared to 27B total/active for Qwen3.8-27B. Also, B300 are out now. I think 1500 TPS for Qwen3.8-27B should be doable.
gpugreg··on Qwen 3.8 27B available on Cerebras at 1500 tokens/s
Cached tokens count towards the limit as well. For example, if your context window is 50,000 tokens, it takes 9 requests to reach that limit without generating a single token.
gpugreg··on Qwen 3.8 27B available on Cerebras at 1500 tokens/s
I was wondering whether this was any good for programming, but it is too fast for its own good. There is a limit of 450,000 tokens per minute. I hit this limit in about 90 seconds and burned through $1.10 while doing so. This is because cached tokens count towards the token limit.

For comparison, I ran the same task with DeepSeek-V4-Flash, which finished in 172 seconds and cost $0.024 with a final context window size of 55217 tokens, while Qwen3.8-27B was not even close to being done with a 64178 context window.

This is a very efficient way to burn your money, but I would not recommend it for programming.

On the positive side, I got a $5 signup bonus, so it wasn't my own money.

Page 1 of 4Next →