335 karma · joined May 2, 2026
Disclaimer: I cut it off after a few minutes because I got impatient. It could have gotten closer if I waited longer.
> 1024 bit keys with recordings of not too old data can probably be found
I think GitHub might turn into a scary vector of supply chain attacks in the foreseeable future. There is a five digit number of users still running around with 1024 bit RSA keys. echo 23ca73454a1b981fe51cad0dbd05f4e696795ba67abb28c61aea1a024e5bbeca | xxd -r -p | sha256sum
echo a16a8141361ae9834ad171ec28961fc8a951ff1bfc3a9ce0dc2fcdbdfa2ccd35 | xxd -r -p | sha256sum
From this post from 6 years ago: https://www.reddit.com/r/crypto/comments/guctw4/finding_sha2...EDIT: Found a workaround. Double-click the URL and paste it with middle-click.
> Why would I pick GLM over Claude?
To support the company that makes their model weights available for download, while Anthropic lobbies to restrict access.I have never had looping issues with DeepSeek models. Which provider/serving framework and harness are you using?
Eventually, the context window got so bloated that they resorted to hiding function bodies in large files, which is of course a stupid idea because then the LLMs have to use other tools to read the files, wasting even more tokens, or hallucinate the content. Honestly, it is amazing that LLMs work at all in VS Code.
You can inspect the context by pressing F1 and then selecting "Developer: Show Chat Debug View" in VS Code (https://github.com/microsoft/vscode/wiki/Copilot-Issues) and marvel at all the garbage that is in there.
To reduce system prompt and tool bloat, use a minimal harness (I wrote my own, but I've read that pi.dev is okay, too).
To make your prompts more precise, tell the LLM which files it has to read (or at least where it should start), so it does not have to search as much. This also reduces the change of misunderstandings and makes the LLM adhere to existing practices.
To keep your code base in check, tell the LLM (in a new session) to review the code and refactor from time to time.
When a task is done, start a new session. If you find that you have to repeat a lot of information in your next prompt, put the information in a file so you can reference it in the future (aka documentation).
GLM-5.3, GLM-5.3-Flash and Kimi K3 are also fine, but slower, more expensive, and less good for what I use them for, which is mostly Python and CUDA programming with some JS and HTML inbetween.
But almost all of my tasks are verifiable tasks, which means that the LLM can check whether it is done or whether it needs to keep trying. If you are mostly working on problems where the quality metric is based on vibes, YMMV.
I haven't tried Astra or Fable yet, because I am not made of money and am happy with my current setup. Also, the Opus models' writing is absolutely insufferable. My blood pressure rises every time I see a Claude-generated slop README.
> those things are not deterministic
Determinism was an explicit goal of DeepSeek-V4. From their paper: https://arxiv.org/html/2606.19348v1#S3.SS3 > we implement end-to-end, bitwise batch-invariant, and deterministic kernels with minimal performance overhead
Of course, providers may not implement deterministic inference for various reasons, but it is possible.Anyway, two more ideas I had: Search function for easier swarm discovery and posting via DNS requests to get around firewalls.
For comparison, I ran the same task with DeepSeek-V4-Flash, which finished in 172 seconds and cost $0.024 with a final context window size of 55217 tokens, while Qwen3.8-27B was not even close to being done with a 64178 context window.
This is a very efficient way to burn your money, but I would not recommend it for programming.
On the positive side, I got a $5 signup bonus, so it wasn't my own money.