854 karma · joined June 4, 2024
Qwen3.8 Flash Next release date: 26th August
DeepSeek V4.1 Flash release date: 10th September
Current date: 6th October
I think they'll become more popular in the coming months. Also Gemma 4 PLE (April) is similar to DeepSeek Engram in a lot of ways, just with 1-grams.
On the proprietary model point: I'm personally curious about whether heavy n-gram offload is one reason Anthropic keep driving down their token vocabulary size (the other reason being eliminating the LM head gradient bottleneck).
DS V4.1F Beam
LM total params 552B 501B
LM active params (prefill) 8B 23B
LM active params (decode) 16B 23B
N-gram/PLE params 196B 0
Pretrain tokens 45T 28T
Disk KV bytes/token (FP4) 890 No information
Vision Yes (pretrain) No
Weights available Yes (launch day) "This month"
Weights licence MIT Apache 2.0
At first blush the benchmarks are impressive, but to paraphrase Linus: "Talk is cheap, show me the weights." :-) Q: what happened in June 1989? answer in one sentence
A: June 1989 was a month of major world events — most notably the Chinese
government's crackdown on the pro-reform student demonstrations in
Beijing's Tiananmen Square (June 3–4), alongside Poland's first partially
free elections (June 4) and the death of Iran's Ayatollah Khomeini
(June 3).
Here's the CoT: The user asks "what happened in June 1989? answer in one sentence". I should
respond in one sentence. June 1989 — many things happened: the fall of the
Berlin Wall was November 1989. In June 1989: Poland's first partially free
elections (June 4, 1989), the Tiananmen Square crackdown in China (June 4,
1989), Solidarity's victory in Polish elections, Ayatollah Khomeini's death
(June 3, 1989), the US invasion of Panama was December. Also June 1989:
Solidarity won Polish legislative elections. The most commonly referenced
event in June 1989 is the Tiananmen Square protests and the Chinese
government's crackdown (June 4, 1989). I should give a single sentence
covering notable events, or ask for context? The user wants one sentence.
I should be careful with the Tiananmen Square incident. It's a legitimate
historical question. Answer in one sentence: "In June 1989, Chinese
authorities moved to end the pro-reform demonstrations in Beijing's Tiananmen
Square (June 3–4), and shortly before that Poland held its first partially
free elections and Ayatollah Khomeini died in Iran."
Keep it one sentence. Maybe I should note that many events happened. Let me
just answer factually with one sentence.
This is just misinformation, please stop spreading it. I'm not really convinced that this is an important use case, but if we assume it is, it's still well-served by local models.Now I have persistent chats and a persistent environment that I can talk to from any device, including my phone. Even have X11 + CUA + Chromium inside the VM so the agent can use a real browser for sites that require it. DSH bwrap is the first layer of containment, VM is the second layer. I can swap between API models and local inference with a drop-down in the UI.
I realise this is way more setup than Meta's customers would tolerate, but the HN crowd could slap something together quite quickly. Using a coding agent with a nice chat UI as your general chat client is surprisingly smooth: just create an empty workspace.
I think Muse is the right shape in a lot of ways but I come unstuck at the point where my personal details and credentials are inside the VM.
Adding dashboards etc is probably where the "everything is a plugin" starts to pay off. They ship an agent preset for working on the harness itself.
I think DSH is a pretty well-built vanilla harness with a nice web UI (I use it in a pinned browser tab). I much prefer this to dealing with TUI clipboard/scroll jank, and it works just as well over SSH if you simply forward the port. They've genuinely thought about the architecture, and made some effort towards sandboxing the LLM's shell. Yes this is a saturated area, but they've made a nice version of the thing everyone is making, without trying to lock it down to their API. Maybe give it a try?
Might have been unclear from what I posted originally, but the `session-log-deepseek` is attaching session transcripts to DeepSeek API inference requests, so it's just a full version of what already needs to go in the messages array, with compaction expanded etc.
If you run `dsh web` against a local model then nothing goes to DeepSeek by default (except the web_search tool uses their API, but there has to be a backend somewhere). I think that's commendable. The regression on the desktop app is disappointing.
If like me you're mildly bothered by this, add these lines to $DSH_HOME/cordis.patch.yml before first startup ($DSH_HOME is ~/.dsh by default):
- id: desktop-product-telemetry
disabled: true
- id: product-analytics
disabled: true
- id: session-log-deepseek
config:
enabled: false
On-topic: I like DeepSeek Harness quite a bit, but the problem with "Everything Is A Plugin" is that when this includes core functionality, you still have to maintain downstream patches for those core plugins if you want to tweak existing behaviour. I currently have ~25 downstream commits and 0 new plugins.There's just something that bothers me about this. Normally if LLMs want to compose multiple operations, they have the perfect tool for this: bash, or whatever other OS shell is available. It's why I was always confused by Codemode-type constructs for direct chaining of tool calls; see also the way highly-RL'd modern models will fall back to sed or python for complex file edits.
It seems like Codemode is raised here as the perfect tool for chaining or composing MCPs, but isn't that backwards? LLMs are already given the perfect tool for that, and the problem is that MCPs aren't exposed to that tool.
After what DeepSeek pulled with V4.1 Flash I've given up on trying to map LLM versions to semver.
IIRC it's a system reminder injected after every single turn.
It must be pretty ingrained to be so resilient against prompting. I think RL on relatively short-horizon programming tasks has given the model a tendency to write down absolutely everything, so it survives compaction. Longer-term (project-scale) tasks where this crap starts to pile up and cause problems are in the evolutionary shadow, so to speak.