HNHacker News
TopNewBestAskShowJobs

wren6991

854 karma · joined June 4, 2024

submissionscomments
wren6991··on Mistral Large 4
The benchmark is saturated. Frontier models are tested with an armadillo in fishnet tights jaywalking on Mars.
wren6991··on Beam: Reflection's 501B open-weight model
DeepSeek Engram paper published: 12th January

Qwen3.8 Flash Next release date: 26th August

DeepSeek V4.1 Flash release date: 10th September

Current date: 6th October

I think they'll become more popular in the coming months. Also Gemma 4 PLE (April) is similar to DeepSeek Engram in a lot of ways, just with 1-grams.

On the proprietary model point: I'm personally curious about whether heavy n-gram offload is one reason Anthropic keep driving down their token vocabulary size (the other reason being eliminating the LM head gradient bottleneck).

wren6991··on Beam: Reflection's 501B open-weight model
I heard you should ask Claude about this. Preferably with thousands of accounts, routed through residential proxies
wren6991··on Beam: Reflection's 501B open-weight model
I thought it would be interesting to look at some key figures vs another contemporary model in the same weight class (DeepSeek V4.1 Flash)

                                DS V4.1F            Beam
    LM total params             552B                501B
    LM active params (prefill)  8B                  23B
    LM active params (decode)   16B                 23B
    N-gram/PLE params           196B                0
    Pretrain tokens             45T                 28T
    Disk KV bytes/token (FP4)   890                 No information
    Vision                      Yes (pretrain)      No
    Weights available           Yes (launch day)    "This month"
    Weights licence             MIT                 Apache 2.0
At first blush the benchmarks are impressive, but to paraphrase Linus: "Talk is cheap, show me the weights." :-)
wren6991··on Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
Great, then they will PII-scrub my sessions before feeding them into the training pool :-)
wren6991··on Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
Claude remove all punctuation so it looks like I wrote it myself. Yeah even the hyphens for adjectival compounds, fuck 'em, it's all punctuation so it's gotta go
wren6991··on Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
I see this kind of message a lot from people who haven't tried it. They have knowledge, just are trained to refuse to talk about it. Here's an abliterated Qwen3.8 Flash Next:

    Q: what happened in June 1989? answer in one sentence

    A: June 1989 was a month of major world events — most notably the Chinese
    government's crackdown on the pro-reform student demonstrations in
    Beijing's Tiananmen Square (June 3–4), alongside Poland's first partially
    free elections (June 4) and the death of Iran's Ayatollah Khomeini
    (June 3).
Here's the CoT:

    The user asks "what happened in June 1989? answer in one sentence". I should
    respond in one sentence. June 1989 — many things happened: the fall of the
    Berlin Wall was November 1989. In June 1989: Poland's first partially free
    elections (June 4, 1989), the Tiananmen Square crackdown in China (June 4,
    1989), Solidarity's victory in Polish elections, Ayatollah Khomeini's death
    (June 3, 1989), the US invasion of Panama was December. Also June 1989:
    Solidarity won Polish legislative elections. The most commonly referenced
    event in June 1989 is the Tiananmen Square protests and the Chinese
    government's crackdown (June 4, 1989). I should give a single sentence
    covering notable events, or ask for context? The user wants one sentence.

    I should be careful with the Tiananmen Square incident. It's a legitimate
    historical question. Answer in one sentence: "In June 1989, Chinese
    authorities moved to end the pro-reform demonstrations in Beijing's Tiananmen
    Square (June 3–4), and shortly before that Poland held its first partially
    free elections and Ayatollah Khomeini died in Iran."

    Keep it one sentence. Maybe I should note that many events happened. Let me
    just answer factually with one sentence.
This is just misinformation, please stop spreading it. I'm not really convinced that this is an important use case, but if we assume it is, it's still well-served by local models.
wren6991··on What Meta got right with Muse
My personal solution to this is to run DSH in an Incus VM on a local machine, then expose the web UI to the internet through my VPS (with some extra auth). I also made a couple of UI tweaks so it works better as a PWA.

Now I have persistent chats and a persistent environment that I can talk to from any device, including my phone. Even have X11 + CUA + Chromium inside the VM so the agent can use a real browser for sites that require it. DSH bwrap is the first layer of containment, VM is the second layer. I can swap between API models and local inference with a drop-down in the UI.

I realise this is way more setup than Meta's customers would tolerate, but the HN crowd could slap something together quite quickly. Using a coding agent with a nice chat UI as your general chat client is surprisingly smooth: just create an empty workspace.

I think Muse is the right shape in a lot of ways but I come unstuck at the point where my personal details and credentials are inside the VM.

wren6991··on DeepSeek Harness Desktop for macOS and Windows
Did you look into the workflow tools? The agent can write a TypeScript program that encodes a subagent graph, so you can codify a review workflow and ensure it's actually followed. Kind of neat, I think other providers also have these now, but it was new to me, and interesting to see how it worked.

Adding dashboards etc is probably where the "everything is a plugin" starts to pay off. They ship an agent preset for working on the harness itself.

wren6991··on DeepSeek Harness Desktop for macOS and Windows
This harness supports pretty much any provider out of the box (it ships Pi LLM SDK), and DeepSeek's models are hosted by a number of providers.
wren6991··on DeepSeek Harness Desktop for macOS and Windows
It can't be long now until humans have to ask (vision) LLMs for help with CAPTCHAs.
wren6991··on DeepSeek Harness Desktop for macOS and Windows
Combination of main character syndrome, and wanting consistency across platforms like MacOS, which doesn't define $XDG_CONFIG_HOME or $XDG_CACHE_HOME.
wren6991··on DeepSeek Harness Desktop for macOS and Windows
Seems like a category error: you can use DSH with openrouter, in fact it ships built-in support using the Pi LLM SDK.

I think DSH is a pretty well-built vanilla harness with a nice web UI (I use it in a pinned browser tab). I much prefer this to dealing with TUI clipboard/scroll jank, and it works just as well over SSH if you simply forward the port. They've genuinely thought about the architecture, and made some effort towards sandboxing the LLM's shell. Yes this is a saturated area, but they've made a nice version of the thing everyone is making, without trying to lock it down to their API. Maybe give it a try?

wren6991··on DeepSeek Harness Desktop for macOS and Windows
Yeah, I am actually ok with this. I'm selective about what goes into it, and the training goes into better open-weight models. From what I've seen the DSH desktop telemetry is actually telemetry (like VS Code), not uploading all your files.

Might have been unclear from what I posted originally, but the `session-log-deepseek` is attaching session transcripts to DeepSeek API inference requests, so it's just a full version of what already needs to go in the messages array, with compaction expanded etc.

If you run `dsh web` against a local model then nothing goes to DeepSeek by default (except the web_search tool uses their API, but there has to be a backend somewhere). I think that's commendable. The regression on the desktop app is disappointing.

wren6991··on DeepSeek Harness Desktop for macOS and Windows
I don't think DSH is doing anything close to what Grok was doing.
wren6991··on DeepSeek Harness Desktop for macOS and Windows
Nature wants to create crab
wren6991··on DeepSeek Harness Desktop for macOS and Windows
The desktop build seems to enable telemetry by default. Regular `dsh web` only has telemetry for explicit user feedback. Hmm :/

If like me you're mildly bothered by this, add these lines to $DSH_HOME/cordis.patch.yml before first startup ($DSH_HOME is ~/.dsh by default):

    - id: desktop-product-telemetry
      disabled: true
    - id: product-analytics
      disabled: true
    - id: session-log-deepseek
      config:
        enabled: false
On-topic: I like DeepSeek Harness quite a bit, but the problem with "Everything Is A Plugin" is that when this includes core functionality, you still have to maintain downstream patches for those core plugins if you want to tweak existing behaviour. I currently have ~25 downstream commits and 0 new plugins.
wren6991··on Context Language Models
Kimi K3 is an example of a modern LLM that doesn't use positional encodings. It uses NoPE'd MLA for global attention and a variant of Gated DeltaNet (KDA) for local attention. I wonder how much the KDA would degrade if you just did a bounded replay over the last few thousand tokens to recover its approximate state when assembling your context window from chunks of known MLA KV.
wren6991··on Gemini 4 Argon
DeepSeek is the immortal whale
wren6991··on You Said No MCP
Right, because MCP != web. In fact most of my exposure to MCP has been local tools. Also, for context, we're talking about a coding agent harness that runs on a local machine, and Codemode applies to all of its tools, not just MCP.
wren6991··on The AI Race Just Got Awkward
The doublethink required to simultaneously believe "our safeguards prevent our models from doing unsanctioned cybersecurity tasks" and "distillation is why Chinese models are getting better at cybersecurity tasks" is genuinely quite funny.
wren6991··on Pi.dev: You Said No MCP
Yeah, I'm probably over-indexing on local use cases due to my own preferences, prejudices, biases etc. For a coding agent like Pi it does seem reasonable to expect some kind of shell access though, unless some people are using it as a CLI chat client with MCP?
wren6991··on Solving Factorio Quality
Very easy to optimise the fun out of Factorio. I have a hard rule: no blueprints carried forward between saves. It's more fun for me to figure out what I can design quickly that kinda works and kinda supports scaling. A lot of the time I'll just build something as a repeatable unit cell and then use copy/paste, skipping the blueprints. I think it's worth trying this out.
wren6991··on You said no MCP
> And while we could have just wired up the metadata to enable better MCP extensions, we also think that MCP with Codemode solves quite a few of the issues that it traditionally had.

There's just something that bothers me about this. Normally if LLMs want to compose multiple operations, they have the perfect tool for this: bash, or whatever other OS shell is available. It's why I was always confused by Codemode-type constructs for direct chaining of tool calls; see also the way highly-RL'd modern models will fall back to sed or python for complex file edits.

It seems like Codemode is raised here as the perfect tool for chaining or composing MCPs, but isn't that backwards? LLMs are already given the perfect tool for that, and the problem is that MCPs aren't exposed to that tool.

wren6991··on GLM-5.3 and the spread of advanced cyber capabilities
I need Anthropic's employees to understand that refusing to fix vulnerabilities in code you just wrote is not a morally neutral position.
wren6991··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
It's just vibe versioning, right? Fable 5 is a beloved product, it gets a .1 bump to feel close. Opus 5 and Sonnet 5 had a mixed reception, they get a .5 bump to create a sense of distance.

After what DeepSeek pulled with V4.1 Flash I've given up on trying to map LLM versions to semver.

wren6991··on Sonnet 5.5
Hey now, if programming is solved they still need to achieve vendor lock-in somehow. Won't somebody please think of the vendors?
wren6991··on Prompting Claude Opus 5.5
Does Claude Code not preempt block-on-output when you send a steering message? This was one of the first things I fixed in my DeepSeek Harness fork
wren6991··on Prompting Claude Opus 5.5
> I am told this is merely something in the system prompt that the model tends not to pay attention to with large contexts.

IIRC it's a system reminder injected after every single turn.

It must be pretty ingrained to be so resilient against prompting. I think RL on relatively short-horizon programming tasks has given the model a tendency to write down absolutely everything, so it survives compaction. Longer-term (project-scale) tasks where this crap starts to pile up and cause problems are in the evolutionary shadow, so to speak.

wren6991··on Meta VR Glasses
I'm honestly not sure. Often when companies lobby for regulation it's to avoid other types of regulation they would find more troublesome. I don't buy the popular narrative that Meta is lobbying for verification laws for the purpose of getting ID data because surely they have a lot of signal there already. At the same time I think it's backwards to say "Meta are behaving this way because of regulation" when they actively lobbied for it.
Page 1 of 8Next →