HNHacker News
TopNewBestAskShowJobs

mksglu

295 karma · joined March 7, 2023

mksg.lu
submissionscomments
mksglu··on [dead]
OpenAI published "Harness Engineering" in February 2026. The thesis: engineers don't write code anymore. They design environments, specify intent, and build feedback loops. Agents do the rest. Codex proved the concept. Symphony showed the architecture in Elixir/OTP. I wanted the same thing with Claude Code.

So I built Hatice.

mksglu··on [dead]
OpenAI published "Harness Engineering" in February 2026. The thesis: engineers don't write code anymore. They design environments, specify intent, and build feedback loops. Agents do the rest.

Codex proved the concept. Symphony showed the architecture in Elixir/OTP. I wanted the same thing with Claude Code.

So I built Hatice.

mksglu··on [dead]
Six days ago I posted a write-up on Hacker News explaining how Context Mode reduces Claude Code's context consumption by 98%. I expected a handful of comments. I got 565 points, 107 comments, and the kind of feedback that makes you stay up until 4 AM shipping features.

What's new:

  - Five platform adapters: Claude Code, Gemini CLI, VS Code Copilot, OpenCode, Codex CLI
  - Session Continuity: survives context compactions, restores state automatically (~3h sessions vs ~30 min)
  - ctx_batch_execute: one call replaces 30+ individual commands
  - Three-layer fuzzy search: Porter stemming, trigram substring, Levenshtein
  - FTS5 deduplication, background mode, SSL cert auto-detection

Breaking: repo renamed claude-context-mode → context-mode, tools prefixed with ctx_.

  Install (Claude Code): /install context-mode
  Install (all platforms): npm install -g context-mode

  GitHub: https://github.com/mksglu/context-mode
  Release notes: https://github.com/mksglu/context-mode/releases/tag/v1.0.0

Four adapters are in beta — PRs welcome. The codebase is designed for Agentic Engineering: point your coding agent at it and let it propose fixes.

Thank you to everyone who commented, starred, and opened issues on the original post. You made this release happen.

mksglu··on MCP server that reduces Claude Code context consumption by 98%
Thanks, really appreciate hearing that! Glad it's working well for your team.
mksglu··on MCP server that reduces Claude Code context consumption by 98%
Yeah it's basically pre-compaction, you're right. The key difference is nothing gets thrown away. The full output sits in a searchable FTS5 index, so if the model realizes it needs some detail it missed in the summary, it can search for it. It's less "decide what's relevant upfront" and more "give me the summary now, let me come back for specifics later."
mksglu··on MCP server that reduces Claude Code context consumption by 98%
That's the theory and it does hold up in practice. When context is 70% raw logs and snapshots, the model starts losing track of the actual task. We haven't run formal benchmarks on answer quality yet, mostly focused on measuring token savings. But anecdotally the biggest win is sessions lasting longer before compaction kicks in, which means the model keeps its full conversation history and makes fewer mistakes from lost context.
mksglu··on MCP server that reduces Claude Code context consumption by 98%
That's a fair point and honestly the ideal approach. But in practice most people don't hand-curate their MCP server list per task. They install 5-6 servers and suddenly have 80 tools loaded by default. Context-mode doesn't solve the tool definition bloat, that's the input side problem. It handles the output side, when those tools actually run and dump data back. Even with a focused set of tools, a single Playwright snapshot or git log can burn 50k tokens. That's what gets sandboxed.
mksglu··on MCP server that reduces Claude Code context consumption by 98%
It doesn't break the cache. The raw data never enters the conversation history, so there's nothing to invalidate. A short summary goes into context instead of the full payload, and the model can search the full data from a local FTS5 index if it needs specifics later. Cache stays intact because you're just appending smaller messages to the conversation.
mksglu··on MCP server that reduces Claude Code context consumption by 98%
Nice approach. Same core idea as context-mode but specialized for your build domain. You're using SQLite as a structured knowledge cache over YAML rule files with keyword lookup. Context-mode does something similar but domain-agnostic, using FTS5 with BM25 ranking so any tool output becomes searchable without needing predefined schemas. Cool to see the pattern emerge independently from a completely different use case.
mksglu··on MCP server that reduces Claude Code context consumption by 98%
That's true, Claude Code does truncate large outputs now. But 25k tokens is still a lot, especially when you're running multiple tools back to back. Three or four Playwright snapshots or a batch of GitHub issues and you've burned 100k tokens on raw data you only needed a few lines from. Context-mode typically brings that down to 1-2k per call while keeping the full output searchable if you need it later.
mksglu··on MCP server that reduces Claude Code context consumption by 98%
Haven't looked at rtk closely but from the description it sounds like it works at the CLI output level, trimming stdout before it reaches the model. Context-mode goes a bit further since it also indexes the full output into a searchable FTS5 database, so the model can query specific parts later instead of just losing them. It's less about trimming and more about replacing a raw dump with a summary plus on-demand retrieval.
mksglu··on MCP server that reduces Claude Code context consumption by 98%
Nope. The raw data never enters the conversation history in the first place, so there's nothing to invalidate. Tool output runs in a sandbox, a short summary comes back, and the full data sits in a local FTS5 index. The conversation cache stays intact because the context itself doesn't change after the fact.
mksglu··on MCP server that reduces Claude Code context consumption by 98%
Right, context-mode doesn't change how MCP tool definitions get loaded into context. That's the "input side" problem that Cloudflare's Code Mode tackles by compressing tool schemas. Context-mode handles the "output side," the data that comes back from tool calls. That said, if you're writing your own MCPs, you could apply the same pattern directly. Instead of returning raw payloads, have your MCP server return a compact summary and store the full output somewhere queryable. Context-mode just generalizes that so you don't have to rebuild it per server.
mksglu··on MCP server that reduces Claude Code context consumption by 98%
The people who spent years doing the work manually are the ones who immediately see where the bottlenecks are.
mksglu··on MCP server that reduces Claude Code context consumption by 98%
Good point on prompt cache invalidation. Context-mode sidesteps this by never letting the bloat in to begin with, rather than snipping it out after. Tool output runs in a sandbox, a short summary enters context, and the raw data sits in a local search index. No cache busting because the big payload never hits the conversation history in the first place.
mksglu··on MCP server that reduces Claude Code context consumption by 98%
That's pretty much the approach we took with context-mode. Tool outputs get processed in a sandbox, only a stub summary comes back into context, and the full details stay in a searchable FTS5 index the model can query on demand. Not trained into the model itself, but gets you most of the way there as a plugin today.
mksglu··on MCP server that reduces Claude Code context consumption by 98%
That's exactly what context-mode does for tool outputs. Instead of dumping raw logs and snapshots into context, it runs them in a sandbox and only returns a summary. The full data stays in a local FTS5 index so you can search it later when you need specifics.
mksglu··on MCP server that reduces Claude Code context consumption by 98%
Totally agree. Failed attempts are just noise once the right path is found. Auto-detecting retry patterns and pruning them down to the final working version feels very doable, especially for clear cases like lint or compilation fixes.
mksglu··on Show HN: SecLaw – Self-hosted AI agents on your machine, Docker-isolated
Author here. I built this after seeing OpenClaw (68K stars) give agents full access to ~/.ssh, ~/.aws, and browser cookies with zero container isolation.

  SecLaw runs 4 Docker containers with strict boundaries: non-root,
  cap_drop ALL, read-only filesystem, 512MB/1CPU limits per container,
  zero inbound ports via Cloudflare Tunnel. API keys are sealed per service,
  not shared across containers.

  The interesting part is multi-agent auto-routing. You install agents as
  templates (npx seclaw add inbox-agent, npx seclaw add research-agent)
  and they stack onto one Telegram bot. The LLM routes each message to the
  right capability — email questions go to Inbox, lead questions go to Sales.
  Every response shows which agent answered.

  Architecture: Node.js agent + Inngest for scheduled workflows + Desktop
  Commander (MCP server, read-only) + Cloudflare Tunnel. All orchestrated
  by a single CLI command.

  Setup is `npx seclaw` — walks you through LLM provider, API key, Telegram
  token, and runs docker compose up. 60 seconds, no YAML editing.
mksglu··on MCP server that reduces Claude Code context consumption by 98%
No magic — standard Unix process inheritance. Each execute() spawns a child process via Node's child_process.spawn() with a curated env built by #buildSafeEnv (https://github.com/mksglu/claude-context-mode/blob/main/cont...). It passes through an explicit allowlist of auth vars (GH_TOKEN, AWS_ACCESS_KEY_ID, GOOGLE_APPLICATION_CREDENTIALS, KUBECONFIG, etc.) plus HOME and XDG paths so CLI tools find their config files on disk. No state persists between calls — each subprocess inherits credentials from the MCP server's environment, runs, and exits. This works because tools like gh and aws resolve auth on every invocation anyway (env vars or ~/.config files). The tradeoff is intentional: allowlist over full process.env so the sandbox doesn't leak unrelated vars.
mksglu··on MCP server that reduces Claude Code context consumption by 98%
Author here. I shared the GitHub repo a few days ago (https://news.ycombinator.com/item?id=47148025) and got great feedback. This is the writeup explaining the architecture.

The core idea: every MCP tool call dumps raw data into your 200K context window. Context Mode spawns isolated subprocesses — only stdout enters context. No LLM calls, purely algorithmic: SQLite FTS5 with BM25 ranking and Porter stemming.

Since the last post we've seen 228 stars and some real-world usage data. The biggest surprise was how much subagent routing matters — auto-upgrading Bash subagents to general-purpose so they can use batch_execute instead of flooding context with raw output.

Source: https://github.com/mksglu/claude-context-mode Happy to answer any architecture questions.

mksglu··on Show HN: Context Mode – 315 KB of MCP output becomes 5.4 KB in Claude Code
Context Mode doesn't replace your other MCP servers — it sits alongside them. Your Context7, Playwright, GitHub servers all stay installed and work normally. The hook intercepts output-heavy tool calls (like WebFetch, curl) and redirects them through the sandbox. For example, instead of WebFetch dumping 56KB of raw HTML into context, the hook blocks it and tells the model to use fetch_and_index instead — which fetches the same URL but indexes it in a local SQLite DB, returning only a 3KB summary.

Your other MCP servers still run. Context Mode just gives the model a more context-efficient way to process their results when the output would be large.

mksglu··on Show HN: Context Mode – 315 KB of MCP output becomes 5.4 KB in Claude Code
That's a known bug in older versions — the WebFetch hook wasn't blocking reliably. Fixed in v0.7.1.

npm install -g context-mode@latest

If you're on the plugin install, re-run:

  /plugin marketplace add mksglu/claude-context-mode
  /plugin install context-mode@claude-context-mode
Then restart Claude Code. Sorry about that.
mksglu··on Show HN: Context Mode – 315 KB of MCP output becomes 5.4 KB in Claude Code
On Tantivy: Agree it's the better search engine, but context-mode is session-scoped — DB is a temp file that dies when the process exits. At that scale (50-200 chunks), FTS5 is zero-config, single-file, <1ms startup, and good enough. If we ever add persistent cross-session indexing, Tantivy would be the move.

On benchmarking: This is the experiment I most want to see. The hypothesis: context-mode benefits smaller models disproportionately — a 32K model with clean context could outperform a 200K model drowning in raw tool output. Would love to see SWE-bench results with context-mode on vs. off across model tiers.

mksglu··on Show HN: Context Mode – 315 KB of MCP output becomes 5.4 KB in Claude Code
Yes — the database is tied to the MCP server process, so it's created fresh on each claude launch and lost when you exit; resuming a session starts a new process with a new empty database.
mksglu··on Show HN: Context Mode – 315 KB of MCP output becomes 5.4 KB in Claude Code
Good question.

The SQLite database is ephemeral — stored in the OS temp directory (/tmp/context-mode-{pid}.db) and scoped to the session process. Nothing persists after the session ends. For sensitive data masking specifically: right now the raw data never leaves the sandbox (it stays in the subprocess or the temp SQLite store), and only stdout summaries enter the conversation. But a dedicated redaction layer (regex-based PII stripping before indexing) is an interesting idea worth exploring. Would be a clean addition to the execute pipeline.

mksglu··on Show HN: Context Mode – 315 KB of MCP output becomes 5.4 KB in Claude Code
That means a lot, thank you! Would love to hear your feedback once you try it — and an upvote would be much appreciated if you find it useful
mksglu··on Show HN: Context Mode – 315 KB of MCP output becomes 5.4 KB in Claude Code
Sure, thank you for your comment!
mksglu··on Show HN: Context Mode – 315 KB of MCP output becomes 5.4 KB in Claude Code
Great questions.

--

On lossy compression and the "unsurfaced signal" problem:

Nothing is thrown away. The full output is indexed into a persistent SQLite FTS5 store — the 310 KB stays in the knowledge base, only the search results enter context. If the first query misses something, you (or the model) can call search(queries: ["different angle", "another term"]) as many times as needed against the same indexed data. The vocabulary of distinctive terms is returned with every intent-search result specifically to help form better follow-up queries.

The fallback chain: if intent-scoped search returns nothing, it splits the intent into individual words and ranks by match count. If that still misses, batch_execute has a three-tier fallback — source-scoped search → boosted search with section titles → global search across all indexed content.

There's no explicit "raw mode" toggle, but if you omit the intent parameter, execute returns the full stdout directly (smart-truncated at 60% head / 40% tail if it exceeds the buffer). So the escape hatch is: don't pass intent, get raw output.

On token counting:

It's a bytes/4 estimate using Buffer.byteLength() (UTF-8), not an actual tokenizer. Marked as "estimated (~)" in stats output. It's a rough proxy — Claude's tokenizer would give slightly different numbers — but directionally accurate for measuring relative savings. The percentage reduction (e.g., "98%") is measured in bytes, not tokens, comparing raw output size vs. what actually enters the conversation context.

mksglu··on Show HN: Context Mode – 315 KB of MCP output becomes 5.4 KB in Claude Code
Thanks! Context Mode is a standard MCP server, so it works with any client that supports MCP — including Codex and opencode.

Codex CLI:

  codex mcp add context-mode -- npx -y context-mode
Or in ~/.codex/config.toml:

  [mcp_servers.context-mode]
  command = "npx"
  args = ["-y", "context-mode"]
opencode:

In opencode.json:

  {
    "mcp": {
      "context-mode": {
        "type": "local",
        "command": ["npx", "-y", "context-mode"],
        "enabled": true
      }
    }
  }
We haven't tested yet — would love to hear if anyone tries it!
Page 1 of 2Next →