So I built Hatice.
295 karma · joined March 7, 2023
So I built Hatice.
Codex proved the concept. Symphony showed the architecture in Elixir/OTP. I wanted the same thing with Claude Code.
So I built Hatice.
What's new:
- Five platform adapters: Claude Code, Gemini CLI, VS Code Copilot, OpenCode, Codex CLI
- Session Continuity: survives context compactions, restores state automatically (~3h sessions vs ~30 min)
- ctx_batch_execute: one call replaces 30+ individual commands
- Three-layer fuzzy search: Porter stemming, trigram substring, Levenshtein
- FTS5 deduplication, background mode, SSL cert auto-detection
Breaking: repo renamed claude-context-mode → context-mode, tools prefixed with ctx_. Install (Claude Code): /install context-mode
Install (all platforms): npm install -g context-mode
GitHub: https://github.com/mksglu/context-mode
Release notes: https://github.com/mksglu/context-mode/releases/tag/v1.0.0
Four adapters are in beta — PRs welcome. The codebase is designed for Agentic Engineering: point your coding agent at it and let it propose fixes.Thank you to everyone who commented, starred, and opened issues on the original post. You made this release happen.
SecLaw runs 4 Docker containers with strict boundaries: non-root,
cap_drop ALL, read-only filesystem, 512MB/1CPU limits per container,
zero inbound ports via Cloudflare Tunnel. API keys are sealed per service,
not shared across containers.
The interesting part is multi-agent auto-routing. You install agents as
templates (npx seclaw add inbox-agent, npx seclaw add research-agent)
and they stack onto one Telegram bot. The LLM routes each message to the
right capability — email questions go to Inbox, lead questions go to Sales.
Every response shows which agent answered.
Architecture: Node.js agent + Inngest for scheduled workflows + Desktop
Commander (MCP server, read-only) + Cloudflare Tunnel. All orchestrated
by a single CLI command.
Setup is `npx seclaw` — walks you through LLM provider, API key, Telegram
token, and runs docker compose up. 60 seconds, no YAML editing.The core idea: every MCP tool call dumps raw data into your 200K context window. Context Mode spawns isolated subprocesses — only stdout enters context. No LLM calls, purely algorithmic: SQLite FTS5 with BM25 ranking and Porter stemming.
Since the last post we've seen 228 stars and some real-world usage data. The biggest surprise was how much subagent routing matters — auto-upgrading Bash subagents to general-purpose so they can use batch_execute instead of flooding context with raw output.
Source: https://github.com/mksglu/claude-context-mode Happy to answer any architecture questions.
Your other MCP servers still run. Context Mode just gives the model a more context-efficient way to process their results when the output would be large.
npm install -g context-mode@latest
If you're on the plugin install, re-run:
/plugin marketplace add mksglu/claude-context-mode
/plugin install context-mode@claude-context-mode
Then restart Claude Code. Sorry about that.On benchmarking: This is the experiment I most want to see. The hypothesis: context-mode benefits smaller models disproportionately — a 32K model with clean context could outperform a 200K model drowning in raw tool output. Would love to see SWE-bench results with context-mode on vs. off across model tiers.
The SQLite database is ephemeral — stored in the OS temp directory (/tmp/context-mode-{pid}.db) and scoped to the session process. Nothing persists after the session ends. For sensitive data masking specifically: right now the raw data never leaves the sandbox (it stays in the subprocess or the temp SQLite store), and only stdout summaries enter the conversation. But a dedicated redaction layer (regex-based PII stripping before indexing) is an interesting idea worth exploring. Would be a clean addition to the execute pipeline.
--
On lossy compression and the "unsurfaced signal" problem:
Nothing is thrown away. The full output is indexed into a persistent SQLite FTS5 store — the 310 KB stays in the knowledge base, only the search results enter context. If the first query misses something, you (or the model) can call search(queries: ["different angle", "another term"]) as many times as needed against the same indexed data. The vocabulary of distinctive terms is returned with every intent-search result specifically to help form better follow-up queries.
The fallback chain: if intent-scoped search returns nothing, it splits the intent into individual words and ranks by match count. If that still misses, batch_execute has a three-tier fallback — source-scoped search → boosted search with section titles → global search across all indexed content.
There's no explicit "raw mode" toggle, but if you omit the intent parameter, execute returns the full stdout directly (smart-truncated at 60% head / 40% tail if it exceeds the buffer). So the escape hatch is: don't pass intent, get raw output.
On token counting:
It's a bytes/4 estimate using Buffer.byteLength() (UTF-8), not an actual tokenizer. Marked as "estimated (~)" in stats output. It's a rough proxy — Claude's tokenizer would give slightly different numbers — but directionally accurate for measuring relative savings. The percentage reduction (e.g., "98%") is measured in bytes, not tokens, comparing raw output size vs. what actually enters the conversation context.
Codex CLI:
codex mcp add context-mode -- npx -y context-mode
Or in ~/.codex/config.toml: [mcp_servers.context-mode]
command = "npx"
args = ["-y", "context-mode"]
opencode:In opencode.json:
{
"mcp": {
"context-mode": {
"type": "local",
"command": ["npx", "-y", "context-mode"],
"enabled": true
}
}
}
We haven't tested yet — would love to hear if anyone tries it!