HNHacker News
TopNewBestAskShowJobs

showhz

5 karma · joined July 13, 2026

submissionscomments
showhz··on [dead]
GitHub: https://github.com/oooscoos/Benzi

Demo: https://varianttech.net/demo

Benchmarks: https://varianttech.net/benchmark

Roughly speaking, the way current AI coding agents/harnesses work is by either:

a) Pulling in appropriate text snippets of code across multiple files and handing them to the agent, or

b) Parsing code to make high dimenional embeddings to approximate a symptom map, and hand that to the agent.

Both of these approaches skyrocket the token count, add to wall clock time, contribute to context drifting, add to the model's thinking tokens to discover the structure of the program, and then FORGET most of it when Claude Code compacts, or ALL of it if it's a multifile refactoring because all line numbers shift and need re-grepping.

Benzi is built from the ground up to AVOID reading source code in the first place. It supplies the artificial intelligence model deterministic intelligence via tool calls. For example, when a model is about to make a code change, it could query "what functions feed this one?" -- half the time it isn't even necessary because the Benzi compiler already informs it of the blast radius before and after making edits, along with a complete static analysis check.

Benzi Sonnet reads far less source code (9,125 lines) than Claude Code Sonnet (20,704), DeepSeek's harness (43,598), and OpenCode (65K+ LOC -- disqualified due to repeated failure) to accomplish the same tasks faster and cheaper. (benchmark link in comments)

Benzi has truth tiers clearly seperating what can be analyzed with static analysis from what can't -- and then adding a runtime tracer on top to bridge the gap between the two (details in FAQ on github).

It also has several bonus features such as a runtime tracer, syntax & semantic verified edits, context aware model written repro, mid task model upgrade if task is too difficult, and SEVERAL more.

It currently supports Python · JavaScript · TypeScript · Java · C# · C++ · C · Go · Rust · Ruby, and can handle HTML, CSS and JS -- deterministically. Claude Code clicks photos, Benzi resolves winners of CSS rules. The CodeIndex and the MarkupIndex are fairly well tested, and if something isn't working, the model is made aware of it first.

On the benchmarks side, 78.2% SWE-bench Verified for <10¢ a fix (using V4flash). This score is noteable because while the rest of the industry is leaning plugin-heavy and pouring millions of dollars into increasing context window sizes, Benzi's approach might prove to be economically more valuable while improving the model's code writing/comprehenion abilities.

Thanks for reading! please let me know what you think.

showhz··on [dead]
Roughly speaking, the way current AI coding agents/harnesses work is by either: a) Pulling in appropriate text snippets of code across multiple files and handing them to the agent, or b) Parsing code to make high dimensional embeddings to approximate a symptom map, and hand that to the agent. Both of these approaches skyrocket the token count, add to wall clock time, contribute to context drifting, add to the model's thinking tokens to discover the structure of the program, and then FORGET most of it when Claude Code compacts, or ALL of it if it's a multifile refactoring because all line numbers shift and need re-grepping. Benzi is built from the ground up to AVOID reading source code in the first place. It supplies the artificial intelligence model deterministic intelligence via tool calls. For example, when a model is about to make a code change, it could query "what functions feed this one?" -- half the time it isn't even necessary because the Benzi compiler already informs it of the blast radius before and after making edits, along with a complete static analysis check.

Benzi Sonnet reads far less source code (9,125 lines) than Claude Code Sonnet (20,704), DeepSeek's harness (43,598), and OpenCode (65K+ LOC -- disqualified due to repeated failure) to accomplish the same tasks faster and cheaper. (Benchmark details: https://benzi.fly.dev/benchmark)

Truth tiers are maintained with clear and transparent boundaries. Explained in the github readme's FAQ. (https://github.com/oooscoos/Benzi)

It also has several bonus features such as a runtime tracer, self-aware model upgrade mid task if it thinks the job is over its pay grade, context aware model written repro, and SEVERAL more.

It currently supports Python · JavaScript · TypeScript · Java · C# · C++ · C · Go · Rust · Ruby, and can handle HTML, CSS and JS -- deterministically. Claude Code clicks photos, Benzi resolves winners of CSS rules. The CodeIndex and the MarkupIndex are fairly well tested, and if something isn't working, the model is made aware of it first.

On the benchmarks side, 78.2% SWE-bench Verified for <10¢ a fix (using V4flash). This score is notable because while the rest of the industry is leaning plugin-heavy and pouring millions of dollars into increasing context window sizes, Benzi's approach might prove to be economically more valuable while improving the model's code writing/comprehenion abilities.

If you're curious to learn more, click https://benzi.fly.dev/about

showhz··on [dead]
Roughly speaking, the way current AI coding agents/harnesses work is by either: a) Pulling in appropriate text snippets of code across multiple files and handing them to the agent, or b) Parsing code to make high dimensional embeddings to approximate a symptom map, and hand that to the agent. Both of these approaches skyrocket the token count, add to wall clock time, contribute to context drifting, add to the model's thinking tokens to discover the structure of the program, and then FORGET most of it when Claude Code compacts, or ALL of it if it's a multifile refactoring because all line numbers shift and need re-grepping. Benzi is built from the ground up to AVOID reading source code in the first place. It supplies the artificial intelligence model deterministic intelligence via tool calls. For example, when a model is about to make a code change, it could query "what functions feed this one?" -- half the time it isn't even necessary because the Benzi compiler already informs it of the blast radius before and after making edits, along with a complete static analysis check.

Benzi Sonnet reads far less source code (9,125 lines) than Claude Code Sonnet (20,704), DeepSeek's harness (43,598), and OpenCode (65K+ LOC -- disqualified due to repeated failure) to accomplish the same tasks faster and cheaper. (Benchmark details: https://benzi.fly.dev/benchmark)

Truth tiers are maintained with clear and transparent boundaries. Explained in the github readme's FAQ. (https://github.com/oooscoos/Benzi)

It also has several bonus features such as a runtime tracer, self-aware model upgrade mid task if it thinks the job is over its pay grade, context aware model written repro, and SEVERAL more.

It currently supports Python · JavaScript · TypeScript · Java · C# · C++ · C · Go · Rust · Ruby, and can handle HTML, CSS and JS -- deterministically. Claude Code clicks photos, Benzi resolves winners of CSS rules. The CodeIndex and the MarkupIndex are fairly well tested, and if something isn't working, the model is made aware of it first.

On the benchmarks side, 78.2% SWE-bench Verified for <10¢ a fix (using V4flash). This score is notable because while the rest of the industry is leaning plugin-heavy and pouring millions of dollars into increasing context window sizes, Benzi's approach might prove to be economically more valuable while improving the model's code writing/comprehenion abilities.

If you're curious to learn more, click https://benzi.fly.dev/about

showhz··on [dead]
Roughly speaking, the way current AI coding agents/harnesses work is by either:

a) Pulling in appropriate text snippets of code across multiple files and handing them to the agent, or

b) Parsing code to make high dimenional embeddings to approximate a symptom map, and hand that to the agent.

Both of these approaches skyrocket the token count, add to wall clock time, contribute to context drifting, add to the model's thinking tokens to discover the structure of the program, and then FORGET most of it when Claude Code compacts, or ALL of it if it's a multifile refactoring because all line numbers shift and need re-grepping.

Benzi is built from the ground up to AVOID reading source code in the first place. It supplies the artificial intelligence model deterministic intelligence via tool calls. For example, when a model is about to make a code change, it could query "what functions feed this one?" -- half the time it isn't even necessary because the Benzi compiler already informs it of the blast radius before and after making edits, along with a complete static analysis check.

Benzi Sonnet reads far less source code (9,125 lines) than Claude Code Sonnet (20,704), DeepSeek's harness (43,598), and OpenCode (65K+ LOC -- disqualified due to repeated failure) to accomplish the same tasks faster and cheaper. (Benchmark details: https://benzi.fly.dev/benchmark)

"But what if the compiler isn't doing its job right! Wouldn't you mislead the AI model?" - Absolutely. Benzi meticulously takes care of this by having 3 truth tiers. RESOLVED has definite evidence, CANDIDATE is what couldn't be resolved by the static analysis, and OBSERVED is what actually happened during an execution. The artificial intelligence and the determinstic intelligence layers coordinate to reduce source hits where possible, without producing incorrect results for the sake of efficiency.

It also has several bonus features such as a runtime tracer, self-aware model upgrade mid task if it thinks the job is over its pay grade, context aware model written repro, and SEVERAL more.

It currently supports Python · JavaScript · TypeScript · Java · C# · C++ · C · Go · Rust · Ruby, and can handle HTML, CSS and JS -- deterministically. Claude Code clicks photos, Benzi resolves winners of CSS rules. The CodeIndex and the MarkupIndex are fairly well tested, and if something isn't working, the model is made aware of it first.

On the benchmarks side, 78.2% SWE-bench Verified for <10¢ a fix (using V4flash). This score is noteable because while the rest of the industry is leaning plugin-heavy and pouring millions of dollars into increasing context window sizes, Benzi's approach might prove to be economically more valuable while improving the model's code writing/comprehenion abilities.

If you're curious to learn more, click https://benzi.fly.dev/about and check out StallionSwipe. probably the best thing i ever made. It's a Fireship inspired horse tinder app greenfielded entirely in Benzi Opus 4-8 and a little bit v4 flash.

and lastly, please star on github if you like where this is headed!

showhz··on Checkout this coding harness that compiles codebase. 78% SWE-Bench (DS v4flash)
original show HN post here: https://news.ycombinator.com/item?id=49226627
showhz··on Show HN: Try Benzi – A coding harness/agent beating Claude Code itself on Sonnet
hi. rewrote the text for more visibility. Benzi remains the same product:

Benzi is a code intelligence software.

It starts with a tree sitter, and produces a unique Benzi ID for all symbols in arbitrarily large codebases.

With all callflow and dataflow resolved, the agent simply queries the codebase using agentic tools instead of reading though it using traditional RAG approaches or trying to "rank" results using embedding-space approaches.

This affords faster inference, cheaoer prices, and less context drift. Tested over 24 repos, Benzi reads 2x less source code than Claude Code, 3x less than the very recently released DeepSeek Harness, and 6.5x less than OpenCode.

Benzi has several bonus features such as syntax+semantic verified code edits, runtime tracer to track actual execution through the Benzi compiler, context aware model generated repros to fast track testing, etc.

On the benchmarks side, Benzi + DeepSeek V4 Flash scores 78% on SWE-bench Verified (benchmark details on the benchmark page). For comparison, DeepSeek reports 73.7% as the baseline scaffolding number for v4flash and self reports their own score to be 78.6%. HOWEVER. Benzi gets to this parity reading 3x less source code than the deepseek harness. (benchmarks page again)

Please try it out, and let me know what you think!

https://benzi.fly.dev/

https://lnkd.in/gtiv7feC

showhz··on Show HN: Try Benzi – A coding harness/agent beating Claude Code itself on Sonnet
77.4% SWE-bench Verified (#4 on the leaderboard) for less than $30. all runs detailed here https://benzi.fly.dev/benchmark
showhz··on Show HN: Defcon and HN, the only two audiences I might find
what have my ears been blessed with??