67 karma · joined December 11, 2017
https://www.consensus.tools
https://clawhub.ai/u/kaicianflone
It seems like governments and large corporations already struggle with governance in general, so I’m skeptical that AI governance will be solved quickly.
Do you expect the next few years to be defined by painful trial and error? I could imagine billion dollar companies disappearing almost overnight due to litigation, compliance failures, security incidents, or outright fraud enabled by AI-assisted development and weak governance.
Or are these risks overstated?
https://www.anthropic.com/engineering/multi-agent-research-s...
Their findings suggest multi-agent systems result in better performance attributed mostly to token usage (80% of variance).
Rumors are that Elon gets spaceX to buy tesla so tele-operated Optimus robots do the hard space work from now on. Not a bad idea per se but I’m not educated on the topic. Curiosity has me asking if we really want humans to go to mars or in space at all.
Just like on a team of high performers, there are a million ways to skin a grape.
In my research, I've found that models perform better when they operate as a collective system with reputation, incentives, and accountability instead of isolated oracles answering alone.
Agreement, dissent, and correctness should all carry rewards and consequences. Just like in real life.
Collective machine intelligence, not AGI.
It's expensive, but it's also naive to believe a single model will consistently produce profoundly correct answers to profoundly novel questions.
import { consensus } from "@consensus-tools/wrapper";
const safeSend = consensus(sendEmail, {
reviewers: [humanReviewer, aiSafetyReviewer],
strategy: { mode: "unanimous" },
hooks: { onBlock: (ctx) => audit.log("blocked", ctx) },
});
await safeSend({ to: "user@example.com", body: "Hello" });
The call to sendEmail doesn't execute until every reviewer votes. Strategy modes handle the consensus logic (unanimous, majority, weighted, etc.), and guards can ALLOW, BLOCK, REWRITE, or escalate to REQUIRE_HUMAN before anything fires.The monorepo has 9 built-in policy types and 7 guard types designed so you can drop governance into an existing agent system without rewriting your orchestration.
Repo's at github.com/consensus-tools if you want to poke around.
The coordination and consistency problems the paper describes are what the monorepo is designed around. Giving agents auditable stake in decisions. Happy to share more if anyone’s working in this space.
But now, thanks to makerworld and 3D printers, I have a stand with integrated neodymium magnets for home that puts them to sleep on my desk and nightstand.
I’m equally surprised I had to print something Apple doesn’t sell and Apple hasn’t improved the design for what feels like a decade (other than USB-C and lossless and now old H2)
The unstable tier is the key result. Models that get it right 70–80% of the time are not “almost correct.” They are nondeterministic decision functions. In production that’s worse than being consistently wrong.
A single sampled output is just a proposal. If you treat it as a final decision, you inherit its variance. If you treat it as one vote inside a simple consensus mechanism, the variance becomes observable and bounded.
For something this trivial you could:
-run N independent samples at low temperature
-extract the goal state (“wash the car”)
-assert the constraint (“car must be at wash location”)
-reject outputs that violate the constraint
-RL against the "decision open ledger"
No model change required. Just structure.The takeaway isn’t that only a few frontier models can reason. It’s that raw inference is stochastic and we’re pretending it’s authoritative.
Reliability will likely come from open, composable consensus layers around models, not from betting everything on a single forward pass.
My wife is trilingual, so now I’m tempted to use her as a manual red team for my own guardrail prompts.
I’m working in LLM guardrails as well, and what worries me is orchestration becoming its own failure layer. We keep assuming a single model or policy can “catch” errors. But even a 1% miss rate, when composed across multi-agent systems, cascades quickly in high-stakes domains.
I suspect we’ll see more K-LLM architectures where models are deliberately specialized, cross-checked, and policy-scored rather than assuming one frontier model can do everything. Guardrails probably need to move from static policy filters to composable decision layers with observability across languages and roles.
Appreciate you publishing the methodology and tooling openly. That’s the kind of work this space needs.
Agents compete and review, then the best proposal gets promoted to me as a PR. I stay in control and sync back to the fork.
It’s not auto-merge. It’s structured pressure before human merge.
You define a policy (majority, weighted vote, quorum), set the confidence level you want, and run enough independent inferences to reach it. Cost is visible because reliability just becomes a function of compute.
The question shifts from “is this output correct?” to “how much certainty do we need, and what are we willing to pay for it?”
Still early, but the goal is to make accuracy and cost explicit and tunable.
With consensus.tools we split things intentionally. The OSS CLI solves the single user case. You can run local "consensus boards" and experiment with policies and agent coordination without asking anyone for permission.
Anything involving teams, staking, hosted infra, or governance sits outside that core.
Open source for us is the entry point and trust layer, not the whole business. Still early, but the federation vs stadium framing is useful.
Instead of just wiring agents together, I require stake and structured review around outputs. The idea is simple: coordination without cost trends toward noise.
Curious how entire.io thinks about incentives and failure modes as systems scale.
The idea is to let multiple agents propose, critique, and stake on decisions before a single action is taken, rather than letting one model silently decide. It’s model-agnostic and runs locally, with no blockchain or financial layer involved.
I’m mostly exploring whether adding explicit disagreement and cost at decision time actually improves outcomes in high-stakes or automated workflows.
https://github.com/consensus-tools/consensus-tools
I've also created an AgentSkill to interact with the cli:
I’m working on an open source CLI that experiments with this at a local, off-chain level. It lets maintainers introduce cost, review pressure, or reputation at submission time without tying anything to money or blockchains. The goal is to reduce low-quality contributions without financializing the workflow or creating new attack surfaces.